Posts

51% Attack

Image
If a single entity were to acquire control of over half of the mining power on the Bitcoin network, or any blockchain network, it would gain the ability to do certain things that would not typically be possible and this poses a known security risk. While acquiring so much mining power would be very expensive, some have cautioned that because a majority of mining power may be located geographically within the borders of one country—China specifically—this poses a unique risk that should concern policymakers. So how exactly would such an attack work and how big of a risk is it? What could you actually do with a majority of the mining power on Bitcoin or any other proof-of-work cryptocurrency? Bitcoin’s blockchain is a list of all valid transactions that have been made since the network’s inception back in 2009. Anyone can update that list by adding a new block of transactions to the chain, but they have to compete and follow the rules of the protocol. The network relies on these compe...

Tweets Analysis with Python and NLP

Image
Introduction You should be already familiar with the concepts of NLP from our previous post , so today we'll see more useful case of analysis the tweets and classifying them into marketing and non-marketing tweets. We won't get into details of tweets retrieval, this can be done with various packages with Tweepy being the most popular one. Baseline For the purpose of the discussion we already have 2 sets of tweets separated into files and are uploaded into GitHub folder . First we download the datasets, add target column as 1 for marketing tweets and unite the datasets. Then we'll check the baseline classification results, without any pre-processing. We do this so later we could understand whether our changes improve the metrics. We'll be using Random Forest for classification, since it doesn't expect linear features or even features that interact linearly and it can handle very well high dimensional spaces as well as large number of training examples. Plu...

Symmetric and Asymmetric Encryption

Image
What is Encryption and Cryptographic Keys? Encryption is actually an age-old practice dating back to the times of the famous Roman king Caesar, who encrypted his messages using a Caesar cipher. The practice can be viewed as a transformation of information whereby the sender uses plain text, which is then encoded into cipher text to ensure that no eavesdropper interferes with the original plain text. On receiving the encoded message, the intended receiver decrypts it to obtain the original plain text message. Once the transaction data encrypted then it can only be decrypted using the appropriate keys, its called a “Cryptographic keys“. A cryptographic key is a password which is used to encrypt and decrypt information. There are two types of cryptographic keys. They are known as symmetric key and asymmetric key cryptography: symmetric and asymmetric encryption. What is Symmetric Encryption? Symmetric Encryption also called Secret Key Cryptography, it employs the same secret key...

Optimizations of Gradient Descent

Image
Introduction Gradient Descent is one of the most popular technique to optimize machine learning algorithm. We've already discussed Gradient Descent in the past in Gradient descent with Python article, and gave some intuitions toward it's behaviour. We've also made an overview about choosing learning rate hyper-parameter for the algorithm in hyperparameter optimization article. So by now, you should have a fair understanding of how it works. Today we'll discuss different ways to optimize the performance of the algorithm itself. Gradient Descent Variants We've already three variants of the Gradient Descent in Gradient Descent with Python article: Batch Gradient Descent, Stochastic Gradient Descent and Mini-Batch Gradient Descent. What we haven't discussed was problems arising when using these techniques. Choosing a proper learning rate is difficult. A too small learning rate leads to tremendously slow convergence, while a very large learning rate that ca...

Overview of Machine Learning Metrics

Image
Introduction One of the core tasks in building a machine learning model is to evaluate its performance. The usual data science pipeline consists of prototyping a model on some historical data, reaching a satisfying model and deploying it into production, where it will go through further testing on live data. The stages are usually called offline and online evaluations, where the former analyses prototyped model on historical data and the latter the deployed model on live data. Surprisingly to some, evaluation is really hard as good measurement are often vague or infeasible. Also generally statistical models assume that the distribution of data stays the same over time. But in practice, the distribution of data changes constantly, sometimes drastically. This is called distribution drift. One way to detect distribution drift is to continue tracking the model’s performance on the validation metric on live data. That's why any data science project cannot just end after the model i...

Markov chain Monte Carlo with PyMC

Image
Markov Chain Monte Carlo (MCMC) is a technique for generating a sample from a distribution, and it works even if all you have is a non-normalized representation of the distribution. Why does a data scientist care about this? Well, in a Bayesian analysis a non-normalized form of the posterior distribution is super easy to come by, being just the product of likelihood and prior - so MCMC can be used to sample from (essentially simulate) a Bayesian posterior. In python one of the most widely used packages for doing exactly this is called PyMC . What is Markov chain Monte Carlo A Markov Chain is a sequence of RVs {X} each of which will have an observed value from the state space of possible values, {x}. A Markov Process is a sequence of such RVs where the distribution of the initial RV's value is specified (Π0) as well as a Transition Rule (P) which gives the probability to transition from one state to another, P(i,j)=P(Xn+1=xj|Xn=xi), for all pairs of states. Notice P doe...

Natural Language Processing with Python

Image
Introduction Natural language processing, or NLP, is a process of analyzing the text and extracting insights from it. It is used everywhere, from search engines such as Google or Bing , to voice interfaces such as Siri or Cortana . The pipeline usually involves tokenization , replacing and correcting words, part-of-speech tagging , named-entity recognition and classification. In this article we'll be describing tokenization, by using a full example from Kaggle notebook . The full code can be found on GitHub repository . Installation For the purposes of NLP, we'll be using NLTK Python library, a leading platform to work with human language data. It provides easy-to-use interfaces to over 50 corpora and lexical resources such as WordNet, along with a suite of text processing libraries for classification, tokenization, stemming, tagging, parsing, and semantic reasoning, wrappers for industrial-strength NLP libraries. Installing the package is easy using the Python p...