Today, data lakes are springing up here and there. And with that, the composition structure of data lakes is changing. As more and more data are moving towards cloud, data lakes are shifting focus towards cutting edge sources, like NoSQL, while cloud data warehouses are emerging across hybrid deployments.
A humongous amount of data is being churned out on digital platform each day. IBM says as much as 2.5 quintillion bytes of data is created on a daily basis. Now, this ever-expanding amount of data needs for proper storage system – for that, data lakes have been constructed to hold data in its raw form. In these vast storehouses, data remain mostly in their unstructured state, which is pulled out by data scientists to remodel and transform them into versatile data sets for future use.
Data is the buzzword. It is conquering the world, but who conquers data: the companies that use them or the servers in which they are stored?
Let’s usher you into the fascinating world of data, and data governance. FYI: the latter is weaving magic around the Business Intelligence community, but to optimize the results to the fullest, it needs to depend heavily on a single factor, i.e. efficient data management. For that, highly-skilled data analysts are called for – to excel on business analytics, opt for Business Analytics Online Certification by DexLab Analytics. It will feed you in the latest trends and meaningful insights surrounding the daunting domain of data analytics.
A substantial part of the Apache project, Hadoop is an open source, Java-based programming software framework that is used for storing data and running applications on different clusters of commodity hardware. Be it any kind of data, Hadoop acts as a massive storage unit backed by gargantuan processing power and an ability to tackle virtually countless tasks and jobs, simultaneously.
In this blogpost, we are going to discuss top 10 Hadoop interview questions – cracking these questions may help you bag the sexiest job of this decade.
What are the components of Hadoop?
There are 3 layers in Hadoop and they are as follows:
Storage layer (HDFS) – Also known as Hadoop Distributed File System, HDFS is responsible for storing various forms of data as blocks of information. It includes NameNode and DataNode.
Batch processing engine (MapReduce) – For parallel processing of large data sets across a standard Hadoop cluster, MapReduce is the key.
Resource management layer (YARN) – Yet Another Resource Negotiator is the powerful processing framework in Hadoop system that keeps a check on the resources.
Why is Hadoop streaming?
Hadoop distribution includes a generic application programming interface for drawing MapReduce jobs in programming languages like Ruby, Python, Perl, etc. and this is known as Hadoop streaming.
What are the different modes to run Hadoop?
Local (standalone) Mode
Pseudo-Distributed Mode
Fully-Distributed Mode
How to restart Namenode?
Begin by clicking on stop-all.sh and then on start-all.sh
Use multiple nodes and the distcp command to ensure smooth copying of files between HDFS clusters.
What do you mean by speculative execution in Hadoop?
In case, a node executes a task slower, the master node has the ability to start the same task on another node. As a result, the task that finishes off first will be accepted and the other one will be rejected. This entire procedure is known as “speculative execution”.
What is “WAL” in HBase?
Here, WAL stands for “Write Ahead Log (WAL)”, which is a file located in every Region Server across the distributed environment. It is mostly used to recover data sets in case of mishaps.
How to do a file system check in HDFS?
FSCK command is your to-go option to do file system check in HDFS. This command is extensively used to block locations or names or check overall health of any files.
A block divides the data, physically without taking into account the logical equations. This signifies you can posses a record that originated in one block and stretches over to another. On the other hand, InputSplit includes the logical boundaries of records, which are crucial too.
Why should you use Storm for Real-Time Processing?
Easy to operate – simple operating system makes it easy
Fast processing – it can process around 100 messages per second per node
Fault detection – it can easily detect faults and restarts functional attributes
Scores high on reliability – expect execution of each data unit at least for once
High scalability – it operates throughout clusters of machines
To learn more about Data Analyst with Advanced excel course – Enrol Now. To learn more about Data Analyst with R Course – Enrol Now. To learn more about Big Data Course – Enrol Now.
To learn more about Machine Learning Using Python and Spark – Enrol Now. To learn more about Data Analyst with SAS Course – Enrol Now. To learn more about Data Analyst with Apache Spark Course – Enrol Now. To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.
Google is strengthening its artificial intelligence base, including China.
And it is so doing by establishing a new AI research center in Beijing. Google is digging deep into China, where it contravened the government in 2010 committing a spectacularly principled act of self-sabotage by refusing to self-censor search content and later found most of its services to be blocked. The company’s decision to return back to China is more about safeguarding its future, and acknowledging the supreme importance of technology’s most competitive field: AI.
In marketing, the analysis of data is a highly established one but the marketers nowadays have a massive amount of public and proprietary data about the preferences, usage, and behavior of a customer. The term ‘big data’ points out to this data explosion and the capability to use the data insights to make informed decisions. Understanding the potential of big data presents various technical challenges but it also needs executive talent devoted to applying the solutions of big data. Today, the marketers are widely embracing big data and are confident in their use of analytics tools and techniques. Let us learn about the ways in which Big data and analytics can improve the marketing efforts of various businesses around the around.
Locating Prospective Customers
Previously, marketers had to frequently make guesses as to which sector of population comes under their ideal market segment but this is no longer the scenario today. The companies can exactly see who is buying and even extract more details about them with the help of big data. The other details include which buttons they generally click while on a website, which websites they visit frequently, and which social media channels they utilize.
Tracking Impact and ROI
Many retailers have introduced loyalty card systems that track the purchases of a customer, but these systems can also track which promotions and incentives are most effective in encouraging a group of customers or a single customer to make another purchase.
Handling Marketing Budgets
Because big data allows companies to optimize and monitor their marketing campaigns for performance, this implies they can allocate their budget for marketing for the highest return-on-investment (ROI).
Personalizing Offers in Real-Time
Marketers can personalize their offers to customers in real time with the combination of big data and machine learning algorithms. Think about the Amazon’s “customers also bought” section or the recommended list of TV shows and movies from Netflix. The organizations can personalize what promotions and products a particular customer views, even down to sending personalized offers and coupons to the mobile phone of a customer when he walks into a physical location. The role of Personalized Merchandising in the ecommerce industry will continue to increase in the years to come.
Improvement in Market Research
Companies can conduct quantitative and qualitative market research much more inexpensively and quickly than ever before. The tools for online survey mean that customer feedback and focus groups are inexpensive and easy to implement, and data analytics make the results easier to take action.
Prediction of Buyer Behavior and Sales
For the past several years, sales teams, in order to rate their hottest leads, have made use of lead scoring. But, with the help of predictive analytics, a model can be generated and it can successfully predict sales and buyer behavior.
Enhanced Content Marketing
Previously, the return-on-investment for a blog post used to be highly difficult to measure. But, with the help of big data and analytics, the marketers can effortlessly analyze which pieces of content are highly effective at moving leads via a sales and marketing funnel. Even a small firm can afford to use tools for implementing content scoring which can highlight the content pieces that are highly responsible for closing sales.
Optimize Customer Engagement
Data can provide more information about your customers which includes who they are, what they want, where they are, how often they purchase on your site, and how, when they prefer to be contacted, and various other major factors. The organizations can also examine how users interact not only with their website, but also their physical store to enhance the experience of the user.
Tracking Competitors
New tools for social monitoring have made it easy to gather and examine data about the competitors and their efforts regarding marketing as well. The organizations that can utilize this data will have a distinct competitive advantage.
Managing Reputation
With the help of big data, organizations can monitor their brand mentions very easily across different social channels and websites to locate unfiltered testimonials, reviews, and opinions about their company and products. The savviest can also utilize social media to offer service to the customers and create a trustworthy brand presence.
Marketing Optimization
It is quite difficult to track direct ROI and impact with traditional advertising. But, big data can help organizations to make optimal marketing buys across various channels and to optimize their marketing efforts continuously through analysis, measurement, and testing.
What is Needed for Big Data?
At this point, talent and leadership are the major things that big data needs. In most of the companies, the marketing teams don’t have the right talent in place to leverage analytics and data. Apart from people who possess analytical skills to understand the capability of big data and where to use it, companies require data scientists who can extract meaningful insights from data and the technologists who can develop include new technologies. Due to this, there is a high demand for experienced analytics talent today.
Big Data Limitations for Marketing
In spite of all the promise, there exist certain limits to the usefulness of big data analytics in its present state. Among them, the major one is the major one is the analytics tools’ and techniques’ complex “black box” nature which makes it hard to trust and interpret the output of the approaches of big data and to assure others of the accuracy and value of the insights generated by the tools. The difficulty of gathering and understanding data also limits the capability of marketing companies to more fully leverage big data. Beyond this, the marketers are identifying many hurdles to expanding their utilization of big data tools and they include lack of sufficient technology investment, the inability of senior team members to leverage big data tools for decision-making, and the lack of credible tools for measuring effectiveness.
Conclusion
Cloud computing is also playing a major role in marketing with the Cloud Marketing process. Cloud Marketing is a process that outlines the efforts of a company to market their services and goods online via integrated digital experiences. Once the data analytics tools become available and accessible to even the smallest businesses, there will be a much higher impact of big data on the marketing sector as there will be much broader utilization of data analytics. This can only be a boon as organizations enhance their marketing and reach their customers in innovative and new ways.
Author’s Bio: Savaram Ravindra was born and raised in Hyderabad, popularly known as the ‘City of Pearls’. He is presently working at Mindmajix.com. His previous professional experience includes Programmer Analyst at Cognizant Technology Solutions. He holds a Masters degree in Nanotechnology from VIT University. He can be contacted at savaramravindra4@gmail.com. Connect with him also on LinkedIn and Twitter.
Interested in a career in Data Analyst?
To learn more about Data Analyst with Advanced excel course – Enrol Now. To learn more about Data Analyst with R Course – Enrol Now. To learn more about Big Data Course – Enrol Now.
To learn more about Machine Learning Using Python and Spark – Enrol Now. To learn more about Data Analyst with SAS Course – Enrol Now. To learn more about Data Analyst with Apache Spark Course – Enrol Now. To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.
Microsoft Excel is a smorgasbord of information – it is a staple technology tool in any business environment. Whether you are crunching business data, organizing client sales inventory or planning an office event, Excel is arguably the most powerful tool entrenched across multiple business domains worldwide.
Aspiring professionals contemplating to make an entry into the workplace are required to excel on Excel tricks – Excel Dashboards Training Pune from DexLab Analytics is a promising gateway to your dreams!
More is always better, isn’t it? But does it always holds true, especially when it comes to customer data? Maybe not, because business is all about extracting meaningful insights from data, and if that cannot be acted upon then it is of no good.
Recently, Accenture concluded that one of the greatest challenges that marketers face nowadays is to discover the right ways to turn data into productive insights and then into action. For that, you would need analytics professionals who do know how to collect, store and integrate information, while mastering the technology aspect.
Want to get to the core of understanding risk within various business frameworks? The answer is Risk Analytics. This new breed of data analytics facilitates organizations in precisely defining, recognizing and managing their risk, and its need is going to increase in the coming few years. New developments in risk analyticsare gaining limelight and bringing a notable transformation in the market, while enhancing its overall capability.
Recently, a team of analysts had eureka moment – they introduced a new concept of real-time risk analytics – it is nothing but a modern, more advanced version of traditional risk analytics methods. Here, the prediction is based on real-time data – it processes, examines and determines risk all on a real-time basis – hence top notch financial institutions are putting real-time risk analytics to best use to manage and mitigate associated risks. Several asset management, portfolio management and hedge fund firms, and investment banks are relying on this mode of risk analytics to modify their operating principles to play in accordance with investment and market shifts.
The root cause for the Financial Crisis which stormed the globe in 2008 was the Sub-prime crisis which appeared in USA during late 2006. A sub-prime lending practice started in USA during 2003-2006. During the later parts of 2003, the housing sector started expanding and housing prices also increased. It has been shown that the housing prices were growing exponentially at that time. As a result, the housing prices followed a super-exponential or hyperbolic growth path. Such super-exponential paths for asset prices are termed as ‘bubbles’ So USA was riding a Housing price bubble. Now the bankers, started giving loans to the sub-prime segments. This segment comprised of customers who hardly had the eligibility to pay back the loans. However, since the loans were backed by mortgages bankers believed that with housing price increases the they could not only recover the loans but earn profits by selling off the houses. The expectations made by the bankers that asset prices always would ride the rising curve was erroneous. Hence, when the housing prices crashed the loans were not recoverable. Many banks sold off these loans to the investment banks who converted the loans into asset based securities. These assets based securities were disbursed all over the globe by the investments banks, the largest being done by Lehmann Brothers. When the underlying assets went valueless and the investors lost their investments, many of the investment banks collapsed. This caused the Financial Crisis and a huge loss of investors and tax-payers wealth. The involvement of Systematically Important Financial Institutions (SIFIs) and Globally Systematically Important Financial Institutions (G-SIFIs) into the frivolous lending process had amplified the intensity and the exposure of the crisis.
SYSTEMATICALLY IMPORTANT FINANCIAL INSTITUTIONS AND THEIR ROLE IN SYSTEMIC STABILITY
A Systematically Important Financial Institution (SIFI) is a bank, insurance company, or other financial institutions whose failure might trigger a financial crisis.
If a SIFI has the capacity to bring in a recession across the globe then it is known as a Globally Systematically Important Financial Institution (G-SIFI). The Basel Committee follows an indicator based approach for assessing the systematic importance of the G-SIFIs. The basic tenets of this approach are:
The BASEL committee is of the view that the global systemic importance should be measured in terms of the impact that a failure of a bank can have on the global financial system and wider economy rather than the risk that the failure can occur. So, the concept is more of a global, system wide, loss given default (LGD) concept rather than a probability of default (PD) problem.
The indicators reflect the following metrics: size of banks, their interconnectedness, the lack of availability of substitutable or financial institution infrastructure for provided services, their global activity, their complexity etc. Each of these are defined as:
(i) Cross-Jurisdiction: The indicator captures the global footprints of the banks. This indicator is divided into two activities: Cross Jurisdictional claims and Cross Jurisdictional liabilities. These two indicators measure the banks activities outside its home relative to overall activity of other banks’ in the sample. The greater the global reach of the bank, the more difficult is it to coordinate its resolution and the more widespread the spill over effects from its failure.
(ii) Size: Size of a bank is measured using the total exposure that it has globally. This is the exposure measure used to calculate Leverage ratio. BASEL III paragraph 157 uses a particular definition of exposure for this purpose. The score of each bank for this criterion is calculated as its amount of total exposure divided by the sum of total exposures of all banks in the sample.
(iii) Interconnectedness: Financial distress at one institution can materially raise the likelihood of distress at other institutions given the contractual obligations in which the firms operate. Interconnectedness is defined in terms of the following parameters: (a) Inter-financial system assets (b) Inter-financial system liabilities (c) The degree to which a bank funds itself from the other financial systems.
(iv) Complexity: The systemic impact of a bank’s distress or failure is expected to be positively related to its overall complexity. Complexity includes: business, structural and operational complexity. The more complex the bank is the greater are the costs and time needed to resolve the banks.
Given these characteristics, it was important to apply different restrictions to keep the lending practices of the banks under control. Frivolous lending done by such SIFIs had resulted in the financial crisis 2008-09. Post the crisis, regulators became more vigilant about maintaining appropriate reserves for banks to survive macroeconomic stress scenarios. Three major sources of risks to which banks are exposed to are: 1. Credit Risk 2. Market Risk 3. Operational Risk. Several regulations
have been imposed on banks to ensure that they are adequately capitalised. The major regulatory requirements to which banks need to be compliant with are:
BASEL 2. Dodd Frank Act Stress Testing 3. Comprehensive Capital Adequacy Review.
Before looking into the Regulatory frameworks and their impact on the Credit Risk modelling, let us form an understanding of the framework of the Bank Capital.
The bank’s capital structure is comprised of two main components: 1. Equity Capital of Banks 2. Supplementary capital of banks. The Equity capital of banks are the purest form of banking capital. This is true or the actual capital that a bank has and it has been raised from the shareholders. The supplementary capital of banks comprises of estimated capital such as allowances, provisions etc. This portion of the capital can easily be tampered by the management to meet undue shareholders expectations or unnecessarily over reserve capital. Thus, there are strong capital norms and regulations around the supplementary capital. The two tiers of capital are: Tier1 and Tier2 capital. Tier1 capital is also decomposed into two parts: Tier1 Common capital and Tier1 capital.
Tier1 common capital = Common shareholder’s equity-goodwill-Intangibles. Goodwill and intangibles are no physical capital. In scenarios, where the goodwill and intangible assets are stressed, the capital in the banks would deteriorate. Therefore, they cannot be added to the company’s tier1 capital. Only the core or the physical amount of capital present in the bank account is the capital.
Tier1 Capital = Total Shareholders’ equity (Common + Preffered stocks) -goodwill -intangibles + Hybrid securities.
Tier 1 is the core equity capital for the bank. The components of Tier1 capital are common across all geographies for the banking system. Equity capital includes issued and fully paid equities. This is the purest form of capital that the bank has.
Tier2 Capital: tier 2 capital comprises of estimated reserves and provisions. This is the part of capital which is used to cushion against expected losses. Tier 2 capital has the following composition:Tier 2 = Subordinated debts +Allowances for Loans and lease losses + Provisions for bad debts -> This portion of the capital is reserved out of profits. Hence,
managers always try to under report these parameters to meet shareholder’s expectations. However, under reserving often poses the chances of bankruptcies or regulatory penalties. Total Capital of a Bank = Tier 1 capital + Tier 2 Capital
Every bank faces three main types of risks: 1. Credit risk 2. Market Risk 3. Operational risk. Credit Risk is the risk that arises from lending out funds to borrowers, given their chances of defaulting on loans. Market Risk is the risk that the bank faces due to market fluctuations like stock price changes, interest rate risk and price level fluctuation etc. Operational risk occurs as a failure of the operational processes. The exposure of the banks to these risks differ from bank to bank. So the capital that they to set aside would differ based on the exposure to risk. Therefore, regulators have defined a metric called Risk Weighted Assets (RWA) to identify the exposure of the bank’s assets to risk. Every bank must keep aside their capital relative to the exposure of their asset to risk. The biggest advantage of RWAs is that they not only include On-balance sheet items but off-balance sheet items as well. Banks need to maintain their Tier1 common capital, tier1 capital and tier2 capital relative to their RWAs. Thus, arises the Capital ratios.
Total RWA = RWA for Credit Risk + RWA for Market Risk + RWA for Operational Risk
Tier1 Common Capital Ratio = tier1 common capital / RWA (CR + MR + OR)
Tier1 Capital Ratio = Tier1 Capital / RWA (CR+MR+OR)
Total Capital Ratio = Total capital/ RWA(CR+MR+OR)
Leverage Ratio = Tier1 Capital / Firms consolidated assets
Regulators require some critical cut-offs for each of these ratios:
Tier1 Common Capital Ratio > = 2% all times
Tier1 Ratio >= 4% all times
Tier 2 capital cannot exceed Tier1 capital
Leverage ratio > = 3% of all times.
In the next blog we explore how the credit risk models help in ensuring the capital adequacy of the banks and in the business risk management.
To learn more about Data Analyst with Advanced excel course – Enrol Now. To learn more about Data Analyst with R Course – Enrol Now. To learn more about Big Data Course – Enrol Now.
To learn more about Machine Learning Using Python and Spark – Enrol Now. To learn more about Data Analyst with SAS Course – Enrol Now. To learn more about Data Analyst with Apache Spark Course – Enrol Now. To learn more about Data Analyst with Market Risk Analytics and Modelling Course – Enrol Now.