Tuesday, 13 March 2018

Computer Science & Engineering: An International Journal (CSEIJ), Vol.8, No.1, February 2018

DISEASE PREDICTION USING MACHINE LEARNING OVER BIG DATA


Vinitha S, Sweetlin S, Vinusha H and Sajini S 

Computer Science and Engineering, S.A. Engineering College, India 

ABSTRACT

Due to big data progress in biomedical and healthcare communities, accurate study of medical data benefits early disease recognition, patient care and community services. When the quality of medical data is incomplete the exactness of study is reduced. Moreover, different regions exhibit unique appearances of certain regional diseases, which may results in weakening the prediction of disease outbreaks. In the proposed system, it provides machine learning algorithms for effective prediction of various disease occurrences in disease-frequent societies. It experiment the altered estimate models over real-life hospital data collected. To overcome the difficulty of incomplete data, it use a latent factor model to rebuild the missing data. It experiment on a regional chronic illness of cerebral infarction. Using structured and unstructured data from hospital it use Machine Learning Decision Tree algorithm and Map Reduce algorithm. To the best of our knowledge in the area of medical big data analytics none of the existing work focused on both data types. Compared to several typical estimate algorithms, the calculation exactness of our proposed algorithm reaches 94.8% with a convergence speed which is faster than that of the CNNbased unimodal disease risk prediction (CNN-UDRP) algorithm.

KEYWORDS

Big data analytics, machine learning, healthcare.  

1. INTRODUCTION  

With the advance of big data analytics equipment, more devotion has been paid to disease expectation from the perception of big data inquiry, various explores have been conducted by choosing the features mechanically from a large number of data to improve the truth of menace classification rather than the formerly selected physiognomies. However, those prevailing work mostly measured structured data. Thus, risk organization based on big data analysis, the following tasks remain: How should the mislaid data be lectured? How should the main chronic diseases in a positive county and the main faces of the disease in the region be gritty? How can big data analysis expertise be used to estimate the disease and generate a better method? 
To solve these problems, it see the structured and unstructured data in healthcare field to assess the risk of disease. First, the system use Decision tree map algorithm to generate the pattern and causes of disease. It clearly shows the diseases and sub diseases. Second, by using Map Reduce algorithm for partitioning the data such that a query will be analyzed only in a specific partition, which will increase the operational efficiency but reduce query retrieval time. Map reducing algorithm is used for partitioning the medical data based on the output of Decision Tree map algorithm. Compared to several typical prediction algorithms, the prediction accuracy of our proposed algorithm increases.


 


No comments:

Post a Comment