Computer Science & Engineering: An International Journal (CSEIJ), Vol.8, No.1, February 2018
DISEASE PREDICTION USING MACHINE LEARNING OVER BIG DATA
Vinitha S, Sweetlin S, Vinusha H and Sajini S
Computer Science and Engineering, S.A. Engineering College, India
ABSTRACT
Due to big data progress in biomedical and healthcare communities, accurate study of medical data
benefits early disease recognition, patient care and community services. When the quality of medical data
is incomplete the exactness of study is reduced. Moreover, different regions exhibit unique appearances of
certain regional diseases, which may results in weakening the prediction of disease outbreaks. In the
proposed system, it provides machine learning algorithms for effective prediction of various disease
occurrences in disease-frequent societies. It experiment the altered estimate models over real-life hospital
data collected. To overcome the difficulty of incomplete data, it use a latent factor model to rebuild the
missing data. It experiment on a regional chronic illness of cerebral infarction. Using structured and
unstructured data from hospital it use Machine Learning Decision Tree algorithm and Map Reduce
algorithm. To the best of our knowledge in the area of medical big data analytics none of the existing work
focused on both data types. Compared to several typical estimate algorithms, the calculation exactness of
our proposed algorithm reaches 94.8% with a convergence speed which is faster than that of the CNNbased
unimodal disease risk prediction (CNN-UDRP) algorithm.
KEYWORDS
Big data analytics, machine learning, healthcare.
1. INTRODUCTION
With the advance of big data analytics equipment, more devotion has been paid to disease
expectation from the perception of big data inquiry, various explores have been conducted by
choosing the features mechanically from a large number of data to improve the truth of menace
classification rather than the formerly selected physiognomies. However, those prevailing work
mostly measured structured data. Thus, risk organization based on big data analysis, the following
tasks remain: How should the mislaid data be lectured? How should the main chronic diseases in
a positive county and the main faces of the disease in the region be gritty? How can big data
analysis expertise be used to estimate the disease and generate a better method?
To solve these problems, it see the structured and unstructured data in healthcare field to assess
the risk of disease. First, the system use Decision tree map algorithm to generate the pattern and
causes of disease. It clearly shows the diseases and sub diseases. Second, by using Map Reduce
algorithm for partitioning the data such that a query will be analyzed only in a specific partition,
which will increase the operational efficiency but reduce query retrieval time. Map reducing
algorithm is used for partitioning the medical data based on the output of Decision Tree map algorithm. Compared to several typical prediction algorithms, the prediction accuracy of our
proposed algorithm increases.
No comments:
Post a Comment