Online inquiry information processing method and system based on big data
Through big data and natural language processing technology, analyzing patient descriptions and generating symptom lists based on disease databases and personal characteristics, the problem of insufficient effectiveness of patient descriptions in online consultations is solved, and the efficiency and accuracy of consultations are improved.
Patent Information
- Application Number
- CN202510216981.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-25
AI Technical Summary
During the online consultation process, the patient's self-described physical condition is limited by language ability, resulting in a decrease in the effectiveness of symptom description and a decrease in communication efficiency and treatment effectiveness.
The online consultation information processing method based on big data is adopted, and the patient description information is analyzed using natural language processing technology, combined with the disease database and personal characteristics, a list of disease symptoms is generated, and the second condition information is confirmed through patient selection to generate online consultation information.
It improves the accessibility and response speed of medical services, enhances the accuracy and completeness of patient information, helps narrow the scope of diagnosis, and improves diagnostic efficiency.
Smart Images

Figure CN120376184A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly to an online consultation information processing method and system based on big data. Background Art
[0002] Online consultation is a remote medical service realized through an Internet platform, which allows patients to conduct health consultations and diagnoses with doctors without having to go to a hospital or clinic in person. Through online consultation, patients can share their symptoms, medical history and other health-related information, and doctors can provide medical advice, diagnosis results and even issue electronic prescriptions based on this information and auxiliary diagnostic tools such as videos and pictures. This service mode not only improves the convenience and accessibility of medical services, but also can relieve the pressure on physical medical institutions to a certain extent, and is particularly suitable for follow-up consultations, chronic disease management and preliminary consultations for minor illnesses.
[0003] However, during the process of online consultation, patients need to organize their own language to describe their physical conditions by themselves, which is often restricted by their own language abilities, thus reducing the effectiveness of symptom description, lowering the communication efficiency of online consultation, and making it difficult to discover other characteristics through online video observation, thereby reducing the effectiveness of online consultation treatment. Summary of the Invention
[0004] The purpose of the present invention is to provide an online consultation information processing method and system based on big data, aiming to improve the accessibility and response speed of medical services, and by effectively utilizing big data, enhancing the accuracy and completeness of information provided by patients.
[0005] To achieve the above purpose, in the first aspect, the present invention provides an online consultation information processing method based on big data, including obtaining basic disease description information input by a patient;
[0006] Using natural language processing technology to parse the basic disease description information and retrieve all relevant disease names in the disease database;
[0007] Retrieving all corresponding disease symptoms in the database based on the disease names, and sorting the disease symptoms according to the correlation scores of personal characteristics in the database, where the personal characteristics include gender, age, and residential area;
[0008] Presenting the sorted list of disease symptoms to the patient, and obtaining second disease information selected by the patient that conforms to the patient's own situation;
[0009] Generating online consultation information based on the basic disease description information and the second disease information.
[0010] Among them, the specific steps of obtaining the basic disease description information input by the patient include:
[0011] Obtain basic medical condition description information;
[0012] Check whether the basic medical condition description information meets the requirements;
[0013] Store the basic medical condition description information that meets the requirements in the database.
[0014] Among them, the specific steps of parsing the basic medical condition description information using natural language processing technology and retrieving all relevant disease names in the disease database include:
[0015] Split the basic medical condition description information into word groups;
[0016] Remove invalid information from the word groups and format them to obtain the recognition text;
[0017] Identify and classify the entity information in the recognition text, where the entity information includes the discomfort site, discomfort feeling description, drug information, and time information;
[0018] Extract the disease core keywords from the entity information;
[0019] Build a disease knowledge base containing various known diseases and their related symptoms;
[0020] Calculate the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity;
[0021] Sort according to the similarity score and retrieve relevant disease names from the knowledge base to obtain a set of disease names.
[0022] Among them, the specific steps of calculating the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity include:
[0023] Use a pre-trained word embedding model to convert each word of the disease core keywords into a vector representation;
[0024] Apply the cosine similarity formula to calculate the symptom similarity.
[0025] Among them, the specific steps of retrieving all corresponding disease symptoms in the database based on the disease name and sorting the disease symptoms according to the relevance score of personal characteristics in the database include:
[0026] Collect the patient's personal information, where the patient's personal information includes but is not limited to gender, age, and residential area;
[0027] Build a characteristic disease database containing various diseases and their patient information;
[0028] Retrieve all patient information corresponding to the disease in the disease database using the set of matched disease names as keywords;
[0029] Calculate the similarity between the patient's personal information and all patient information to screen for similar patient groups;
[0030] Sort the disease symptoms corresponding to the similar patient groups in descending order based on the frequency of occurrence.
[0031] Among them, the specific steps of retrieving all patient information corresponding to the disease in the disease database using the set of matched disease names as keywords include:
[0032] Construct a corresponding SQL query statement based on the determined set of disease names to retrieve relevant patient information;
[0033] Run the SQL query statement in the database management system to obtain all patient information related to the specified disease;
[0034] Remove duplicate or incomplete records from the obtained patient information.
[0035] Among them, the specific steps of sorting the disease symptoms corresponding to the similar patient groups in descending order based on the frequency of occurrence include:
[0036] Create a comprehensive symptom list that includes all symptoms mentioned in the similar patient groups;
[0037] For each symptom, count the number of times it appears within the entire patient group;
[0038] Sort the symptoms in descending order based on the frequency of occurrence.
[0039] Among them, the specific steps of presenting the sorted list of disease symptoms to the patient for them to select the condition information that matches their own situation to obtain the second condition information include:
[0040] Generate a brief description based on the disease symptoms;
[0041] Obtain the first mark of the patient for the brief description that matches their own situation;
[0042] Expand the detailed information corresponding to the brief description with the first mark;
[0043] Obtain the second mark of the patient for the detailed information to obtain the second condition information.
[0044] In a second aspect, the present invention also provides an online consultation information processing system based on big data, including a condition information acquisition module, a disease name retrieval module, a disease symptom sorting module, a screening module, and a consultation information generation module;
[0045] The medical condition information acquisition module is used to obtain the basic medical condition description information input by the patient;
[0046] The disease name retrieval module is used to parse the basic medical condition description information by using natural language processing technology and retrieve all relevant disease names in the disease database;
[0047] The disease symptom sorting module is used to retrieve all corresponding disease symptoms in the database based on the disease name, and sort the disease symptoms according to the relevance score of personal characteristics in the database, where the personal characteristics include gender, age, and residential area;
[0048] The screening module is used to present the sorted list of disease symptoms to the patient and obtain the second medical condition information selected by the patient that conforms to their own situation;
[0049] The consultation information generation module is used to generate online consultation information based on the basic medical condition description information and the second medical condition information.
[0050] The online consultation information processing method and system based on big data of the present invention
[0051] First, obtain the basic medical condition description information input by the patient, which can be achieved through an online form or an intelligent chatbot, allowing the patient to describe in detail their symptoms, the duration of discomfort, and any relevant health background information. Subsequently, use natural language processing technology to parse these basic medical condition descriptions. This process involves multiple levels such as text analysis and semantic understanding, aiming to accurately capture the key points of the patient's condition and convert them into structured data. Next, the system will retrieve all disease names associated with the parsed medical condition description in a pre-established disease database. After determining the list of potential diseases, the system will further search for the corresponding typical symptoms in the database based on each disease name. To make the results more in line with the actual situation of the patient, the system will sort the disease symptoms according to a series of personal characteristics. These personal characteristics include but are not limited to gender, age, residential area, etc., because different populations may show different symptom intensities or types for the same disease. Then, the system presents the sorted list of disease symptoms to the patient, allowing them to select the symptoms that conform to their own situation. This is a key step in collecting the second medical condition information, which allows the patient to provide more accurate feedback on their health status, thereby helping to narrow down the scope of potential diseases. Finally, based on the initially provided basic medical condition description information and the second medical condition information confirmed by the patient, the system generates a detailed online consultation information. This information can be used for subsequent doctor diagnosis reference, or directly provide preliminary health management suggestions through artificial intelligence algorithms. The whole process not only improves the accessibility and response speed of medical services, but also enhances the accuracy and completeness of the information provided by patients through the effective use of big data. Description of the Drawings
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 is a flowchart of the online consultation information processing method based on big data of the present invention.
[0054] Figure 2 is a flowchart of obtaining the basic condition description information input by the patient of the present invention.
[0055] Figure 3 is a flowchart of parsing the basic condition description information using natural language processing technology and retrieving all relevant disease names in the disease database of the present invention.
[0056] Figure 4 is a flowchart of calculating the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity based on the present invention.
[0057] Figure 5 is a flowchart of retrieving all corresponding disease symptoms of a disease in the database based on the disease name and sorting the disease symptoms according to the relevance score of personal characteristics in the database according to the present invention.
[0058] Figure 6 is a flowchart of retrieving all patient information corresponding to the disease in the disease database using the set of matched disease names as keywords according to the present invention.
[0059] Figure 7 is a flowchart of sorting the disease symptoms corresponding to the similar patient group in descending order based on the occurrence frequency according to the present invention.
[0060] Figure 8 is a flowchart of presenting the sorted list of disease symptoms to the patient for them to select the condition information that suits their own situation to obtain the second condition information according to the present invention. Detailed implementation manners
[0061] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.
[0062] The first embodiment
[0063] Please refer to Figures 1 to 8 , the present invention provides an online consultation information processing method based on big data, including:
[0064] S101 Obtain the basic disease description information input by the patient;
[0065] The specific steps include:
[0066] S201 Obtain the basic disease description information;
[0067] At this stage, the system will guide the patient to input their basic disease description information through the user interface. This information includes, but is not limited to, the time when the symptoms appeared, the specific manifestations of the symptoms (such as pain, fever, etc.), the severity of the symptoms, and any relevant medical history information. To ensure the comprehensiveness and accuracy of the information, the system will provide some prompting questions or options for the patient to choose.
[0068] S202 Detect whether the basic disease description information meets the requirements;
[0069] Once the patient submits their basic information, the system will automatically conduct a preliminary analysis and verification of this information to ensure its integrity and reasonableness. In this step, the system will check whether all required fields have been filled, confirm whether the input information is within a reasonable range (for example, the body temperature value should be within the normal physiological range), and determine whether there are logical contradictions (such as the typical symptoms of a certain disease do not match the symptoms described by the patient). If it is found that the information is incomplete or there are doubts, the system will feedback to the patient and request to supplement or correct the relevant information.
[0070] S203 Store the basic disease description information that meets the requirements in the database.
[0071] When the disease description information provided by the patient passes the above detection steps, this data will be encrypted and securely stored in the system's database.
[0072] S102 Use natural language processing technology to parse the basic disease description information and retrieve all relevant disease names in the disease database;
[0073] The specific steps include:
[0074] S301 Split the basic disease description information into word groups;
[0075] In this initial stage, the system first needs to decompose the patient's disease description text into word or phrase groups that are easy to process. This step usually involves the application of word segmentation technology, and corresponding algorithms are adopted according to different language characteristics. For example, in the Chinese environment, methods based on statistical models or deep learning will be used to perform accurate word segmentation operations.
[0076] S302 removes the invalid information in the word group and formats it to obtain the recognized text;
[0077] Next, remove the information that is not directly helpful for diagnosis, such as redundant adjectives, adverbs, or non-medical-related expressions. At the same time, the remaining word groups will also be formatted, such as unifying the case, standardizing the term expression, etc., to ensure the consistency and accuracy of subsequent processing. The purpose of doing this is to improve the quality of the data and make the finally generated recognized text clearer and more targeted.
[0078] S303 identifies and classifies the entity information in the recognized text, and the entity information includes the discomfort site, the description of discomfort feeling, drug information, and time information;
[0079] After obtaining the clean recognized text, the system will use the named entity recognition (NER) technology to extract the key medical entity information, including but not limited to the discomfort site (such as the head, chest), the description of discomfort feeling (such as pain, burning sensation), drug information (if mentioned), and time information (the time when the symptom starts). This step is crucial for understanding the specific details of the condition because it directly affects the accuracy of subsequent disease prediction.
[0080] S304 extracts the disease core keywords from the entity information;
[0081] Based on the various entity information identified in the previous step, the system will further refine the keywords closely related to potential diseases. These keywords are specific combinations of symptoms, frequently occurring site descriptions, or certain specific drug reaction patterns. They form a bridge connecting the patient's condition description and specific diseases.
[0082] S305 establishes a disease knowledge base containing various known diseases and their related symptoms;
[0083] In order to effectively compare and match the patient's symptom descriptions, the system relies on a comprehensive and detailed disease knowledge base. This knowledge base not only includes a wide range of disease entries but also details the common symptom manifestations of each disease. Building such a knowledge base involves integrating data from multiple sources, such as medical literature, clinical guidelines, and publicly available health database resources.
[0084] S306 calculates the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity;
[0085] The specific steps include:
[0086] S401 uses a pre-trained word embedding model to convert each word of the disease core keywords into a vector representation;
[0087] In this step, the system first needs to use a pre-trained word embedding model (such as Word2Vec, GloVe, or BERT, etc.) to convert the core keywords extracted from the patient's medical condition description into vector representations in a high-dimensional space. These word embedding models learn the semantic relationships between words through a large amount of text data, so they can effectively capture the meaning of words and represent them in vector form. For each disease core keyword, a corresponding vector will be generated, and this vector can reflect the position of the word in the semantic space.
[0088] S402 Apply the cosine similarity formula to calculate the symptom similarity.
[0089] Once all keywords have been converted into vector representations, the cosine similarity can then be used to measure the similarity between these vectors and the disease symptom vectors stored in the knowledge base. The cosine similarity determines whether the directions of two vectors are similar by calculating the cosine value of the angle between them. Its value ranges from -1 to 1, where 1 means exactly the same, 0 means no similarity, and -1 means exactly the opposite. In this scenario, since we are concerned with positive correlations, we are more interested in results close to 1. Specifically, for each pair of keyword vectors from the patient's medical condition description and the symptom vectors in the knowledge base, apply the cosine similarity formula to calculate the similarity score between them, thereby quantifying the matching degree between the two.
[0090] S307 Sort according to the similarity scores and retrieve relevant disease names from the knowledge base to obtain a set of disease names.
[0091] After calculating the similarity of all potential matches, the next step is to sort the results according to the similarity scores. This means that the system will rank all disease options from high to low according to the similarity scores. In this way, those diseases that best match the patient's medical condition description can be presented first. Finally, the system will extract several disease names at the top from the knowledge base to form a preliminary set of disease names for further analysis or for doctors' reference. This step not only helps to narrow down the diagnosis scope but also improves the diagnosis efficiency, enabling medical resources to be more accurately applied to the patient's treatment process.
[0092] S103 Retrieve all corresponding disease symptoms in the database based on the disease names and sort the disease symptoms according to the relevance scores of personal characteristics in the database. The personal characteristics include gender, age, and residential area.
[0093] The specific steps include:
[0094] S501 Collect the patient's personal information, which includes but is not limited to gender, age, and residential area.
[0095] Collect detailed personal information of patients, which includes but is not limited to gender, age, and residential area. This data is crucial for subsequent analysis because the susceptibility of people in different genders, age groups, and regions to certain diseases varies. For example, certain diseases are more common in specific age groups or genders, or due to different geographical environments, residents in certain areas are more prone to certain endemic diseases.
[0096] S502 Construct a characteristic disease database containing various diseases and their patient information;
[0097] Construct a comprehensive characteristic disease database that not only includes detailed descriptions and typical symptoms of various diseases but also associates patient information from a large number of actual cases. This information covers a wide range of individual characteristics such as gender, age distribution, and geographical features. In this way, the database can support multi-dimensional queries and analyses, providing strong data support for personalized medicine.
[0098] S503 Based on the set of matched disease names as keywords, retrieve all patient information corresponding to the disease in the disease database;
[0099] The specific steps include:
[0100] S601 Construct a corresponding SQL query statement according to the determined set of disease names to retrieve relevant patient information;
[0101] The system needs to carefully design an SQL query statement that can accurately extract the required information based on the set of disease names determined in the previous steps. This involves understanding the structure of the database, including table names, field names, and their relationships. For example, assume there is a table named diseases in the database to store detailed information about diseases, another table named patients to record all patients' personal information, and an associated table named patient_disease to connect patients with the diseases they have.
[0102] S602 Run the SQL query statement in the database management system to obtain all patient information related to the specified disease;
[0103] Once the SQL query statement is ready, the next step is to execute this query in the database management system (DBMS). This step is usually completed by the system's backend logic, which is responsible for establishing a secure connection to the database, sending the query request, and receiving the returned data results. During the execution of the query, the database engine scans the relevant tables, applies the necessary filtering conditions, and finally returns all patient records that meet the conditions.
[0104] Remove duplicate or incomplete records from the patient information obtained by S603.
[0105] The raw data retrieved from the database may contain some problems, such as duplicate records or entries missing key information. Therefore, these data need to be cleaned up next. For duplicates, they can be resolved by identifying and deleting duplicate records with the same unique identifier (such as patient_id). For those incomplete records lacking important information (such as empty fields for keywords like gender, age, etc.), a decision needs to be made whether to directly exclude them or attempt to fill in the missing values. In some cases, statistical methods or other algorithms can also be used to estimate the missing information, but this requires careful handling to avoid introducing biases.
[0106] S504 Calculate the similarity between the patient's personal information and all patient information to screen for similar patient groups;
[0107] At this stage, the system will compare the patient's personal information (such as gender, age, residential area, etc.) with the data of other patients to screen out the group of patients most similar to the current patient's situation. This step is crucial for improving the accuracy of diagnosis because it allows the system to infer symptoms or diseases based on the experiences of patients with similar backgrounds.
[0108] S505 Sort the disease symptoms corresponding to the similar patient groups in descending order based on the frequency of occurrence.
[0109] The specific steps include:
[0110] S701 Create a comprehensive symptom list that includes all the symptoms mentioned in the similar patient groups;
[0111] In this step, the system needs to iterate through each member of the similar patient group and collect all the symptoms recorded by them. This is done to build a comprehensive symptom list to ensure that no important symptoms are missed. This list includes not only common physical symptoms (such as fever, headache), but also covers some less common or disease-specific symptoms. Creating such an exhaustive list is the basis for subsequent statistical work.
[0112] S702 For each symptom, count the number of times it appears within the entire patient group;
[0113] Next, for the symptom list generated in step S701, the system needs to count the number of occurrences of each symptom within the entire group of similar patients. This means checking each patient's record one by one, marking and counting the frequency of each symptom. This process can be automated by writing a dedicated algorithm or script to ensure the accuracy and efficiency of data processing. In this way, the prevalence of different symptoms in this specific patient group can be quantified.
[0114] S703 sorts the symptoms in descending order according to the frequency of symptom occurrence.
[0115] The last step is to sort the statistically obtained symptom frequencies, usually from high to low. The purpose of this is to highlight those symptoms that most frequently occur among similar patients, helping doctors quickly identify the most relevant problems. This sorting not only guides the further diagnostic process but also provides a reference basis for the design of personalized treatment plans. For example, if a certain symptom frequently appears in the group of similar patients, it is a key symptom unique to this patient group and worthy of special attention.
[0116] S104 presents the sorted list of disease symptoms to the patient and obtains the second condition information selected by the patient that matches their own situation;
[0117] The specific steps include:
[0118] S801 generates a brief description based on the disease symptoms;
[0119] At this stage, the system will generate an easy-to-understand brief description for each symptom based on the disease symptom list screened and sorted in the previous steps. These descriptions not only describe the symptom itself (such as "persistent headache"), but also include some background information or common causes, as well as which diseases the symptom is usually associated with. The purpose of this is to enable patients to better understand and identify whether they have experienced similar symptoms. In addition, to improve the user experience, these descriptions should be as concise and clear as possible and avoid using overly professional medical terms.
[0120] S802 obtains the first mark of the patient for the brief description that matches their own situation;
[0121] Next, the system will provide a user-friendly interface that allows patients to browse all the generated brief descriptions and mark those symptoms that they think match their own conditions. This step can be achieved through simple checkboxes, buttons or other interactive elements, enabling patients to easily express their choices. The first marking process is mainly to initially screen out those symptoms that most reflect the patient's current health status, thereby narrowing the scope of subsequent analysis.
[0122] S803 Expand the detailed information corresponding to the brief description with the first mark;
[0123] Once the patient has completed the initial selection, the system will expand a more detailed explanation for those marked brief descriptions. This part includes a more in-depth description of symptoms, cause analysis, relationships with other symptoms, and potential influencing factors, etc. Such detailed information helps the patient make a more accurate judgment because it provides more context and details than the brief description. For some complex symptoms, relevant pictures or video materials can also be included to help the patient better understand their physical condition.
[0124] S804 Obtain the second mark of the patient for the detailed information to get the second condition information.
[0125] Finally, based on the provided detailed information, the system will request the patient to confirm again, that is, make a second mark on which of these detailed descriptions actually match their actual situation. This step is crucial because it directly relates to the quality of the final formed second condition information. Through this double-confirmation mechanism, not only can the possibility of misdiagnosis be reduced, but also it can ensure that the information collected reflects the patient's actual situation as accurately as possible. After completing this step, the system will integrate all the feedback to form a comprehensive second condition information report, which can be used as an important reference for doctors to formulate treatment plans.
[0126] S105 Generate online consultation information based on the basic condition description information and the second condition information.
[0127] It is necessary to integrate the basic condition description information collected from the patient (such as the initial symptom description, duration, etc.) with the second condition information obtained through the interaction process (that is, the relevant symptoms and details further confirmed by the patient). This step requires ensuring that all relevant information is accurately recorded and organized in an easy-to-understand manner. For example, the information can be classified and sorted according to symptom types, severity levels, or the time sequence of occurrence.
[0128] Next, the system will automatically generate a standardized online consultation information form according to a pre-set template. This form not only includes the patient's basic personal information (name, gender, age, contact information, etc.), but also details the patient's condition description and symptom details. To facilitate doctors to quickly grasp the key points, some key information, such as the most severe symptoms, whether there are emergency situations, etc., can also be highlighted in the form. In addition, considering the differences in the needs of different medical institutions, the system should support customizing and adjusting the content format of the consultation form to meet specific requirements.
[0129] In addition to information directly from patients, the system can also add some auxiliary analysis and suggestions to the consultation information based on existing medical knowledge bases and algorithms. For example, the system can suggest the disease direction based on the symptom patterns reported by the patient; or, if certain symptom combinations are found to indicate a higher risk in a specific population, it can recommend arranging further examinations as soon as possible. Although these analyses and suggestions are not the final diagnosis, they can provide valuable references for doctors and help them make judgments more quickly.
[0130] Second Embodiment
[0131] The present invention also provides an online consultation information processing system based on big data, including a disease condition information acquisition module, a disease name retrieval module, a disease symptom sorting module, a screening module, and a consultation information generation module; the disease condition information acquisition module is used to acquire the basic disease condition description information input by the patient; the disease name retrieval module is used to parse the basic disease condition description information by using natural language processing technology and retrieve all relevant disease names in the disease database; the disease symptom sorting module is used to sort all the corresponding disease symptoms in the disease name retrieval database based on the relevance scores of personal characteristics in the database, and the personal characteristics include gender, age, and residential area; the screening module is used to present the sorted disease symptom list to the patient and obtain the second disease condition information selected by the patient that conforms to the patient's own situation; the consultation information generation module is used to generate online consultation information based on the basic disease condition description information and the second disease condition information.
[0132] In this embodiment, the medical condition information acquisition module is used to receive and record the basic medical condition description information input by the patient. The basic medical condition description here includes key information such as the nature, duration, and severity of symptoms, providing basic data support for subsequent diagnosis. The disease name retrieval module uses natural language processing technology to parse the above basic medical condition description information and retrieve all relevant disease names in a vast disease database. This process not only requires high-precision text recognition ability but also needs the algorithm to be intelligent enough to understand the implicit meaning in the patient's expression, so as to cover all potential diseases as comprehensively as possible. The disease symptom sorting module extracts all corresponding disease symptoms from the retrieval database according to the disease name and sorts these symptoms according to the correlation scores regarding personal characteristics (such as gender, age, living area, etc.) in the database. Such a design takes into account the different symptom manifestations of different populations due to physiological differences and environmental factors, helping to more accurately locate the problem. The screening module will present the sorted list of disease symptoms to the patient and invite them to select the symptoms that match their actual situation to form the second medical condition information. This step is crucial for narrowing the diagnosis scope and focusing on the core problem because it directly depends on the patient's subjective judgment of their own health condition, increasing the accuracy of the diagnosis result. The consultation information generation module will automatically generate a detailed online consultation information by integrating the basic medical condition description information and the second medical condition information confirmed by the patient. This document not only contains all the original materials provided by the patient but also integrates the professional insights obtained from system analysis, providing a clear and systematic reference basis for doctors and greatly simplifying the subsequent diagnosis and treatment process.
[0133] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand the implementation of all or part of the above processes and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. An online consultation information processing method based on big data, characterized in that it includes: Obtain the basic disease description information input by the patient; Use natural language processing technology to parse the basic disease description information and retrieve all relevant disease names in the disease database; Retrieve all corresponding disease symptoms in the database based on the disease names, and sort the disease symptoms according to the correlation scores of personal characteristics in the database, where the personal characteristics include gender, age, and residential area; Present the sorted list of disease symptoms to the patient, and obtain the second disease information selected by the patient that conforms to their own situation; Generate online consultation information based on the basic disease description information and the second disease information.
2. The online consultation information processing method based on big data according to claim 1, characterized in that The specific steps of obtaining the basic disease description information input by the patient include: Obtain the basic disease description information; Detect whether the basic disease description information meets the requirements; Store the basic disease description information that meets the requirements in the database.
3. The online consultation information processing method based on big data according to claim 2, characterized in that The specific steps of using natural language processing technology to parse the basic disease description information and retrieve all relevant disease names in the disease database include: Segment the basic disease description information into word groups; Remove the invalid information in the word groups and format them to obtain the recognition text; Identify and classify the entity information in the recognition text, where the entity information includes discomfort location, discomfort feeling description, drug information, and time information; Extract the disease core keywords from the entity information; Establish a disease knowledge base containing various known diseases and their related symptoms; Calculate the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity; Sort according to the similarity scores, and retrieve relevant disease names from the knowledge base to obtain a set of disease names.
4. The online consultation information processing method based on big data according to claim 3, characterized in that The specific steps of calculating the similarity between the disease core keywords and the symptoms recorded in the knowledge base using cosine similarity include: Use a pre-trained word embedding model to convert each word of the disease core keywords into a vector representation; Apply the cosine similarity formula to calculate the symptom similarity.
5. The online consultation information processing method based on big data according to claim 4, characterized in that The specific steps of retrieving all corresponding disease symptoms in the database based on the disease names and sorting the disease symptoms according to the correlation scores of personal characteristics in the database include: Collect the patient's personal information, where the patient's personal information includes but is not limited to gender, age, and residential area; Construct a characteristic disease database containing various diseases and their patient information; Based on the set of disease names obtained as keywords, retrieve all patient information corresponding to the disease in the disease database; Calculate the similarity between the patient's personal information and all patient information to screen for similar patient groups; Sort the disease symptoms corresponding to the similar patient groups in descending order based on the occurrence frequency.
6. The method for processing online consultation information based on big data according to claim 5, wherein the specific steps of retrieving all patient information corresponding to the disease in the disease database by using the matched disease name set as a keyword include: constructing a corresponding SQL query statement according to the determined disease name set to retrieve relevant patient information; running the SQL query statement in the database management system to obtain all patient information related to the specified disease; removing duplicate or incomplete records from the obtained patient information.
7. The method for processing online consultation information based on big data according to claim 6, wherein the specific steps of sorting the disease symptoms corresponding to the similar patient group in descending order based on the occurrence frequency include: creating a comprehensive symptom list including all symptoms mentioned in the similar patient group; counting the number of times each symptom appears within the entire patient group; sorting the symptoms in descending order according to the occurrence frequency of the symptoms.
8. The method for processing online consultation information based on big data according to claim 7, wherein the specific steps of presenting the sorted disease symptom list to the patient for selecting the condition information that conforms to their own situation to obtain the second condition information include: generating a brief description based on the disease symptoms; obtaining the first mark of the patient for the brief description that conforms to their own situation; expanding the detailed information corresponding to the brief description with the first mark; obtaining the second mark of the patient for the detailed information to obtain the second condition information.
9. An online consultation information processing system based on big data, which is applied to the method for processing online consultation information based on big data according to any one of claims 1 to 8, wherein it includes a condition information acquisition module, a disease name retrieval module, a disease symptom sorting module, a screening module and an online consultation information generation module; the condition information acquisition module is used to acquire the basic condition description information input by the patient; the disease name retrieval module is used to parse the basic condition description information by using natural language processing technology and retrieve all relevant disease names in the disease database; the disease symptom sorting module is used to retrieve all corresponding disease symptoms in the database based on the disease name and sort the disease symptoms according to the relevance score of personal characteristics in the database, and the personal characteristics include gender, age, and residential area; the screening module is used to present the sorted disease symptom list to the patient and obtain the second condition information selected by the patient that conforms to their own situation; the online consultation information generation module is used to generate online consultation information based on the basic condition description information and the second condition information.