Methods and devices for analyzing consultation data
By identifying user demand characteristics and performing data table matching and numerical analysis, the problem of insufficient data analysis in internet healthcare platforms has been solved, enabling rapid and accurate analysis of consultation data and improving operational efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2026-06-30
AI Technical Summary
Existing internet healthcare platforms lack sufficient data analysis capabilities, making it impossible to conduct real-time, accurate, multi-dimensional, and multi-level data analysis. This leads to difficulties in intelligent analysis and operational decision-making for online consultations, impacting user experience.
By identifying user needs, extracting their characteristics and matching them with data tables, generating a set of query statements, executing the queries and performing numerical analysis, and using a diagnostic knowledge base and large models to generate natural language analysis results, the causes of problems can be quickly located and analyzed.
It enables rapid and accurate analysis of consultation data, improves the operational efficiency and user experience of online medical platforms, and supports intelligent analysis and decision-making in online consultations.
Smart Images

Figure CN122314286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network medical data analysis technology, and in particular to methods and devices for analyzing consultation data. Background Technology
[0002] With the rapid development of internet technology, internet healthcare platforms have greatly improved the accessibility and efficiency of medical services through online consultations and telemedicine. However, the digital operation of existing internet healthcare platforms still mainly relies on manual statistics and simple reports, requiring manual analysis of the impact of each dimension attribute on the results. Even with experienced analysts, problems remain with inaccurate, incomplete, inefficient, and limited-dimensional analysis. This hinders real-time, accurate, multi-dimensional, and multi-level data analysis, failing to meet the ever-changing operational needs of internet healthcare platforms. For example, in the intelligent analysis and operational decision-making aspects of online consultations, insufficient data analysis capabilities prevent the rapid identification and resolution of problems, impacting user experience. Summary of the Invention
[0003] This invention provides a method and apparatus for analyzing online medical consultation data, aiming to improve the analytical capabilities of online medical consultation data and quickly and accurately locate and analyze problems in the digital operation process of online medical consultation.
[0004] To achieve the above objectives, according to a first aspect of the present invention, a method for analyzing medical history data is provided, comprising:
[0005] In response to identifying user analysis requests from user statements, request features are extracted from the user analysis requests, and the request features are matched with a data table. The request features include user intent, dimensions, and metrics.
[0006] In response to matching the target data table, a set of query statements corresponding to the requested features is generated;
[0007] Execute the set of query statements to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator.
[0008] Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and to generate natural language analysis results of the user analysis demand corresponding to the demand feature.
[0009] Furthermore, the method is based on a predefined consultation knowledge base, which includes definitions of consultation terminology and table structure definitions of data tables. The definitions of consultation terminology include user intent definitions, dimension definitions, and indicator definitions.
[0010] Extracting request features from the user analysis requests and matching these features with the data table includes:
[0011] The user analysis request is segmented into words, and the segmentation results are compared with the professional medical terminology definitions in the medical knowledge base to calculate the first similarity. In response, user intents, dimensions, and indicators with a first similarity greater than the first similarity threshold are selected and combined to obtain the request features.
[0012] The second similarity is calculated between each of the stated claim features and the table structure definition of the data table, and the data tables with a second similarity greater than the second similarity threshold are selected as the target data tables.
[0013] Furthermore, the consultation knowledge base includes table structure definitions and mappings to query statement set templates;
[0014] In response to a match with the target data table, a set of query statements corresponding to the requested features is generated, including:
[0015] Matching is performed based on the table structure definition of the target data table and the mapping between the table structure definition and the query statement set template;
[0016] In response to a match to the mapping, a set of query statements for the target data table is generated based on the query statement set template corresponding to the mapping;
[0017] In response to the failure to match the mapping, a large model prompt is generated based on the user's analysis request and the request characteristics. The large model prompt is then input into the question-and-answer large model to generate a set of query statements for the target data table.
[0018] Furthermore, based on the user analysis requests and their characteristics, a large model prompt is generated. This large model prompt is then input into the question-and-answer large model to generate a query statement for the target data table, including:
[0019] The user analysis requests and their characteristics, along with the table structure definition of the target data table, are used as the first major model prompts for input into the question-and-answer model;
[0020] The question-answering big model generates and outputs a first answer based on the prompts of the first big model and the question-and-answer knowledge base. The first answer includes a first set of query statements corresponding to the prompts of the first big model and its query result instances, dimensions and indicators in the first set of query statements, and natural language interpretation of the first set of query statements.
[0021] In response to the user's confirmation of the first answer, the first set of query statements is used as the query statement for the target data table;
[0022] In response to the user's correction instruction for the first answer, the large model generates a second large model prompt based on the first large model prompt, the correction prompt, and the consultation knowledge base. Based on the second large model prompt and the consultation knowledge base, a second answer is generated and output. The second answer includes the second large model prompt, a second query statement set corresponding to the second large model prompt and its query result instance, the dimensions and indicators in the second query statement set, and the natural language interpretation of the second query statement set.
[0023] In response to the user's confirmation of the second answer, the second set of query statements is used as the query statement for the target data table.
[0024] Furthermore, the indicator in the user's request is the timeliness of consultation, and the dimensions in the user's request include consultation scenario, consultation channel, consultation category, consultation department and consultation time period. The user's intent includes the target dimension and target dimension value that affect the timeliness of consultation.
[0025] Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that influence each indicator in the claimed feature from the dimensions and dimension values in the dataset, including:
[0026] Obtain the consultation timeliness value and the dimension value of the dimension from the dataset;
[0027] Construct a combination of multiple said dimension values that corresponds to the consultation timeliness value;
[0028] Based on the consultation timeliness value corresponding to each combination of dimension values, calculate the influence factor of each combination of dimension values on each consultation timeliness value;
[0029] Based on the magnitude of the influencing factors, the combination of target dimension values affecting the timeliness of patient reception is determined, and the target dimension and the target dimension value are obtained.
[0030] Furthermore, the metrics in the user demands are aggregated metrics calculated by aggregating order volume and number of doctors. The dimensions in the user demands include consultation scenarios, consultation channels, consultation categories, consultation departments, and consultation time periods. The metrics include consultation timeliness. The user intent includes target dimensions and target dimension values that affect the aggregated metrics.
[0031] Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that influence each indicator in the claimed feature from the dimensions and dimension values in the dataset, including:
[0032] The order volume and doctor volume values are obtained from the dataset, and the aggregated index value of the aggregated index is obtained through the aggregation calculation.
[0033] Obtain the dimension values of the dimension from the dataset, and construct a dimension value combination consisting of multiple dimension values and corresponding to the aggregated index value;
[0034] Based on the aggregated index value corresponding to the combination of values of each dimension, calculate the influence factor of the combination of values of each dimension on the aggregated index;
[0035] Based on the magnitude of the influencing factors, determine the combination of target dimension values that affect the aggregated index, and obtain the target dimension and the target dimension value.
[0036] Furthermore, generating natural language analysis results for the user analysis request corresponding to the requested features includes:
[0037] Based on the indicators, the influencing factors, and their corresponding target dimensions and target dimension values, a corresponding third major model suggestion is generated;
[0038] The third major model prompt is input into the pre-trained large model for analyzing consultation data to obtain the natural language analysis results of the statistical claims of the indicators.
[0039] According to a second aspect of the present invention, a medical consultation data analysis device is provided, comprising:
[0040] The semantic matching module is used to respond to the user analysis request identified from the user statement, extract the request features from the user analysis request, and match the request features with the data table. The request features include user intent, dimensions, and metrics.
[0041] The statement generation module is used to generate a set of query statements corresponding to the requested features in response to matching the target data table.
[0042] The data acquisition module is used to execute the query statement set to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator.
[0043] The data analysis module is used to perform numerical analysis on the dataset, filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and generate natural language analysis results of the user analysis demand corresponding to the demand feature.
[0044] According to a third aspect of the present invention, an electronic processing apparatus is provided, comprising:
[0045] One or more processors; and,
[0046] A memory communicatively connected to the at least one processor; wherein,
[0047] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, cause the one or more processors to implement the method provided in the first aspect of the present invention.
[0048] According to a fourth aspect of the present invention, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method provided in the first aspect of the present invention.
[0049] According to a fifth aspect of the present invention, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided in the first aspect of the present invention.
[0050] One embodiment of the invention has the following advantages or beneficial effects:
[0051] By employing the demand features extracted from user analysis requests in user statements, matching data tables, and generating and executing query sets corresponding to the demand features, relevant data scattered across multiple data tables can be quickly and accurately matched. This allows for the identification of the data source for online consultation data analysis and the acquisition of the corresponding dataset. Numerical analysis is then performed on this dataset, filtering out the target dimensions and target dimension values that influence each indicator within the demand features from the dimensions and dimension values in the dataset. Corresponding natural language analysis results are generated, providing both the data analysis results and corresponding natural language explanations for the user analysis requests. This enables rapid and accurate analysis and location of the causes of digital operational problems in user analysis requests, providing interpretations and suggestions based on the data analysis. This enhances the analytical capabilities of online medical consultation data, facilitating intelligent analysis and operational decision-making for online consultations on the platform, improving operational methods, and enhancing user experience. Attached Figure Description
[0052] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:
[0053] Figure 1 This is a schematic diagram of the main process of the consultation data analysis method according to an embodiment of the present invention;
[0054] Figure 2 This is a schematic diagram of the data extraction and matching process in the consultation data analysis method according to an embodiment of the present invention;
[0055] Figure 3 This is a flowchart illustrating semantic matching in a consultation data analysis method according to another embodiment of the present invention;
[0056] Figure 4 This is a flowchart illustrating the process of generating query statements from a large model in a consultation data analysis method according to another embodiment of the present invention;
[0057] Figure 5 This is a flowchart illustrating the numerical analysis process in a consultation data analysis method according to an embodiment of the present invention.
[0058] Figure 6 This is a flowchart illustrating the numerical analysis process in a consultation data analysis method according to another embodiment of the present invention;
[0059] Figure 7 This is a schematic diagram of the constituent modules of the consultation data analysis device according to an embodiment of the present invention;
[0060] Figure 8 This is a schematic diagram of the constituent modules of a medical history data analysis device according to another embodiment of the present invention;
[0061] Figure 9 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;
[0062] Figure 10 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation
[0063] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0064] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0065] Figure 1 This is a schematic diagram of the main flow of the consultation data analysis method according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the data extraction and matching process in the consultation data analysis method according to an embodiment of the present invention. Figure 1 and Figure 2 As shown, the consultation data analysis method in this embodiment of the invention includes the following steps S101 to S104.
[0066] Step S101: In response to identifying user analysis requests from user statements, extract request features from the user analysis requests and match the request features with a data table. The request features include user intent, dimensions, and metrics.
[0067] Understandably, user intent refers to the goal or need expressed in a user's analytical request. Dimensions are used to describe the characteristics or attributes of data. Metrics are used to measure and evaluate the numerical value, proportion, or attribute of a specific analytical dimension, typically including the statistical period, dimension, and measurement method. This step includes two aspects: firstly, identifying the user's analytical request in the user's statement; and secondly, if the user's analytical request is identified, further identifying the characteristics of that request. In this embodiment and some other embodiments of the present invention, identifying the user's analytical request from the user's statement and extracting the request features from the user's analytical request can be achieved using intelligent technologies such as question-answering large models or intelligent agents.
[0068] Specifically, such as Figure 2 As shown, the question-answering model or intelligent agent identifies key content from the user's statement "Please check the consultation time of dermatologists in the last 30 minutes" including "check the consultation time of dermatologists in the last 30 minutes". Based on the key content, it determines that the user's analysis request includes "check consultation time" and then extracts the user's intent as "check consultation time", with the dimensions being department (dimensional value: dermatology) and time (dimensional value: within 30 minutes) and the indicator being "consultation time".
[0069] Understandably, after extracting the characteristics of the user's request, feature matching can be used to find the data table. In this embodiment and some other embodiments of the present invention, the user's request characteristics are analyzed and similarity is calculated with predefined data table characteristics. The matching data table is then selected based on the similarity. Table 1 shows examples of table characteristics of the patient reception details table and doctor information table matched in this embodiment. As shown in Table 1, the predefined table characteristics in this embodiment include the data table identifier and natural language meaning, as well as the field identifier and natural language meaning of the data table.
[0070]
[0071] The data table is searched using feature matching. Specifically, the dimension "Department" is matched against the field "dep" and its meaning "Order Department" in the table structure definition. If an unknown dimension is encountered, the match will fail. Some embodiments of this invention will indicate that the analysis dimension does not exist, allowing users to adjust their analysis requirements or add data tables to meet their analytical needs.
[0072] Step S102: In response to matching the target data table, generate a set of query statements corresponding to the requested features.
[0073] Specifically, in this embodiment and some other embodiments of the present invention, a query statement set template is obtained based on a pre-defined mapping relationship between a data table and a query statement set. The query statement set template is a standardized set of query statements pre-defined according to the table fields of the data table. For example... Figure 2 As shown, based on the user request analysis of "querying the consultation time of dermatology clinics in the last 30 minutes," the request features were extracted. Feature matching was used to match the relevant consultation details table. The consultation details table contains dimension-related fields "order department" and "order time," and indicator-related field "consultation time." Furthermore, the consultation details table and doctor information table have pre-defined SQL query statement templates. By substituting the dimension values and / or indicator values corresponding to the user request into the template, the SQL query statement set corresponding to the request features can be generated. Specifically, as shown... Figure 2 As shown, Figure 2 The template only shows the SQL query statement set template for the patient reception details table. The dimension values "dermatology" and "30 minutes" are inserted into the SQL query statement set template for this template.
[0074] Step S103: Execute the query statement set to query the target data table, and filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator.
[0075] Understandable, execution as Figure 2 The set of SQL queries shown can obtain the dataset corresponding to the requested features (dimensions are department (dimension value: dermatology) and time (dimension value: 30 minutes), and the indicator is "consultation timeliness") from the patient visit details table and the doctor information table. Specifically, an example of some data is shown in Table 2 (current time is 8:00 AM).
[0076]
[0077] Referring to Table 2, it can be understood that the dataset retrieved from the patient details table and doctor information table by executing the SQL query statement set corresponding to the requested feature includes, in addition to the dimensions and corresponding dimension values, indicators and corresponding indicator values directly corresponding to the requested feature (department, order time, and patient reception time shown in Table 1), other dimensions and dimension values of the data table records determined by the requested feature (patient doctor identifier, hospital, and channel shown in Table 2). That is, the dimensions and indicators in the dataset are no less than the dimensions and indicators of the requested feature.
[0078] Step S104: Perform numerical analysis using the dataset, filter out the target dimensions and target dimension values that affect each indicator in the request feature from the dimensions and dimension values in the dataset, and generate the natural language analysis result of the user analysis request corresponding to the request feature.
[0079] Understandably, once the dataset is obtained, numerical analysis can be performed. This numerical analysis allows for the selection of target dimensions and target dimension values from the dimensions and values within the dataset that influence each indicator in the stated characteristics. Specifically, numerical analysis methods include regression analysis, correlation analysis, and cluster analysis. These methods can be used to determine whether dimensions such as the attending physician and the channel have an impact on the timeliness of consultation.
[0080] Specifically, in this embodiment and some embodiments of the present invention, after obtaining the target dimension and target dimension value of each indicator influencing the demand feature, the natural language analysis result of the user analysis demand corresponding to the demand feature is generated using a question-and-answer big model. Specifically, this includes: generating corresponding big model prompts based on the indicator and its corresponding target dimension and target dimension value; inputting the big model prompts into a pre-trained consultation data analysis big model to obtain the natural language analysis result of the indicator statistical demand. For example, after determining that the target dimension affecting consultation timeliness is scenario and the target dimension value is webpage, according to the big model prompt rules (prompting the objects or ranges of big model input and output), the big model prompt could be "Please provide methods to improve consultation timeliness in webpage scenarios based on the data records in the data table shown in 'Table 1' where the scenario is webpage." Inputting the above big model prompt into the question-and-answer big model will yield the corresponding natural language analysis result.
[0081] Optionally, in other embodiments of the present invention, step S101, identifying user analysis requests from user statements and extracting request features from said user analysis requests, can be implemented based on a knowledge base and artificial intelligence methods. In these embodiments of the present invention, firstly, by defining user analysis requests, user intentions, dimensions, and indicators as predefined in the knowledge base, the expression forms and meanings of user analysis requests, user intentions, dimensions, and indicators are unified and standardized, facilitating computer understanding. Secondly, based on predefined professional knowledge in the knowledge base and a pre-trained semantic recognition neural network or question-answering model, one or more user requests in the user statements are identified, and then it is determined whether each user request is a user analysis request related to consultation data analysis. In response to the identification that a user request is a user analysis request related to consultation data analysis, request features are extracted from each user analysis request. Each user analysis request is segmented into words, and the resulting word groups are compared with the professional knowledge in the knowledge base, for example, through similarity calculation, thereby extracting the request features of each user analysis request. Specifically, for example, one user analysis request is "to check the consultation timeliness of dermatology clinics within the last 30 minutes." Based on a pre-built professional knowledge base for analyzing consultation data, the user intent is "to check consultation timeliness," with dimensions being department (dimension value: dermatology) and time (dimension value: 30 minutes), and the metric being "consultation timeliness." Based on dimensions and metrics, user intent can also be defined in a more specific way, such as using statistical analysis of one or more dimensions, to ensure consistency and computer understanding. Therefore, based on the user analysis request "to check the consultation timeliness of dermatology clinics within the last 30 minutes," we can also derive the user intent as "to check the consultation timeliness of [department] within [time]."
[0082] Therefore, in another embodiment of the present invention, a predefined consultation knowledge base is provided. This knowledge base includes definitions of professional consultation terms and table structure definitions for data tables. The definitions of professional consultation terms include user intent definitions, dimension definitions, and indicator definitions. The user intent definitions include their natural language meanings, while the dimension and indicator definitions each include their natural language meanings and values, i.e., dimension values or indicator values. Dimension and indicator values can be discrete or continuous values. For example, the dimension "department" can be "surgery," "internal medicine," etc., and the indicator "conduct timeliness" can be a continuous time value. The table structure definitions for the data tables include the source table structure definition and / or the view table structure definition. The view table is a temporary table formed based on the source table.
[0083] In this embodiment, specifically, such as Figure 3 As shown, extracting request features from the user analysis requests and matching the request features with the data table includes:
[0084] Step S301: The user's analysis request is segmented into words, and the segmentation results are compared with the definitions of professional medical terms in the medical knowledge base to calculate a first similarity. In response to filtering out user intentions, dimensions, and indicators with a first similarity greater than a first similarity threshold, these are combined to obtain the request features. Note that the term "first" in similarity calculation and similarity threshold is used only to indicate purpose and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. This invention defines features defined by terms such as "first" and "second" as explicitly or implicitly including at least one of those features. Therefore, the first similarity threshold and the first similarity calculation represent the target of the similarity calculation, not a limitation on the similarity calculation.
[0085] Step S302: Perform a second similarity calculation on each of the stated claim features and the table structure definition of the data table, and select the data table with a second similarity greater than the second similarity threshold as the target data table.
[0086] Specifically, in this embodiment, a diagnostic knowledge base is used to accurately identify user analysis requests in user statements and extract request features. If no user analysis request is identified, the correct input can be prompted to the user through a question-and-answer model. If the user analysis request lacks dimensions, default dimensions and default dimension values can be defined in the diagnostic knowledge base to handle situations where the user analysis request lacks dimensions.
[0087] Furthermore, in this embodiment and other embodiments of the present invention, the consultation knowledge base also includes table structure definitions and mappings to query statement set templates. Therefore, in response to matching a target data table, generating a query statement set corresponding to the request feature includes: matching the table structure definition of the target data table with the mapping between the table structure definition and the query statement set template; and in response to matching the mapping, generating a query statement set for the target data table based on the query statement set template corresponding to the mapping. By setting table structure definitions and mappings to query statement set templates in the knowledge base, query statement sets can be quickly obtained for repeated user analysis requests without repetitive processing, resulting in high efficiency.
[0088] Furthermore, in this embodiment and some other embodiments of the present invention, in response to the failure to match the mapping, a large model prompt is generated based on the user analysis request and its request characteristics. The large model prompt is then input into the question-and-answer large model to generate a set of query statements for the target data table, such as... Figure 4 As shown, it specifically includes:
[0089] Step S401: The user analysis request, its request characteristics, and the table structure definition of the target data table are used as the first major model prompt input into the question-and-answer major model.
[0090] In step S402, the question-answering big model generates and outputs a first answer based on the prompts from the first big model and the diagnostic knowledge base. The first answer includes a first query statement set corresponding to the prompts from the first big model and instances of its query results, dimensions and indicators in the first query statement set, and natural language interpretations of the first query statement set. Specifically, in this embodiment, the question-answering big model queries the diagnostic knowledge base to generate and output the first answer. The user then judges whether the first answer is correct, complete, and meets their needs based on the output. If it does, it is confirmed; otherwise, it is corrected based on the first answer.
[0091] Step S403: In response to the user's confirmation of the first answer, the first set of query statements is used as the query statement of the target data table.
[0092] Step S404: In response to the user's correction instruction for the first answer, the large model generates a second large model prompt based on the first large model prompt, the correction prompt, and the consultation knowledge base. Then, it generates and outputs a second answer based on the second large model prompt and the consultation knowledge base. The second answer includes the second large model prompt, a second set of query statements corresponding to the second large model prompt and its query result instances, dimensions and metrics in the second set of query statements, and natural language interpretations of the second set of query statements. It is understood that in this embodiment, the question-answering large model generates and outputs the second answer by querying the consultation knowledge base.
[0093] Step S405: In response to the user's confirmation of the second answer, the second set of query statements is used as the query statement for the target data table. It is understood that the user judges whether the second answer is correct, complete, and meets their requirements based on the output. If it does, the user confirms; otherwise, the user corrects the answer based on the first answer.
[0094] Specifically, for example, the large model prompt in step S401 is "Query the average consultation time of doctors in different hospitals in the last 30 minutes. Please automatically link the consultation details table and the doctor information table, and group them according to the doctor and hospital." Based on this large model prompt, the large model will automatically create SQL query statements as shown in Table 3.
[0095]
[0096] Executing the SQL query shown in Table 3 will automatically link the patient visit details table and the doctor information table, grouping them by doctor and hospital, and calculating the average patient visit time for each hospital's doctors. In this way, the large-scale model can automatically create SQL query statements based on the user-defined large-scale model, achieving structured queries across tables, thereby improving the efficiency and accuracy of data analysis.
[0097] Understandably, users judge whether the answer is correct, complete, and meets their needs based on the output. If it does, they confirm it; otherwise, they respond to the user's correction instruction for the current answer. The big model generates the next big model prompt based on the current big model prompt, the correction prompt, and the consultation knowledge base. It then generates the next answer based on the next big model prompt and the consultation knowledge base and outputs it for the user to confirm or correct.
[0098] Furthermore, in one embodiment of the present invention, for a specific scenario of online medical consultation, the indicator in the user's demand is the timeliness of consultation. The dimensions of the user's demand include consultation scenario, consultation channel, consultation category, consultation department, and consultation time period. Consultation scenario, for example, is consultation or order consultation; consultation channel, for example, is a mini-program, APP, or webpage; consultation category, for example, is the respiratory system or digestive system; consultation department, for example, is dermatology or surgery; and consultation time period, for example, is a time period in hours. In addition, user intent includes the target dimension and target dimension value that affect the timeliness of consultation.
[0099] Therefore, optionally, step S104 uses the dataset for numerical analysis, filtering out the target dimensions and target dimension values that affect each indicator in the claim feature from the dimensions and dimension values in the dataset, such as... Figure 5 As shown, it includes:
[0100] Step S501: Obtain the consultation timeliness value and the dimension value of the dimension from the dataset; specifically, the obtained dataset is shown in Table 4.
[0101]
[0102] Step S502: Construct a combination of multiple dimension values and corresponding to the consultation timeliness value; specifically, for example, combine A1 in the consultation scenario and B1 in the consultation category, or combine A1 in the consultation scenario, B1 in the consultation category and C2 in the consultation department.
[0103] Step S503: Calculate the influence factor of each dimension value combination on each consultation timeliness value based on the consultation timeliness value corresponding to each dimension value combination; specifically, this includes: using the consultation timeliness value corresponding to the dimension value combination, calculating the ratio of the standard deviation to the root mean square value of the consultation timeliness value; and calculating the influence factor of the dimension value combination on the consultation timeliness value based on the positive correlation with the square of the ratio.
[0104] Step S504: Based on the magnitude of the influencing factors, determine the target dimension value combinations that affect the timeliness of patient reception, and obtain the target dimensions and their values. Specifically, sort the target dimension value combinations according to the magnitude of the influencing factors, and identify the several dimension value combinations with the greatest impact.
[0105] Furthermore, in another embodiment of the present invention, for a specific scenario of online medical consultation, the indicators in the user's demands are aggregated indicators calculated by aggregating order volume and doctor volume. The dimensions in the user's demands include consultation scenario, consultation channel, consultation category, consultation department, and consultation time period. The indicators include consultation timeliness, and the user's intent includes target dimensions and target dimension values that affect the aggregated indicators; for example... Figure 6 As shown, numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that influence each indicator in the claimed feature from the dimensions and dimension values in the dataset, including:
[0106] Step S601: Obtain the value of the order volume and the value of the number of doctors from the dataset, and obtain the aggregated index value of the aggregated index through the aggregation calculation; for example, the obtained data is shown in Table 5, and the aggregated index is the ratio of the order volume to the number of doctors.
[0107]
[0108] Step S602: Obtain the dimension values of the dimension from the dataset, and construct a dimension value combination consisting of multiple dimension values and corresponding to the aggregated index value. Specifically, combine multiple dimension values under different dimensions.
[0109] Step S603: Calculate the influence factor of each dimension value combination on the aggregated index based on the aggregated index value corresponding to each dimension value combination; specifically, this includes using the aggregated index value corresponding to the dimension value combination to calculate the ratio of the standard deviation to the root mean square value of the aggregated index; and calculating the influence factor of the dimension value combination on the aggregated index based on the positive correlation with the square of the ratio.
[0110] Step S604: Based on the magnitude of the influencing factors, determine the target dimension value combinations that affect the aggregated index, and obtain the target dimension and the target dimension value. Specifically, sort the target dimension value combinations according to the magnitude of the influencing factors, and identify the several dimension value combinations with the greatest influence.
[0111] Specifically, in some embodiments of the present invention, the method for calculating the impact factor is as follows:
[0112]
[0113] in, This represents the combination of the j-th dimension values on dimension combination d. Influencing factors on index or aggregate index I, This represents the value of indicator I in the k-th data record; This represents the average value obtained based on all index values of index I or aggregate index I; Indicates the number of data records.
[0114] In these embodiments of the present invention, specifically, generating the natural language analysis results of the user analysis requests corresponding to the request features includes: generating a corresponding third major model prompt based on the indicators, the influencing factors and their corresponding target dimensions and target dimension values; inputting the third major model prompt into a pre-trained large-scale diagnostic data analysis model to obtain the natural language analysis results of the indicator statistical requests.
[0115] Understandably, the consultation data method of this invention can quickly analyze and locate problem dimensions and values in consultation data analysis scenarios with a large number of dimensions and combinations of dimension values, determine the impact on indicators and their values, and quickly provide data interpretation, thereby improving analysis efficiency.
[0116] According to another aspect of the embodiments of the present invention, such as Figure 7 As shown, a medical history data analysis device 700 is provided, comprising:
[0117] The semantic matching module is used to respond to the user analysis request identified from the user statement, extract the request features from the user analysis request, and match the request features with the data table. The request features include user intent, dimensions, and metrics.
[0118] The statement generation module is used to generate a set of query statements corresponding to the requested features in response to matching the target data table.
[0119] The data acquisition module is used to execute the query statement set to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator.
[0120] The data analysis module is used to perform numerical analysis on the dataset, filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and generate natural language analysis results of the user analysis demand corresponding to the demand feature.
[0121] Specifically, Figure 8 A medical history data analysis device 800 applying the medical history data analysis method of the present invention is provided in one embodiment of the present invention. For example... Figure 8As shown, the consultation data analysis device 800 includes a semantic matching module 81, a statement generation module 82, a data acquisition module 83, and a data analysis module 84; it also includes a data source 85, a large model 86, and a consultation knowledge base 87. The consultation knowledge base 87 includes definitions of consultation professional terms, table structure definitions of data tables, and the mapping between the table structure definitions and query statement set templates. The definitions of consultation professional terms include a medical professional terminology database, user intent definitions, dimension definitions, and indicator definitions.
[0122] The semantic matching module 81 is used to extract request features from the user analysis request in response to the user statement; the statement generation module 82 is used to generate a set of query statements corresponding to the request features in response to the matching of the target data table; the data acquisition module 83 is used to execute the query statement set to query the target data table to obtain the dataset corresponding to the request features; and the data analysis module 84 is used to perform numerical analysis using the dataset, filter out the target dimensions and target dimension values that affect each indicator in the request features from the dimensions and dimension values in the dataset, and generate the natural language analysis result of the user analysis request corresponding to the request features.
[0123] Specifically, in this embodiment, the consultation data analysis method of the present invention, implemented based on the consultation data analysis device 800, includes:
[0124] In step S801, the semantic matching module 81 identifies user statements using the large model 86 and based on the consultation knowledge base 87 (in some embodiments, a medical terminology database is mainly used). In response to identifying user analysis requests from user statements, the module extracts request features from the user analysis requests using the large model 86 and based on the user intent definition, dimension definition, and indicator definition in the consultation knowledge base 87. Based on the request features, the module performs matching using the large model 86 and based on the data table structure definition in the consultation knowledge base 87. If the data matching result of the large model 86 matches the target data table, the module returns the table structure of the target data table.
[0125] Step S802: The large model 86 matches the table structure definition of the target data table with the mapping between the table structure definition and the query statement set template; in response to the matching of the mapping, the statement generation module 82 generates the query statement set of the target data table based on the query statement set template corresponding to the mapping.
[0126] In step S803, in response to the lack of a matching mapping, the statement generation module 82 generates a large model prompt based on the user's analysis request and its characteristics, and the table structure of the target data table returned in step S801. The large model prompt is then input into the question-and-answer large model to generate a set of query statements for the target data table. Specifically, this involves... Figure 4 The steps are shown.
[0127] In step S804, the data acquisition module 83 executes the set of query statements generated in step S802 or step S804 to query the target data table from the database 85 and obtain the dataset corresponding to the requested feature.
[0128] In step S805, the data analysis module 84 performs numerical analysis using the dataset, and filters out the target dimension and target dimension value that affect each indicator in the demand feature from the dimensions and dimension values in the dataset.
[0129] In step S806, the data analysis module 84 generates a prompt input model 86 based on the numerical analysis results, namely the target dimension and target dimension value of each indicator in the demand feature, and generates the natural language analysis results of the user analysis demand corresponding to the demand feature.
[0130] In this embodiment of the invention, by combining a knowledge base, a large model, and a data source, self-service attribution analysis is achieved. Through multidimensional data analysis, the factors affecting the analysis indicators can be fully revealed, the degree of influence of each dimension on the analysis indicators can be quantified, and the causes of data anomalies can be located.
[0131] Understandably, the consultation data analysis device in this embodiment of the invention can serve as an intelligent assistant for the operation of consultation data on internet healthcare platforms. Through intelligent data analysis and operational decision-making, it improves the platform's consultation efficiency and service quality. In this embodiment, automated data analysis is achieved through semantic analysis, association of dimensional indicators across multiple tables, indicator aggregation, and dimensional analysis. It provides real-time, multi-dimensional internet healthcare consultation data analysis functions, supporting multi-dimensional data analysis and factor localization. It also provides intelligent doctor operation suggestions, meeting the operational data requirements of doctors in different departments, scenarios, and categories, thereby reducing the workload of operational staff.
[0132] Figure 9 An exemplary system architecture 900 is shown that can be applied to the consultation data analysis method or consultation data analysis device of the present invention.
[0133] like Figure 9 As shown, system architecture 900 may include terminal devices 901, 902, and 903, network 904, and server 905. Network 904 is used as a medium to provide a communication link between terminal devices 901, 902, and 903 and server 905. Network 904 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0134] Users can use terminal devices 901, 902, and 903 to interact with server 905 via network 904 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 901, 902, and 903, such as medical consultation applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0135] Terminal devices 901, 902, and 903 can be various electronic devices with displays that support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0136] Server 905 can be a server providing various services, such as a backend management server supporting internet medical websites browsed by users using terminal devices 901, 902, and 903 (this is just an example). The backend management server can analyze and process data such as received consultation orders and feed the processing results back to the terminal devices.
[0137] It should be noted that the consultation data analysis method provided in this embodiment of the invention is generally executed by server 905, and correspondingly, the consultation data analysis device is generally set in server 905.
[0138] It should be understood that Figure 9 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0139] The following is for reference. Figure 10 It shows a schematic diagram of the structure of a computer system 1000 suitable for implementing a terminal device of the present invention. Figure 10 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0140] like Figure 10 As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1002 or programs loaded from storage section 1008 into random access memory (RAM) 1003. The RAM 1003 also stores various programs and data required for the operation of the system 1000. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0141] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. A removable medium 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 1010 as needed so that computer programs read from it can be installed into storage section 1008 as needed.
[0142] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs the functions defined above in the system of this invention.
[0143] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0145] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described modules can also be housed in a processor; for example, a processor can be described as including a matching module, a deduplication module, and a recombination module. The names of these modules do not necessarily limit the module itself; for example, the recombination module can also be described as "a module that recombines modules according to relationships."
[0146] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include:
[0147] In response to identifying user analysis requests from user statements, request features are extracted from the user analysis requests, and the request features are matched with a data table. The request features include user intent, dimensions, and metrics.
[0148] In response to matching the target data table, a set of query statements corresponding to the requested features is generated;
[0149] Execute the set of query statements to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator.
[0150] Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and to generate natural language analysis results of the user analysis demand corresponding to the demand feature.
[0151] In another aspect, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the above embodiments.
[0152] According to the technical solutions of the embodiments of the present invention, the present invention has the following advantages or beneficial effects:
[0153] By employing the demand features extracted from user analysis requests in user statements, matching data tables, and generating and executing query sets corresponding to the demand features, relevant data scattered across multiple data tables can be quickly and accurately matched. This allows for the identification of the data source for online consultation data analysis and the acquisition of the corresponding dataset. Numerical analysis is then performed on this dataset, filtering out the target dimensions and target dimension values that influence each indicator within the demand features from the dimensions and dimension values in the dataset. Corresponding natural language analysis results are generated, providing both the data analysis results and corresponding natural language explanations for the user analysis requests. This enables rapid and accurate analysis and location of the causes of digital operational problems in user analysis requests, providing interpretations and suggestions based on the data analysis. This enhances the analytical capabilities of online medical consultation data, facilitating intelligent analysis and operational decision-making for online consultations on the platform, improving operational methods, and enhancing user experience.
[0154] The specific embodiments described herein do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for analyzing medical history data, characterized in that, include: In response to identifying user analysis requests from user statements, request features are extracted from the user analysis requests, and the request features are matched with a data table. The request features include user intent, dimensions, and metrics. In response to matching the target data table, a set of query statements corresponding to the requested features is generated; Execute the set of query statements to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator. Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and to generate natural language analysis results of the user analysis demand corresponding to the demand feature.
2. The method according to claim 1, characterized in that, The method is based on a predefined consultation knowledge base, which includes definitions of consultation terminology and table structure definitions of data tables. The definitions of consultation terminology include definitions of user intent, dimensions, and metrics. Extracting request features from the user analysis requests and matching these features with the data table includes: The user's analysis request is segmented into words, and the segmentation results are compared with the professional medical terminology definitions in the medical knowledge base to calculate the first similarity. In response, user intents, dimensions, and indicators with a first similarity greater than the first similarity threshold are selected and combined to obtain the request features. The second similarity is calculated between each of the stated claim features and the table structure definition of the data table, and the data tables with a second similarity greater than the second similarity threshold are selected as the target data tables.
3. The method according to claim 2, characterized in that, The consultation knowledge base includes table structure definitions and mappings to query statement set templates; In response to a match with the target data table, a set of query statements corresponding to the requested features is generated, including: Matching is performed based on the table structure definition of the target data table and the mapping between the table structure definition and the query statement set template; In response to a match to the mapping, a set of query statements for the target data table is generated based on the query statement set template corresponding to the mapping; In response to the failure to match the mapping, a large model prompt is generated based on the user's analysis request and its request characteristics. The large model prompt is then input into the question-and-answer large model to generate a set of query statements for the target data table.
4. The method according to claim 3, characterized in that, Based on the user analysis requests and their characteristics, a large model prompt is generated. This large model prompt is then input into the question-and-answer large model to generate a query statement for the target data table, including: The user analysis requests and their characteristics, along with the table structure definition of the target data table, are used as the first major model prompts for input into the question-and-answer model; The question-answering big model generates and outputs a first answer based on the prompts of the first big model and the question-and-answer knowledge base. The first answer includes a first set of query statements corresponding to the prompts of the first big model and its query result instances, dimensions and indicators in the first set of query statements, and natural language interpretation of the first set of query statements. In response to the user's confirmation of the first answer, the first set of query statements is used as the query statement for the target data table; In response to the user's correction instruction for the first answer, the large model generates a second large model prompt based on the first large model prompt, the correction prompt, and the consultation knowledge base. Based on the second large model prompt and the consultation knowledge base, a second answer is generated and output. The second answer includes the second large model prompt, a second query statement set corresponding to the second large model prompt and its query result instance, the dimensions and indicators in the second query statement set, and the natural language interpretation of the second query statement set. In response to the user's confirmation of the second answer, the second set of query statements is used as the query statement for the target data table.
5. The method according to claim 1, characterized in that, The indicator in the user's request is the timeliness of consultation; the dimensions in the user's request include consultation scenario, consultation channel, consultation category, consultation department, and consultation time period; and the user's intent includes the target dimension and target dimension value of the timeliness of consultation. Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that influence each indicator in the claimed feature from the dimensions and dimension values in the dataset, including: Obtain the consultation timeliness value and the dimension value of the dimension from the dataset; Construct a combination of multiple said dimension values that corresponds to the consultation timeliness value; Based on the consultation timeliness value corresponding to each combination of dimension values, calculate the influence factor of each combination of dimension values on each consultation timeliness value; Based on the magnitude of the influencing factors, the combination of target dimension values affecting the timeliness of patient reception is determined, and the target dimension and the target dimension value are obtained.
6. The method according to claim 1, characterized in that, The metrics mentioned in the user requests are aggregated metrics calculated by aggregating order volume and number of doctors. The dimensions mentioned in the user requests include consultation scenarios, consultation channels, consultation categories, consultation departments, and consultation time periods. The metrics include consultation timeliness. The user intent includes target dimensions and target dimension values that affect the aggregated metrics. Numerical analysis is performed using the dataset to filter out the target dimensions and target dimension values that influence each indicator in the claimed feature from the dimensions and dimension values in the dataset, including: The order volume and doctor volume values are obtained from the dataset, and the aggregated index value of the aggregated index is obtained through the aggregation calculation. Obtain the dimension values of the dimension from the dataset, and construct a dimension value combination consisting of multiple dimension values and corresponding to the aggregated index value; Based on the aggregated index value corresponding to the combination of values of each dimension, calculate the influence factor of the combination of values of each dimension on the aggregated index; Based on the magnitude of the influencing factors, determine the combination of target dimension values that affect the aggregated index, and obtain the target dimension and the target dimension value.
7. The method according to claim 5 or 6, characterized in that, Generate natural language analysis results for the user analysis request corresponding to the requested features, including: Based on the indicators, the influencing factors, and their corresponding target dimensions and target dimension values, a corresponding third major model suggestion is generated; The third major model prompt is input into the pre-trained large model for analyzing consultation data to obtain the natural language analysis results of the statistical claims of the indicators.
8. A medical consultation data analysis device, characterized in that, For use in an internet healthcare platform, the device includes: The semantic matching module is used to respond to the user analysis request identified from the user statement, extract the request features from the user analysis request, and match the request features with the data table. The request features include user intent, dimensions, and metrics. The statement generation module is used to generate a set of query statements corresponding to the requested features in response to matching the target data table. The data acquisition module is used to execute the query statement set to query the target data table and obtain the dataset corresponding to the demand feature. The dataset includes the dimension value corresponding to the dimension and the indicator value corresponding to the indicator. The data analysis module is used to perform numerical analysis on the dataset, filter out the target dimensions and target dimension values that affect each indicator in the demand feature from the dimensions and dimension values in the dataset, and generate natural language analysis results of the user analysis demand corresponding to the demand feature.
9. An electronic processing device, characterized in that, include: One or more processors; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.