Interaction method and system for man-machine collaborative agent

By training a general entity relationship recognition model in a human-computer collaborative agent and optimizing the model based on the differences in medical data between regions, the adaptation problems of medical data complexity and diversity are solved, and interaction efficiency and accuracy are improved.

CN119993547AActive Publication Date: 2025-05-13WENZHOU MEDICAL UNIV

Patent Information

Application Number
CN202510480541.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The complexity and diversity of medical data increase the difficulty of adapting to human-machine collaborative agents, resulting in limited interaction efficiency and accuracy.

Method used

Collect medical data from multiple regions, train general entity relationship identification models, and calculate loss functions based on the differences in data between regions, and optimize the model to adapt to local characteristics.

Benefits of technology

It improves the accuracy and adaptability of the model in specific regions, enhances the efficiency and accuracy of human-computer interaction, can better understand user needs and provide personalized medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993547A_ABST
    Figure CN119993547A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses an interaction method and system for a man-machine cooperative agent, and the method comprises the steps: collecting medical data of a plurality of regions, and carrying out the preprocessing; general entity relationship recognition models of all regions are trained, the medical data are input, and the relationship between the medical entities is output; calculating a loss function of each region for performing respective corresponding model training based on the difference of the medical data between the regions, taking the parameters of the general entity relationship identification model as pre-training weights, training an optimal model by using the medical data of the region and the loss function corresponding to the region for each region, and obtaining the optimal model; and deploying the optimal model of each region to the corresponding region for man-machine interaction. By processing and analyzing the medical data, optimizing the model training process and deploying models for different regions, high efficiency and accuracy of human-computer interaction in the medical field are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an interaction method and system for a human-machine collaborative intelligent body. Background Art

[0002] With the development of science and technology, especially breakthroughs in artificial intelligence, robotics, the Internet of Things and big data, the medical industry is also changing. The complexity of medical equipment and information systems is increasing, and clinical workers are facing a large number of data processing and decision-making tasks. In this context, human-machine collaborative intelligent agents have emerged as an important means to improve medical efficiency and quality.

[0003] Human-machine collaborative agents are an innovative medical model that emphasizes deep cooperation and interaction between humans and intelligent systems. In this model, agents provide doctors with all-round support and assistance with advanced machine learning, natural language processing, computer vision and other technical means. Whether in the diagnosis of a disease, doctors can make more accurate judgments by analyzing massive medical data, or in the selection of treatment plans, doctors can be provided with personalized treatment suggestions using complex algorithms. This collaborative model not only makes up for the shortcomings of humans in information processing and computing capabilities, but also enhances the scientificity and accuracy of decision-making. At the same time, human-machine collaboration also has the ability to dynamically adjust, and can optimize treatment plans in real time according to the patient's specific situation and treatment progress.

[0004] However, due to the complexity and diversity of medical data, different types of data require different processing methods and technologies. In addition, considering that the sources of data collection are wide, including different hospitals, different equipment, different medical staff, etc., the quality of medical data is also uneven, which increases the difficulty of intelligent adaptation. Summary of the invention

[0005] The purpose of the present invention is to reduce the difficulty of human-computer interaction.

[0006] In order to achieve the above object, the present invention provides an interaction method for a human-machine collaborative intelligent agent, comprising: Collect medical data from multiple regions and perform pre-processing; Training a general entity relationship recognition model for all regions, wherein the medical data is input and the medical entities and the relationships between the medical entities are output; Based on the differences in medical data between regions, the loss function for model training corresponding to each region is calculated, and the parameters of the universal entity relationship recognition model are used as pre-training weights. For each region, the medical data of the region and the corresponding loss function of the region are used to train the optimal model, and the optimal model of each region is deployed to the corresponding region for human-computer interaction.

[0007] The present invention first trains a universal entity relationship recognition model for all regions as a basic model, which can learn the common characteristics of entities and relationships in medical data as a whole. Then, based on the differences in medical data between regions, the corresponding loss function of each region is calculated, and the universal model is optimized using the data and loss function of each region. This can make the optimal model for each region better adapt to the local medical data characteristics and needs. Finally, the optimal model for each region is deployed to the corresponding region for human-computer interaction, which can improve the interaction efficiency between the medical information system and the user. The system can understand the user's input more accurately (such as querying the patient's medical history information, recommending treatment plans, etc.), and give answers and suggestions that are more in line with local medical practices, thereby improving user experience and medical work efficiency.

[0008] Preferably, before calculating the difference in medical data between regions, the method further includes: The word segmentation tool is used to segment the text content in the medical data of all regions and extract keywords.

[0009] Medical data from different regions may come from different hospitals, clinics and other institutions, and their text formats and expressions may vary. After word segmentation and keyword extraction, text data from different sources can be unified into a standard vocabulary level.

[0010] Preferably, the process of obtaining the loss function includes: Traverse the keywords of all medical data, select any region as the target region, and any medical data in the target region as the target data, filter out the valid keywords that appear in the target data, normalize the regional differences of all valid keywords, and take the product of the normalized sum of regional differences of valid keywords and the cross entropy loss as the loss of target data recognition error in the target region.

[0011] By calculating the difference between keywords in the target region and all other regions, the uniqueness of keyword distribution in different regions can be reflected. This difference loss can help the model better capture the feature differences between regions and avoid the model using the same treatment for all regions.

[0012] Among them, by judging whether the keywords appear in the target data, the model can focus only on the keyword differences related to the current medical data and avoid introducing irrelevant noise information.

[0013] Preferably, the process of obtaining the regional differences includes: Get the number of times the keyword appears in the target region, and get the total number of times the keyword appears in all regions; The ratio of the number of times the keyword appears in the target region to the total number of times the keyword appears in all regions is calculated to obtain the difference between the keyword in the target region and all other regions.

[0014] By calculating the difference between a keyword in a certain region and all other regions, you can intuitively understand the distribution of the keyword in different regions. If the number of times a keyword appears in a certain region accounts for a higher proportion of its total number of appearances, it means that the keyword has a higher concentration in this region, which may reflect that the region has significant differences from other regions in terms of attention, activity frequency or resource distribution in related fields.

[0015] Preferably, the process of obtaining the regional differences of the keywords further includes: Calculate the frequency of the keyword in the target area, and calculate the relative frequency of the keyword in each area; The product of the frequency of the keyword appearing in the target area and the relative frequency of the keyword appearing in various regions is taken as the difference between the keyword in the target area and all other regions.

[0016] The product of word frequency and relative frequency is used as the difference value, which takes into account both the popularity of the keyword in a specific region and its prevalence in the entire range. When a keyword has a high word frequency in a certain region and a low relative frequency in other regions, its product value will be larger, thus highlighting the uniqueness and difference of the keyword in the region.

[0017] Preferably, the process of obtaining the frequency of the keyword in the target area is as follows: The total number of keywords appearing in the target area is counted, and the ratio of the number of times the keyword appears in the target area to the total number of keywords appearing in the target area is used as the word frequency of the keyword appearing in the target area.

[0018] Preferably, the process of obtaining the relative frequency of the keywords in various regions includes: Calculate the total number of regions where the keyword appears in all regions, substitute the ratio of the total number of all regions to the total number of regions where the keyword appears in all regions into a logarithmic function with base 2 for calculation, and the result of the calculation is the relative frequency of the keyword appearing in each region.

[0019] Preferably, the general entity relationship recognition model is constructed based on the Bert model.

[0020] Preferably, when the training reaches a preset number of training times or the loss value of the model is less than a preset loss threshold, the model training is stopped. After the training is completed, the optimal model corresponding to the region is selected according to the accuracy of the model.

[0021] In a second aspect, an interactive system for a human-machine collaborative intelligent body comprises: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, any one of the interactive methods for a human-machine collaborative intelligent body is implemented.

[0022] Compared with the prior art, the interactive method and system for human-machine collaborative intelligent body in the embodiment of the present invention has the following beneficial effects: By collecting medical data from multiple regions and training a general entity relationship recognition model, we can better capture the common features of medical data from different regions. At the same time, targeted training and optimization based on the differences in medical data from different regions can improve the accuracy and adaptability of the model in specific regions.

[0023] Using models optimized for regional differences for human-computer interaction can more accurately understand users' medical needs and provide more precise medical service information, thereby improving interaction efficiency and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a method flow chart of steps S1 to S3 in the interaction method for a human-machine collaborative intelligent body according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0026] like Figure 1 As shown, the interaction method for a human-machine collaborative intelligent agent in a preferred embodiment of the present invention includes steps S1 to S3, which are specifically as follows: S1: Collect medical data from multiple regions and perform preprocessing.

[0027] In one embodiment, an electronic health record system or a medical data platform is used to collect medical data from different regions, including medical records, diagnosis results, treatment plans, etc. In addition, the region where the data was collected is also recorded, such as the south or the north (this is because the medical data in different regions may be slightly different due to geographical differences).

[0028] The collected data is preprocessed, mainly to delete the data with incomplete information or ambiguous content.

[0029] Then, the collected data is labeled. These labels include entities in the data (such as "diabetes" as a disease entity and "insulin" as a drug entity) and the relationships between entities (such as the therapeutic relationship between "diabetes" and "insulin").

[0030] By labeling the data with such entities and entity relationships as described above, medical data can be better analyzed.

[0031] S2: Train a universal entity relationship recognition model for all regions, wherein the medical data is input and medical entities and relationships between the medical entities are output.

[0032] Medical data from different regions may differ in details, but the basic medical knowledge and entity relationships are similar. By using this general model as a starting point, we can avoid training from scratch and greatly speed up the subsequent model training for specific regions.

[0033] In one embodiment, the data collected by S1 is used to train a general entity relationship recognition model, and all collected medical data are input, and the medical entities and the relationships between medical entities are output after processing. For example, a medical text describing the patient's symptoms and diagnosis results is input, and the model can recognize medical entities such as "symptoms" and "diseases", and determine that there is a relationship such as "cause" between "symptoms" and "diseases".

[0034] In the process of training the general entity relationship recognition model, you can choose from a variety of model architectures, such as Bert (Bidirectional Encoder Representations from Transformers) and BiLSTM (Bidirectional Long Short-Term Memory). These models perform well in natural language processing tasks and can effectively process text information in medical data and extract entities and relationships.

[0035] After the above training process, we finally obtain an entity relationship recognition model that is common to different regions.

[0036] S3: Based on the differences in medical data between regions, the loss function for model training corresponding to each region is calculated, and the parameters of the universal entity relationship recognition model are used as pre-training weights. For each region, the medical data of the region and the corresponding loss function of the region are used to train the optimal model, and the optimal model of each region is deployed to the corresponding region for human-computer interaction.

[0037] The climate in the south is hot and humid, which is easy to breed mosquitoes and flies, leading to the spread of diseases such as vesicular stomatitis. Skin diseases such as eczema are more common. In the north, the temperature is lower in winter and the air is dry, making people more susceptible to respiratory infections such as influenza and pneumonia. The dry climate can easily lead to skin problems such as chapped skin and eczema.

[0038] Jieba word segmentation is a Chinese word segmentation tool used to split Chinese text into meaningful words or keywords. By using Jieba word segmentation, the text content in medical data from different regions is converted into keywords.

[0039] After processing medical data using Jieba word segmentation, a series of words will be obtained. However, not all words play an equally important role in model construction. Some common but meaningless words (such as stop words like "的", "是", "在", etc.) will interfere with the analysis. Therefore, it is necessary to further screen out the keywords that can represent the core content of the text.

[0040] These keywords can represent the common diseases and symptoms in that region, which helps the model better understand and identify the disease characteristics in different regions.

[0041] In medical data analysis, if the frequency of a disease-related keyword appears significantly higher in a certain region than in other regions, it may mean that the incidence of this disease in that region is higher, the medical focus is different, or there are other special circumstances. Moreover, the greater the difference, the greater the potential loss of recognition errors if the data in this region is confused with that of other regions during the recognition process, because the performance of this keyword in this region is significantly different from that in other regions.

[0042] In one embodiment, among all medical data, one type of keyword is selected, and the number of times this keyword appears in each region and the total number of times this keyword appears in all regions are calculated, and then the difference of this keyword among regions is calculated. Exemplarily, any one region is selected as the target region, and any one medical data in the target region is used as the target data.

[0043] Exemplarily, the th keyword in the th region and the difference between it and all other regions satisfy the relational expression as: In the formula, is the difference between the th keyword in the th region and all other regions, is the number of times the th keyword appears in the th region, is the total number of times the th keyword appears in all regions.

[0044] Among them, if is larger, it indicates that the frequency of this keyword appears much higher in this region than in other regions, meaning that this keyword is unique in this region and has a large difference compared with other regions; on the contrary, The smaller it is, the more likely it is that the keyword appears in that region at a similar frequency to other regions, with a smaller difference.

[0045] It should be noted that relying solely on the number of times a keyword appears in a certain region is not enough to comprehensively measure the differences in keywords in different regions. Therefore, it is also necessary to consider the distribution of keywords in all regions. If a keyword only appears in a few regions, then its frequency will be higher, indicating that it is more unique in these regions; if a keyword appears in all regions, then its frequency will be lower, indicating that it is more common.

[0046] In one embodiment, still with the first Take the keyword as an example, first count the The total number of keywords appearing in the region, Keywords in the The number of times a region appears is the same as the The ratio of the total number of keywords in the regions is taken as the Keywords in the The word frequency that appears in a region. The word frequency satisfies the relationship: In the formula, For the Keywords in the The frequency of words appearing in the region, For the Keywords in the The number of times a region appears, For the The total number of keywords appearing in a region.

[0047] Then, a vector function is used to determine the Is the keyword in the first The number of regions where the keyword appears is further summed up to get the number of regions where the keyword appears, and then the number of regions where the keyword appears is calculated. The relative frequency of the keywords in each region, The relative frequency of keywords in different regions satisfies the following relationship: In the formula, For the The relative frequency of keywords in different regions. is the total number of all regions, is a vector function, indicating the The total number of regions where this keyword appears in all regions. Represents the base 2 logarithmic function.

[0048] If a keyword appears in all regions, then ,at this time, The value of is close to 0, which means that the frequency of this keyword in different regions is very low, because it has appeared in all regions and has no regional characteristics; on the contrary, if a keyword only appears in a few regions, then ,at this time, The value of is greater than 0, which means that the keyword appears frequently in different regions, because it only appears in a few regions and has regional characteristics.

[0049] Above The focus is on the comparison of the number of occurrences of different keywords in a single region, which is used to measure the importance or commonness of the keyword in the region; What is considered is the distribution of the keyword's occurrence among multiple regions, which reflects the prevalence or scarcity of the keyword among different regions, with the focus on the cross-regional keyword occurrence characteristics.

[0050] Finally, according to and Calculate the Keywords in the The difference between a region and all other regions, that is, the relationship is: In the formula, For the Keywords in the The difference between the region and all other regions, For the Keywords in the The frequency of words appearing in the region, For the The relative frequency of the keywords in different regions.

[0051] In one embodiment, the differences in keywords among medical data in different regions are taken into consideration, thereby improving the adaptability of the model in different regions.

[0052] Traverse the keywords of all medical data, select any region as the target region, and any medical data in the target region as the target data, filter out the valid keywords that appear in the target data, normalize the regional differences of all valid keywords, and take the product of the normalized sum of regional differences of valid keywords and the cross entropy loss as the loss of target data recognition error in the target region.

[0053] Then the loss calculated above satisfies the relationship: In the formula, For the Region No. The loss of medical data identification errors, represents normalization processing, For the Keywords in the The sum of the differences between a region and all other regions, For the Region No. The cross entropy loss when training medical data, To judge the Is the keyword in the first Appears in medical data.

[0054] According to the above The relevant calculation operations can be used to obtain the loss of incorrect identification of all medical data in all regions in a similar way.

[0055] In the training of the medical entity relationship recognition model, if the difference between different regions of a medical data is greater, the importance of the data to a specific region will be higher. Therefore, when the model is trained in this specific region, if the prediction of the medical data is wrong, the loss will also increase accordingly.

[0056] In one embodiment, the parameters of the general entity relationship recognition model in S2 above are used as pre-training weights. For each region, the medical data of the region and the loss function corresponding to the region are used to train the recognition model. When the training reaches the set maximum number of training times (set according to actual conditions) or the loss value of the model is less than a preset loss threshold (in actual applications, the loss threshold can be adjusted based on past experience and the results of similar tasks), the model training is stopped. After the training is completed, the optimal model corresponding to the region is selected according to the accuracy of the model, and the optimal model of each region is deployed to the corresponding region for human-computer interaction.

[0057] In actual applications, when the region needs to identify entity relationships in medical data, the optimal model deployed locally can be used for processing to realize human-computer interaction functions, such as assisting doctors in extracting medical information and analyzing medical records.

[0058] The system includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the interaction method for a human-machine collaborative intelligent agent according to the first aspect of the present invention is implemented.

[0059] The system also includes other components familiar to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.

[0060] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary counting personnel in this technical field, several improvements and substitutions can be made without departing from the counting principle of the present invention. These improvements and substitutions should also be regarded as the scope of protection of the present invention.

Claims

1. An interactive method for a human-machine collaborative agent, characterized in that: include: Collect medical data from multiple regions and perform pre-processing; Training a general entity relationship recognition model for all regions, wherein the medical data is input and the medical entities and the relationships between the medical entities are output; Based on the differences in medical data between regions, the loss function for model training corresponding to each region is calculated, and the parameters of the universal entity relationship recognition model are used as pre-training weights. For each region, the medical data of the region and the corresponding loss function of the region are used to train the optimal model, and the optimal model of each region is deployed to the corresponding region for human-computer interaction.

2. The interactive method for human-machine collaborative intelligent agent according to claim 1, characterized in that: Before calculating the difference in medical data between regions, it also includes: The word segmentation tool is used to segment the text content in the medical data of all regions and extract keywords.

3. The interactive method for human-machine collaborative intelligent agent according to claim 2, characterized in that: The process of obtaining the loss function includes: Traverse the keywords of all medical data, select any region as the target region, and any medical data in the target region as the target data, filter out the valid keywords that appear in the target data, normalize the regional differences of all valid keywords, and take the product of the normalized sum of regional differences of valid keywords and the cross entropy loss as the loss of target data recognition error in the target region.

4. The interactive method for human-machine collaborative intelligent agent according to claim 3, characterized in that: The process of obtaining the regional differences includes: Get the number of times the keyword appears in the target region, and get the total number of times the keyword appears in all regions; The ratio of the number of times the keyword appears in the target region to the total number of times the keyword appears in all regions is calculated to obtain the difference between the keyword in the target region and all other regions.

5. The interactive method for human-machine collaborative intelligent agent according to claim 4, characterized in that: The process of obtaining the regional differences of the keywords also includes: Calculate the frequency of the keyword in the target area, and calculate the relative frequency of the keyword in each area; The product of the frequency of the keyword appearing in the target area and the relative frequency of the keyword appearing in various regions is taken as the difference between the keyword in the target area and all other regions.

6. The interactive method for human-machine collaborative intelligent agent according to claim 5, characterized in that: The process of obtaining the frequency of the keyword in the target area is as follows: The total number of keywords appearing in the target area is counted, and the ratio of the number of times the keyword appears in the target area to the total number of keywords appearing in the target area is used as the word frequency of the keyword appearing in the target area.

7. The interaction method for a human-machine collaborative agent according to claim 6, characterized in that: The process of obtaining the relative frequency of the keywords in various regions includes: Calculate the total number of regions where the keyword appears in all regions, substitute the ratio of the total number of all regions to the total number of regions where the keyword appears in all regions into a logarithmic function with base 2 for calculation, and the result of the calculation is the relative frequency of the keyword appearing in each region.

8. The interactive method for human-machine collaborative intelligent agent according to claim 7, characterized in that: The general entity relationship recognition model is constructed based on the Bert model.

9. The interactive method for human-machine collaborative intelligent agent according to claim 8, characterized in that: When the training reaches the preset number of training times or the model loss value is less than the preset loss threshold, the model training is stopped. After the training is completed, the optimal model corresponding to the region is selected according to the model accuracy.

10. An interactive system for a human-machine collaborative agent, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an interaction method for a human-machine collaborative intelligent body according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Medical examination pre-training model construction method and identification method

    CN115759073A

  • Rapid fine tuning method and system for address classification based on text classification

    CN118535737A

  • Medical record data processing method and system based on multiple agents

    CN119207685A

  • Model training method and apparatus, text classification method and apparatus, computer device and medium

    WO2021189974A1

Cited By

  • Intelligent scheduling method and system for skin disease prevention and control resources

    CN121171521A