Interaction Method and System for Human-Machine Collaborative Agent
By training a general entity relationship recognition model in a human-computer collaborative agent and optimizing the model according to regional differences, the adaptation problem caused by the diversity of medical data is solved, and more efficient and accurate human-computer interaction is achieved.
Patent Information
- Application Number
- CN202510480541.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Due to the complexity and diversity of medical data, human-machine collaborative agents face difficulties in adapting to and processing medical data of different types and sources, resulting in limited interaction efficiency and accuracy.
By collecting medical data from multiple regions, training a general entity relationship recognition model, and calculating the loss function based on the data differences between regions, targeted model training is carried out for each region, and finally deploying the optimal model of each region for human-computer interaction.
It improves the accuracy and adaptability of human-computer interaction, so that the system can better understand users' medical needs and provide more accurate medical service information, thereby improving interaction efficiency and user experience.
Smart Images

Figure CN119993547B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an interaction method and system for a human-machine collaborative agent. Background Art
[0002] With the development of technology, especially the breakthroughs in fields such as artificial intelligence, robotics, the Internet of Things, and big data, the medical industry is also constantly changing. The complexity of medical devices and information systems is increasing continuously, and clinical workers are faced with a large number of data processing and decision-making tasks. Against this background, the human-machine collaborative agent emerges as the times require and becomes an important means to improve medical efficiency and quality.
[0003] The human-machine collaborative agent is an innovative medical model that emphasizes the deep cooperation and interaction between humans and intelligent systems. In this model, the agent provides comprehensive support and assistance to doctors by means of advanced technologies such as machine learning, natural language processing, and computer vision. Whether in the process of disease diagnosis, by analyzing a large amount of medical data to assist doctors in making more accurate judgments; or in the selection of treatment plans, using complex algorithms to provide personalized treatment suggestions for doctors. This collaborative model not only makes up for the deficiencies of humans in information processing and computing power, but also enhances the scientific nature and accuracy of decision-making. At the same time, human-machine collaboration also has the ability of dynamic adjustment and can optimize the treatment plan in real time according to the specific situation of the patient and the treatment progress.
[0004] However, due to the complexity and diversity of medical data, different types of data require different processing methods and technologies. In addition, considering the wide range of data collection sources, including different hospitals, different devices, different medical staff, etc., the quality of medical data is uneven, which further increases the difficulty for the agent to adapt. Summary of the Invention
[0005] The object of the present invention is to: reduce the difficulty of human-machine interaction.
[0006] To achieve the above object, the present invention provides an interaction method for a human-machine collaborative agent, including:
[0007] Collect medical data from multiple regions and perform preprocessing;
[0008] Train a general entity relationship recognition model for all regions, where the medical data is input and the medical entities and the relationships between medical entities are output;
[0009] Calculate the loss function for each region to perform its corresponding model training based on the differences in medical data between regions, and use the parameters of the general entity relationship recognition model as pre-trained weights. For each region, train an optimal model using the medical data of that region and the loss function corresponding to that region, and deploy the optimal models of each region to the corresponding regions for human-computer interaction.
[0010] The present invention first trains a general entity relationship recognition model for all regions as a basic model, which can learn the common features of entities and relationships in medical data as a whole. Then, based on the differences in medical data between regions, calculate the loss functions corresponding to each region, and use the data and loss functions of each region to optimize the general model, which can make the final obtained optimal models of each region better adapt to the characteristics and needs of local medical data. Finally, deploy the optimal models of each region to the corresponding regions for human-computer interaction, which can improve the interaction efficiency between the medical information system and users. The system can more accurately understand user inputs (such as querying patient medical records, recommending treatment plans, etc.) and give answers and suggestions more in line with local medical practices, thereby enhancing the user experience and medical work efficiency.
[0011] Preferably, before calculating the differences in medical data between regions, it further includes:
[0012] Use a word segmentation tool to perform word segmentation on the text content in the medical data of all regions and extract keywords.
[0013] The medical data of different regions may come from different institutions such as hospitals and clinics, and their text formats and expression methods may vary. After word segmentation and keyword extraction, the text data from different sources can be unified to a standard vocabulary level.
[0014] Preferably, the process of obtaining the loss function includes:
[0015] Traverse the keywords in all medical data, select any region as the target region, and any medical data in the target region as the target data. Screen out the valid keywords that appear in the target data, normalize the regional differences of all valid keywords, and use the product of the sum of the normalized regional differences of valid keywords and the cross-entropy loss as the loss of incorrect recognition of the target data in the target region.
[0016] By calculating the differences between keywords in the target region and all other regions, it can reflect the uniqueness of keyword distribution in different regions. This difference loss can help the model better capture the feature differences between regions and avoid the model using the same processing method for all regions.
[0017] Among them, by determining whether the keyword appears in the target data, in this way, the model can only focus on the keyword differences related to the current medical data, avoiding the introduction of irrelevant noise information.
[0018] Preferably, the process of obtaining the regional difference includes:
[0019] Obtain the number of times the keyword appears in the target region, and obtain the total number of times the keyword appears in all regions;
[0020] Calculate the ratio of the number of times the keyword appears in the target region to the total number of times the keyword appears in all regions, and obtain the difference between the keyword in the target region and all other regions.
[0021] By calculating the difference between a keyword in a certain region and all other regions, the distribution of the keyword in different regions can be intuitively understood. If the proportion of the number of times a keyword appears in a certain region to its total number of appearances is relatively high, it indicates that the keyword has a high concentration in this region, which may reflect significant differences in the attention, activity frequency, or resource distribution in the relevant field in this region compared to other regions.
[0022] Preferably, the process of obtaining the regional difference of the keyword further includes:
[0023] Calculate the word frequency of the keyword appearing in the target region, and calculate the relative frequency of the keyword appearing among regions;
[0024] Take the product of the word frequency of the keyword appearing in the target region and the relative frequency of the keyword appearing among regions as the difference between the keyword in the target region and all other regions.
[0025] Taking the product of the word frequency and the relative frequency as the difference value takes into account both the attention of the keyword in a specific region and its universality within the whole range. When a keyword has a high word frequency in a certain region and a low relative frequency in other regions, its product value will be large, thus highlighting the uniqueness and difference of the keyword in this region.
[0026] Preferably, the process of obtaining the word frequency of the keyword appearing in the target region is:
[0027] Count the total number of times the keyword appears in the target region, and take the ratio of the number of times the keyword appears in the target region to the total number of times the keyword appears in the target region as the word frequency of the keyword appearing in the target region.
[0028] Preferably, the process of obtaining the relative frequency of the keyword appearing among regions includes:
[0029] Calculate the total number of regions in which the keyword appears among all regions, substitute the ratio of the total number of all regions to the total number of regions in which the keyword appears among all regions into the logarithmic function with base 2 for calculation, and the calculation result is the relative frequency of the keyword appearing among regions.
[0030] Preferably, the general entity relationship recognition model is constructed based on the Bert model.
[0031] Preferably, when the training reaches the preset number of training times or the loss value of the model is less than the preset loss threshold, stop the model training. After the training is completed, select the optimal model corresponding to the region according to the precision rate of the model.
[0032] In a second aspect, an interaction system for a human-machine collaborative agent includes: a processor and a memory, where the memory stores computer program instructions, and when the computer program instructions are executed by the processor, any one of the interaction methods for the human-machine collaborative agent is implemented.
[0033] Compared with the prior art, the interaction method and system for a human-machine collaborative agent in the embodiments of the present invention have the beneficial effects that:
[0034] By collecting medical data from multiple regions and training a general entity relationship recognition model, the common characteristics of medical data in different regions can be better captured. At the same time, targeted training and optimization are carried out according to the differences in medical data in each region, which can improve the accuracy and adaptability of the model in a specific region.
[0035] Using the model optimized by regional differences for human-machine interaction can more accurately understand the medical needs of users, provide more accurate medical service information, and thus improve the interaction efficiency and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the method of steps S1 - S3 in the interaction method for a human-machine collaborative agent in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.
[0038] As Figure 1 shown, the interaction method for a human-machine collaborative agent in the preferred embodiment of the embodiments of the present invention includes steps S1 - S3, specifically as follows:
[0039] S1: Collect medical data from multiple regions and perform preprocessing.
[0040] In one embodiment, an electronic health record system or a medical data platform is reasonably used to collect medical data from different regions. This data includes medical records, diagnostic results, treatment plans, etc. In addition, the regional information where the data is collected is also recorded, such as the south or the north (this is because medical data in different regions may vary due to regional differences).
[0041] The collected data is pre - processed, mainly by deleting the data with incomplete information or ambiguous content.
[0042] Then, the collected data is labeled. These labels include entities in the data (such as "diabetes" as a disease entity and "insulin" as a drug entity) and the relationships between entities (such as the treatment relationship between "diabetes" and "insulin").
[0043] By annotating such entities and entity relationships to the data, the medical data can be better analyzed.
[0044] S2: Train a general entity - relationship recognition model for all regions. Among them, the medical data is input, and the medical entities and the relationships between medical entities are output.
[0045] Medical data in different regions may vary in details, but the basic medical knowledge and entity relationships are similar. By using this general model as a starting point, training from scratch can be avoided, and the training speed of the subsequent model for a specific region can be greatly accelerated.
[0046] In one embodiment, the data collected in the above - mentioned S1 is used to train a general entity - relationship recognition model. All the collected medical data is input, and after processing, the relationships between medical entities and medical entities are output. For example, when inputting a medical text describing a patient's symptoms and diagnostic results, the model can identify medical entities such as "symptoms" and "diseases", and determine relationships such as "causing" between "symptoms" and "diseases".
[0047] During the training process of the general entity - relationship recognition model, multiple model architectures can be selected, such as Bert (Bidirectional Encoder Representations from Transformers, bidirectional encoder representations based on transformers) and BiLSTM (Bidirectional Long Short - Term Memory, bidirectional long - short - term memory network). These models perform well in natural language processing tasks, can effectively process the text information in medical data, and extract entities and relationships.
[0048] After the above - mentioned training process, a general entity - relationship recognition model for different regions is finally obtained.
[0049] S3: Calculate the loss function for each region to perform its corresponding model training based on the differences in medical data between regions, and use the parameters of the general entity relationship recognition model as pre-trained weights. For each region, train the optimal model using the medical data of that region and the loss function corresponding to that region, and deploy the optimal models of each region to the corresponding regions for human-computer interaction.
[0050] In the southern region, the climate is hot and humid, which is prone to breeding mosquitoes and flies, leading to the spread of diseases such as vesicular stomatitis. Skin diseases such as eczema are relatively common; in the northern region, the winter temperature is relatively low and the air is dry, and people are more likely to suffer from upper respiratory tract infectious diseases such as influenza and pneumonia. Dry climate is likely to cause skin problems such as chapped skin and eczema.
[0051] Jieba segmentation is a Chinese word segmentation tool used to split Chinese text into meaningful words or keywords. By using Jieba segmentation, the text content in the medical data of different regions is transformed into keywords.
[0052] After processing the medical data using Jieba segmentation, a series of words will be obtained. However, not all words play an equally important role in model construction. Some common but meaningless words (such as stop words like "de", "shi", "zai", etc.) will interfere with the analysis. Therefore, it is necessary to further screen out the keywords that can represent the core content of the text.
[0053] These keywords can represent the common diseases and symptoms in that region, which helps the model better understand and identify the disease characteristics in different regions.
[0054] In medical data analysis, if the frequency of a keyword related to a certain disease appears significantly higher in a certain region than in other regions, it may mean that the incidence of this disease in this region is higher, the medical focus is different, or there are other special circumstances. Moreover, the greater the difference, the greater the recognition error loss that may occur if the data of this region is confused with that of other regions during the recognition process, because the performance of this keyword in this region is significantly different from that in other regions.
[0055] In one embodiment, among all the medical data, select one type of keyword, calculate the number of times this keyword appears in each region and the total number of times this keyword appears in all regions, and then calculate the difference of this keyword between regions. Exemplarily, select any one region as the target region and any one medical data in the target region as the target data.
[0056] Exemplarily, the th keyword in the th region and the differences between it and all other regions satisfy the relational expression as:
[0057]
[0058] Wherein, is the difference between the th keyword between the th region and all other regions, is the th keyword appears times in the th region, is the total number of times the th keyword appears in all regions.
[0059] Among them, if is larger, it means that the frequency of this keyword appearing in this region is much higher than that in other regions, indicating that this keyword is unique in this region and has a large difference compared with other regions; on the contrary, is smaller, it means that the frequency of this keyword appearing in this region is similar to that in other regions, with a small difference.
[0060] It should be noted that relying only on the number of times a keyword appears in a certain region cannot comprehensively measure the differences of keywords in different regions. Therefore, it is also necessary to consider the distribution of keywords in all regions. If a keyword appears only in a few regions, its frequency will be higher, indicating that it is more unique in these regions; if a keyword appears in all regions, its frequency will be lower, indicating that it is more common.
[0061] In one embodiment, still taking the th keyword as an example, first count the total number of times the keyword appears in the th region, and take the ratio of the number of times the th keyword appears in the th region to the total number of times the keyword appears in the th region as the word frequency of the th keyword in the th region. The word frequency satisfies the relational expression:
[0062]
[0063] Wherein, is the word frequency of the th keyword in the th region, is the th keyword appears times in the th region, is the total number of times the keyword appears in the th region.
[0064] Then, use a vector function to judge the Whether the keyword appears in the th region, further sum over all regions to obtain the number of regions where the keyword appears, and then calculate the relative frequency of the
[0065]
[0066] keyword among regions. Then the relative frequency of the th keyword among regions satisfies the relationship: where is the relative frequency of the th keyword among regions, is the total number of all regions, is the vector function representing the total number of regions where the
[0067] If a keyword appears in all regions, then , and at this time, is close to 0, indicating that the frequency of this keyword in different regions is very low because it appears in all regions and has no regional characteristics; conversely, if a keyword appears only in a few regions, then , and at this time, is greater than 0, indicating that the frequency of this keyword in different regions is very high because it appears only in a few regions and has regional characteristics.
[0068] The above focuses on the comparison of the occurrence times of different keywords within a single region, and is used to measure the importance or commonness of the keyword in this region; considers the occurrence distribution of the keyword among multiple regions, reflecting the universality or scarcity of the keyword among different regions, and the focus is on the keyword occurrence characteristics across regions.
[0069] Finally, according to and calculate the difference between the th keyword in the th region and all other regions, that is, the relationship is satisfied as:
[0070]
[0071] where is the difference between the th keyword in the th region and all other regions, is the The word frequency of a kind of keyword appearing in the th region, and
[0072] is the relative frequency of the appearance of a
[0073] kind of keyword among different regions. In one embodiment, considering the differences in keywords of medical data in different regions, the adaptability of the model in different regions is improved. Traverse the keywords of all medical data, select any region as the target region, and any medical data in the target region as the target data. Screen out the effective keywords that appear in the target data, normalize the regional differences of all effective keywords, and take the product of the sum of the normalized regional differences of the effective keywords and the cross-entropy loss as the loss of misidentifying the target data in the target region.
[0074] Then the loss calculated above satisfies the relational expression:
[0075]
[0076] In the formula, is the loss of misidentifying the th medical data in the th region, represents the normalization process, is the sum of the differences between the th region and all other regions for a kind of keyword, is the cross-entropy loss during the training of the th medical data in the
[0077] th region, and
[0078] is used to determine whether the
[0079] th kind of keyword appears in the th medical data. According to the above related calculation operations, the losses of misidentifying all medical data in all regions can be obtained in the same way. In the training of the medical entity relationship recognition model, if the difference of a certain medical data between different regions is larger, then the importance of this data for a specific region is higher. Therefore, when the model is trained in this specific region, if the prediction of this medical data is incorrect, the resulting loss will also increase accordingly.In one embodiment, the parameters of the general entity relationship recognition model in S2 above are used as pre-trained weights. For each region, the medical data of the region and the corresponding loss function of the region are used to train the recognition model. When the training reaches the set maximum number of training times (set according to the actual situation) or the loss value of the model is less than the preset loss threshold (in practical applications, the loss threshold can be adjusted according to past experience and the results of similar tasks), the model training is stopped. After the training is completed, the optimal model corresponding to the region is selected according to the precision rate of the model, and the optimal models of each region are deployed to the corresponding regions for human-computer interaction.
[0080] In practical applications, when entity relationship recognition of medical data is required in the region, the locally deployed optimal model can be used for processing to implement the human-computer interaction function, such as assisting doctors in medical information extraction, medical record analysis, etc.
[0081] The system includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the interaction method for the human-machine collaborative agent according to the first aspect of the present invention is implemented.
[0082] The system further includes other components well-known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.
[0083] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the counting principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.
Claims
1. An interactive method for a human-machine collaborative agent, characterized in that: include: Collect medical data from multiple regions and perform pre-processing; Training a general entity relationship recognition model for all regions, wherein the medical data is input and the medical entities and the relationships between the medical entities are output; Based on the difference in medical data between regions, the loss function for model training corresponding to each region is calculated, and the parameters of the universal entity relationship recognition model are used as pre-training weights. For each region, the medical data of the region and the loss function corresponding to the region are used to train the optimal model, and the optimal model of each region is deployed to the corresponding region for human-computer interaction; Before calculating the differences in medical data between regions, it also includes: Use word segmentation tools to segment text content in medical data from all regions and extract keywords; The calculated loss satisfies the relationship: , In the formula, For the Region No. The loss of medical data identification errors, represents normalization processing, For the Keywords in the The sum of the differences between a region and all other regions, For the Region No. The cross entropy loss when training medical data, To judge the Is the keyword in the first Appear in medical data; The process of obtaining regional differences in keywords includes: Calculate the frequency of the keyword in the target area, and calculate the relative frequency of the keyword in each area; The product of the keyword's frequency in the target region and the keyword's relative frequency among regions is taken as the difference between the keyword in the target region and all other regions; The word frequency satisfies the relationship: , In the formula, For the Keywords in the The frequency of words appearing in the region, For the Keywords in the The number of times a region appears, For the The total number of keywords appearing in each region; The relative frequency satisfies the relationship: , In the formula, For the The relative frequency of keywords in different regions. is the total number of all regions, is a vector function, indicating the The total number of regions where this keyword appears in all regions. Represents the base 2 logarithmic function.
2. The interactive method for human-machine collaborative intelligent agent according to claim 1, characterized in that: The general entity relationship recognition model is constructed based on the Bert model.
3. The interactive method for human-machine collaborative intelligent agent according to claim 2, characterized in that: When the training reaches the preset number of training times or the model loss value is less than the preset loss threshold, the model training is stopped. After the training is completed, the optimal model corresponding to the region is selected according to the model accuracy.
4. An interactive system for a human-machine collaborative agent, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, an interaction method for a human-machine collaborative intelligent body according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Rapid fine tuning method and system for address classification based on text classification
CN118535737A
Medical record data processing method and system based on multiple agents
CN119207685A