Processing method, device, computer equipment and storage medium for medical record classification model
By adjusting and training the medical record classification model and utilizing the target number of sample medical records and similarity differences, the problem of inaccurate classification in medical record management is solved, and efficient medical record classification management is achieved.
Patent Information
- Application Number
- CN202110932738.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-08-13
AI Technical Summary
The existing medical record management methods are difficult to effectively manage a large number of various medical records and lack efficient classification management methods.
By adjusting the medical record classification model, we obtain the target number of sample medical records and conduct training. By combining word similarity and prediction information differences, we adjust the sample medical record set to balance the number of categories, and further train the model to improve accuracy.
The accuracy of the medical record classification model is improved, the amount of calculation is reduced, the impact of differences in the number of categories on the model is weakened, and more efficient medical record classification management is achieved.
Smart Images

Figure CN114297374B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a method, apparatus, computer equipment, and storage medium for processing a medical record classification model. Background Art
[0002] In healthcare settings, medical records are records of medical personnel's activities, including the onset, development, and outcome of a patient's illness, as well as their examination, diagnosis, and treatment. Due to the large number and variety of medical records, and the increasing number of these types over time, there is an urgent need for a method to effectively manage medical records. Summary of the Invention
[0003] The embodiments of the present application provide a method, apparatus, computer device, and storage medium for processing a medical record classification model, which can improve the accuracy of the medical record classification model. The technical solution is as follows:
[0004] In one aspect, a method for processing a medical record classification model is provided, the method comprising:
[0005] Obtaining a first medical record classification model, where the first medical record classification model is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category;
[0006] A first sample medical record set is formed by first sample medical records belonging to the newly added category and second sample medical records belonging to the second category, where the second sample medical records are selected from sample medical records used to train the second medical record classification model, and the number of second sample medical records belonging to the same second category is a target number; the first medical record classification model is trained based on the first sample medical record set;
[0007] The target number of the first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set form a second sample medical record set;
[0008] Based on the second sample medical record set, the first medical record classification model is trained again.
[0009] In one possible implementation, the training of the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word includes:
[0010] The first medical record classification model is trained based on the difference between the sample label and the first predicted label of each first word in the same first category.
[0011] In another possible implementation, obtaining second prediction information for each second word in the fourth sample medical record based on the first medical record classification model includes:
[0012] For each of the second words, obtaining a third similarity between the second word and each of the first categories based on the first medical record classification model;
[0013] Based on the first medical record classification model and the third similarity corresponding to each of the first categories, each of the third similarities is normalized to obtain multiple second prediction labels for the second word, and the multiple second prediction labels constitute the second prediction information of the second word.
[0014] In another possible implementation, the method further includes:
[0015] Creating a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0016] For each of the second words, obtaining a fourth similarity between the second word and each of the first categories based on the third medical record classification model;
[0017] The training of the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word includes:
[0018] The first medical record classification model is trained based on the difference between the second prediction information and the annotation information of each second word, and the second difference information of each second word, wherein the second difference information includes multiple second similarity differences, and each second similarity difference includes the difference between the third similarity and the fourth similarity of the corresponding second word in one of the first categories.
[0019] In another possible implementation, the first medical record classification model includes weight features of the first categories; and obtaining, based on the first medical record classification model, a third similarity between the second word and each of the first categories includes:
[0020] Based on the first medical record classification model, a third word feature of the second word is obtained, and a third similarity between the third word feature and the weight feature of each first category is obtained.
[0021] In another possible implementation, the method further includes:
[0022] Creating a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0023] For each of the second words, obtaining a fourth word feature of the second word based on the third medical record classification model;
[0024] The training of the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word includes:
[0025] The first medical record classification model is trained based on the difference between the second prediction information and the annotation information of each second word, and the difference between the third word feature and the fourth word feature of each second word.
[0026] In another possible implementation, after retraining the first medical record classification model based on the second sample medical record set, the method further includes:
[0027] Based on the first medical record classification model, obtaining a first category to which each target word contained in the target medical record belongs;
[0028] Based on the first category to which each target word belongs, the first category to which the target medical record belongs is determined.
[0029] In another possible implementation, forming a second sample medical record set from the target number of the first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set includes:
[0030] Obtaining the medical record feature of each of the first sample medical records and the feature center of the first sample medical records in the first sample medical record set;
[0031] Based on the distance between the medical record feature of each first sample medical record and the feature center, selecting the target number of first sample medical records from the first sample medical record set, wherein the distance corresponding to the selected first sample medical records is smaller than the distances corresponding to the remaining first sample medical records in the first sample medical record set;
[0032] The selected first sample medical records and the second sample medical records in the first sample medical record set constitute a second sample medical record set.
[0033] In another possible implementation, obtaining the medical record feature of each of the first sample medical records and the feature center of the first sample medical records in the first sample medical record set includes:
[0034] For each of the first sample medical records, obtaining a word feature of each third word in the first sample medical record based on the first medical record classification model;
[0035] Acquire a medical record feature of the first sample medical record based on a word feature of each of the third words in the first sample medical record;
[0036] The feature center is obtained based on the medical record feature of each of the first sample medical records.
[0037] In another aspect, a device for processing a medical record classification model is provided, the device comprising:
[0038] an acquisition module, configured to acquire a first medical record classification model, wherein the first medical record classification model is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category;
[0039] a training module, configured to form a first sample medical record set from first sample medical records belonging to the newly added category and second sample medical records belonging to the second category, where the second sample medical records are selected from sample medical records for training the second medical record classification model, and the number of second sample medical records belonging to the same second category is a target number; and to train the first medical record classification model based on the first sample medical record set;
[0040] a forming module, configured to form a second sample medical record set by combining the target number of the first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set;
[0041] The training module is further used to retrain the first medical record classification model based on the second sample medical record set.
[0042] In one possible implementation, the first sample medical record set includes labeling information for each word contained in each sample medical record, the labeling information includes a plurality of sample labels, each sample label indicating whether the corresponding word belongs to one of the first categories; the training module includes:
[0043] an acquiring unit, configured to acquire, based on the first medical record classification model, first prediction information for each first word in a third sample medical record, the first prediction information comprising a plurality of first prediction labels, each first prediction label indicating a likelihood that the corresponding first word belongs to one of the first categories, the third sample medical record being any sample medical record in the first sample medical record set;
[0044] A training unit is used to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word.
[0045] In another possible implementation, the acquisition unit is used to obtain, for each first word, a first similarity between the first word and each first category based on the first medical record classification model; and normalize each first similarity based on the first medical record classification model and the first similarity corresponding to each first category to obtain multiple first prediction labels for the first word, and use the multiple first prediction labels to constitute the first prediction information of the first word.
[0046] In another possible implementation, the training unit is configured to train the first medical record classification model based on the difference between the sample label and the first predicted label of each first word in the same first category.
[0047] In another possible implementation, the acquisition module is further configured to acquire, for each first word, a second similarity between the first word and each second category based on the second medical record classification model;
[0048] The training unit is used to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the first difference information of each first word, wherein the first difference information includes multiple first similarity differences, and each first similarity difference includes the difference between the first similarity and the second similarity of the corresponding first word in a second category.
[0049] In another possible implementation, the first medical record classification model includes the weight feature of the first category; the acquisition unit is used to obtain the first word feature of the first word based on the first medical record classification model, and obtain the first similarity between the first word feature and the weight feature of each of the first categories.
[0050] In another possible implementation, the acquisition module is further configured to acquire, for each first word, a second word feature of the first word based on the second medical record classification model;
[0051] The training unit is used to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word.
[0052] In another possible implementation, the second sample medical record set includes labeling information for each word contained in each sample medical record, the labeling information includes a plurality of sample labels, each of the sample labels indicating whether the corresponding word belongs to one of the first categories; and the training module includes:
[0053] an acquiring unit, configured to acquire, based on the first medical record classification model, second prediction information for each second word in a fourth sample medical record, the second prediction information comprising a plurality of second prediction labels, each second prediction label indicating a likelihood that the corresponding second word belongs to one of the first categories, the fourth sample medical record being any sample medical record in the second sample medical record set;
[0054] A training unit is used to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word.
[0055] In another possible implementation, the acquisition unit is used to obtain, for each second word, the third similarity between the second word and each first category based on the first medical record classification model; based on the first medical record classification model and the third similarity corresponding to each first category, normalize each third similarity separately to obtain multiple second prediction labels for the second word, and use the multiple second prediction labels to constitute the second prediction information of the second word.
[0056] In another possible implementation, the apparatus further includes:
[0057] A creation module, configured to create a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0058] The acquisition module is further configured to acquire, for each second word, a fourth similarity between the second word and each first category based on the third medical record classification model;
[0059] The training unit is used to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the second difference information of each second word, wherein the second difference information includes multiple second similarity differences, and each second similarity difference includes the difference between the third similarity and the fourth similarity of the corresponding second word in one of the first categories.
[0060] In another possible implementation, the first medical record classification model includes the weight feature of the first category; the acquisition unit is used to obtain the third word feature of the second word based on the first medical record classification model, and obtain the third similarity between the third word feature and the weight feature of each first category.
[0061] In another possible implementation, the apparatus further includes:
[0062] A creation module, configured to create a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0063] The acquisition module is further configured to acquire, for each second word, a fourth word feature of the second word based on the third medical record classification model;
[0064] The training unit is used to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the difference between the third word feature and the fourth word feature of each second word.
[0065] In another possible implementation, the construction module is used to obtain the medical record characteristics of each first sample medical record and the characteristic center of the first sample medical record in the first sample medical record set; based on the distance between the medical record characteristics of each first sample medical record and the characteristic center, select the target number of first sample medical records from the first sample medical record set, wherein the distance corresponding to the selected first sample medical record is smaller than the distance corresponding to the remaining first sample medical records in the first sample medical record set; the selected first sample medical records and the second sample medical records in the first sample medical record set constitute a second sample medical record set.
[0066] In another possible implementation, the constituent module is used to obtain, for each first sample medical record, the word features of each third word in the first sample medical record based on the first medical record classification model; obtain the medical record features of the first sample medical record based on the word features of each third word in the first sample medical record; and obtain the feature center based on the medical record features of each first sample medical record.
[0067] In another possible implementation, the apparatus further includes:
[0068] The acquisition module is further configured to acquire, based on the first medical record classification model, the category to which each target word contained in the target medical record belongs;
[0069] A determination module is used to determine the category to which the target medical record belongs based on the category to which each target word belongs.
[0070] On the other hand, a computer device is provided, which includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the processing method of the medical record classification model as described in the above aspects.
[0071] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed in the processing method of the medical record classification model as described in the above aspects.
[0072] In another aspect, a computer program product or computer program is provided, comprising computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to implement the operations performed in the method for processing a medical record classification model as described in the above aspects.
[0073] The beneficial effects of the technical solutions provided in the embodiments of the present application include at least:
[0074] The method, apparatus, computer equipment, and storage medium provided in the embodiments of the present application provide a training method for a medical record classification model, and classify and manage medical records based on the trained medical record classification model. In addition, in the process of training the first medical record classification model, only the target number of second sample medical records from the sample medical records previously used to train the second medical record classification model is required. The first medical record classification model is trained based on the obtained second sample medical records and the first sample medical records belonging to the newly added category to improve the accuracy of the first medical record classification model. There is no need to train the first medical record classification model based on all the sample medical records previously used to train the second medical record classification model, thereby saving computational effort. Afterwards, the first sample medical record set is adjusted to balance the number of sample medical records belonging to different first categories in the obtained second sample medical record set. The first medical record classification model is trained again based on the second sample medical record set, thereby reducing the impact of the excessive difference in the number of sample medical records belonging to different first categories and further improving the accuracy of the first medical record classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0076] Figure 1 This is a schematic diagram of the structure of an implementation environment provided by an embodiment of the present application;
[0077] Figure 2 This is a flowchart of a method for processing a medical record classification model provided in an embodiment of the present application;
[0078] Figure 3 This is a flowchart of a method for processing a medical record classification model provided in an embodiment of the present application;
[0079] Figure 4 This is a flowchart of a first medical record classification model trained based on a first sample medical record set provided by an embodiment of the present application;
[0080] Figure 5 This is a flowchart of a method for processing a medical record classification model provided in an embodiment of the present application;
[0081] Figure 6 This is a schematic diagram showing how the accuracy of a medical record classification model provided in an embodiment of the present application changes with category;
[0082] Figure 7 This is a flow chart for classifying medical records provided in an embodiment of the present application;
[0083] Figure 8 It is a structural diagram of a processing device for a medical record classification model provided in an embodiment of the present application;
[0084] Figure 9 It is a structural diagram of a processing device for a medical record classification model provided in an embodiment of the present application;
[0085] Figure 10 This is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0086] Figure 11 This is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0087] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0088] As used herein, the terms "first," "second," "third," "fourth," and the like may be used to describe various concepts herein, but unless otherwise specified, these concepts are not limited by these terms. These terms are merely used to distinguish one concept from another. For example, a first similarity can be referred to as a second similarity, and similarly, a second similarity can be referred to as a first similarity without departing from the scope of this application.
[0089] As used herein, the terms "at least one," "a plurality," "each," and "any" include one, two, or more, "a plurality" includes two or more, "each" refers to each of the corresponding plurality, and "any" refers to any one of the plurality. For example, if the plurality of sample labels includes three sample labels, "each" refers to each of the three sample labels, and "any" refers to any one of the three sample labels, which can be the first sample label, the second sample label, or the third sample label.
[0090] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0091] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.
[0092] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0093] The solution provided in the embodiment of the present application is based on artificial intelligence machine learning technology, which can train a medical record classification model. The trained medical record classification model can then be used to classify medical records, thereby realizing classified management of medical records.
[0094] The processing method of the medical record classification model provided in the embodiment of the present application is executed by a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this.
[0095] In some embodiments, the computer program involved in the embodiments of the present application can be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.
[0096] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present application. Figure 1 , the implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a wireless or wired network.
[0097] Server 102 can train a first medical record classification model based on sample medical records and store the trained first medical record classification model, for example, storing the trained first medical record classification model in the server or other device for subsequent access. Terminal 101 is configured to send a medical record classification request to server 102. Server 102 is configured to receive the medical record classification request sent by terminal 101, classify the medical records included in the medical record classification request based on the first medical record classification model, and manage the medical records based on the categories to which they belong.
[0098] Figure 2 This is a flowchart of a method for processing a medical record classification model provided by an embodiment of the present application, which is executed by a computer device, such as Figure 2 As shown, the method includes:
[0099] 201. A computer device obtains a first medical record classification model, where the first medical record classification model is obtained by adjusting the second medical record classification model.
[0100] The first medical record classification model includes a first category, meaning that the first medical record classification model can be used to classify medical records into multiple first categories. The second medical record classification model includes a second category, meaning that the second medical record classification model can be used to classify medical records into multiple second categories. Furthermore, the first category includes the second category and the newly added category, and the second category is different from the newly added category, meaning that the first medical record classification model can classify more categories than the second medical record classification model.
[0101] The first category is used to indicate the category to which the medical record belongs. For example, the first category includes gastroenterology category, respiratory category, neurology category, etc. The second category includes gastroenterology category, respiratory category, etc., and the newly added category is neurology category. In an embodiment of the present application, the second medical record classification model is a currently trained model, which is used to classify medical records. The second medical record classification model includes a second category, and any medical record can be classified based on the second medical record classification model to determine the second category to which the medical record belongs. The first medical record classification model is the model to be trained currently, and the first medical record classification model is obtained by adjusting the second medical record classification model, that is, the first medical record classification model is obtained by adding the newly added category to the second medical record classification model on the basis of the second medical record classification model. Subsequently, any medical record can be classified based on the trained first medical record classification model to determine the first category to which the medical record belongs.
[0102] 202. The computer device forms a first sample medical record set with the first sample medical records belonging to the newly added category and the second sample medical records belonging to the second category. The second sample medical records are selected from the sample medical records for training the second medical record classification model, and the number of second sample medical records belonging to the same second category is the target number. Based on the first sample medical record set, the first medical record classification model is trained.
[0103] In an embodiment of the present application, the second medical record classification model is obtained based on training of sample medical records belonging to the second category. From the sample medical records used to train the second medical record classification model, a target number of second sample medical records belonging to each second category are selected, and the selected second sample medical records and the first sample medical records belonging to the newly added category constitute a first sample medical record set, so that the first sample medical record set includes sample medical records belonging to each first category, so that the first medical record classification model can be subsequently trained based on the first sample medical record set.
[0104] The first medical record classification model is trained using the first sample medical record set so that the trained first medical record classification model can enhance its ability to distinguish different first categories, so that the trained first medical record classification model can determine the first category to which the medical record belongs, thereby ensuring the accuracy of the trained first medical record classification model.
[0105] 203. The computer device combines the target number of first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set to form a second sample medical record set.
[0106] In an embodiment of the present application, since there may be differences between the numbers of sample medical records belonging to different first categories in the first sample medical record set, the accuracy of the first medical record classification model trained based on the first sample medical record set may be poor. For example, the number of first sample medical records belonging to the newly added category in the first sample medical record set is greater than the target number. Since the number of sample medical records in the first sample medical record set is unbalanced, it may lead to low accuracy of the first medical record classification model. Therefore, the first sample medical record set is adjusted so that the number of sample medical records belonging to different first categories in the obtained second sample medical record set is equal, that is, the number of sample medical records belonging to different first categories in the second sample medical record set is balanced.
[0107] 204. The computer device trains the first medical record classification model again based on the second sample medical record set.
[0108] Since the number of sample medical records belonging to different first categories in the second sample medical record set is balanced, the first medical record classification model is trained again based on the second sample medical record set to weaken the impact caused by the imbalance in the number of sample medical records, so that when the first medical record classification model classifies a medical record, the probability of assigning the medical record to any first category is balanced, thereby improving the accuracy of the first medical record classification model.
[0109] The method provided in the embodiment of the present application provides a training method for a medical record classification model, and classifies and manages medical records based on the trained medical record classification model. In addition, in the process of training the first medical record classification model, only the target number of second sample medical records from the sample medical records used to train the second medical record classification model is obtained. The first medical record classification model is trained based on the obtained second sample medical records and the first sample medical records belonging to the newly added category to improve the accuracy of the first medical record classification model. There is no need to train the first medical record classification model based on all the sample medical records used to train the second medical record classification model, thereby saving computational complexity. Afterwards, the first sample medical record set is adjusted to balance the number of sample medical records belonging to different first categories in the obtained second sample medical record set. The first medical record classification model is trained again based on the second sample medical record set, thereby reducing the impact of the excessive difference in the number of sample medical records belonging to different first categories and further improving the accuracy of the first medical record classification model.
[0110] exist Figure 2 On the basis of the illustrated embodiment, in the two training stages of the first medical record classification model, the similarity differences of the words are obtained respectively, and training is performed in combination with the differences between the predicted information and the annotation information of the words, as well as the difference information that can represent the similarity differences. The training process is detailed in the following embodiment.
[0111] Figure 3 This is a flowchart of a method for processing a medical record classification model provided by an embodiment of the present application, which is executed by a computer device, such as Figure 3 As shown, the method includes:
[0112] 301. The computer device obtains a first medical record classification model, which is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category.
[0113] In an embodiment of the present application, the second medical record classification model is a trained model, and the second medical record classification model includes the second category. The first medical record classification model is the model currently to be trained, and the second medical record classification model includes the first category. The first category includes the second category and the newly added category. That is, the trained second medical record classification model is adjusted, and the newly added category is added to the second medical record classification model to obtain the first medical record classification model.
[0114] In one possible implementation, medical records are used to describe a patient's symptoms, disease progression, treatment process, etc. In this embodiment of the present application, medical records belonging to different first categories describe different symptoms, diseases, and treatment processes. For example, multiple first categories include: gastroenterology category, respiratory category, neurology category, etc. The symptoms described in medical records belonging to the gastroenterology category are different from those described in medical records belonging to the respiratory category.
[0115] In one possible implementation, the first medical record classification model includes weight features for each first category, and the second medical record classification model includes weight features for each second category.
[0116] Since the first medical record classification model is obtained by adjusting the second medical record classification model, the weight feature corresponding to any second category in the first medical record classification model is the same as the weight feature corresponding to the second category in the second medical record classification model. That is, when adjusting the second medical record classification model, only the newly added category is added to the second medical record classification model and the weight feature is determined for the newly added category. The weight feature of each second category in the second medical record classification model remains unchanged. After the adjustment, the first medical record classification model is obtained. Afterwards, the first medical record classification model is trained based on the sample medical record set to adjust the weight feature of each first category or other model parameters in the first medical record classification model.
[0117] 302. The computer device forms a first sample medical record set with first sample medical records belonging to the newly added category and second sample medical records belonging to the second category. The second sample medical records are selected from the sample medical records for training the second medical record classification model, and the number of second sample medical records belonging to the same second category is the target number.
[0118] The target number is an arbitrary number. In the first sample medical record set, the number of first sample medical records belonging to the newly added category may be different from the target number.
[0119] In an embodiment of the present application, each sample medical record in the first sample medical record set contains a word, and the first sample medical record set includes annotation information for each word contained in each sample medical record, and the annotation information includes multiple sample labels, and each sample label indicates whether the corresponding word belongs to a first category. For example, for any sample medical record, the sample medical record includes 5 words, and the first medical record classification model includes 4 first categories, that is, the annotation information of each word includes 4 sample labels, the first sample label indicates whether the corresponding word belongs to the first category, the second sample label indicates whether the corresponding word belongs to the second category, and so on. Based on the multiple sample labels corresponding to the annotation information of the word, it can be determined whether the word belongs to each first category. Optionally, the annotation information of each word is obtained by manual annotation.
[0120] In one possible implementation, step 302 includes: obtaining a first sample medical record belonging to a newly added category, obtaining a second sample medical record belonging to a second category stored in a target storage area, and forming a first sample medical record set with the obtained first sample medical record and the second sample medical record.
[0121] The target storage area is used to store sample medical records for training a medical record classification model. For example, the target storage area is a memory area for the medical record classification model. The target storage area stores sample medical records previously used in training the medical record classification model, and the number of sample medical records belonging to each category stored in the target storage area is the target number. In an embodiment of the present application, after training a new medical record classification model, the sample medical records used to train the new medical record classification model are stored in the target storage area so that other medical record classification models can be trained based on the sample medical records stored in the target storage area.
[0122] For example, after training the first medical record classification model, a target number of sample medical records belonging to each category are selected from the sample medical records used to train the first medical record classification model and stored in the target storage area; thereafter, if a new category needs to be added to the first medical record classification model, the first medical record classification model is adjusted to obtain a second medical record classification model, and the second medical record classification model is trained based on the sample medical records stored in the target storage area and the sample medical records belonging to the newly added category. Thereafter, the target number of sample medical records belonging to the newly added category are stored in the target storage area. With this category, after adding a category to the medical record classification model, the medical record classification model with the added category is trained based on the sample medical records stored in the target storage area and the sample medical records belonging to the newly added category. Thereafter, the target number of sample medical records belonging to the newly added category are stored in the target storage area.
[0123] 303. The computer device obtains, for each first word in the third sample medical record, a first similarity between the first word and each first category based on the first medical record classification model.
[0124] Among them, the third sample medical record is any sample medical record in the first sample medical record set, and the third sample medical record includes at least one first word. The first similarity between the first word and any first category is used to indicate the degree of similarity between the first word and the first category, which can reflect the possibility that the first word belongs to the first category.
[0125] For any word in the third sample medical record, the word is input into the first medical record classification model to obtain the first similarity between the first word and each first category. According to the above method, each first word in the third sample medical record is input into the first medical record classification model to obtain the first similarity between each first word in the third sample medical record and each first category.
[0126] In one possible implementation, the first medical record classification model includes weight features for each first category. Step 303 includes: for each first word, based on the first medical record classification model, performing feature extraction on the first word to obtain the word feature of the first word, and obtaining the first similarity between the word feature and the weight feature of each first category.
[0127] When obtaining the first similarity between the word feature and each weight feature of the first category, cosine similarity or other similarities can be used.
[0128] Optionally, the process of obtaining the word feature of the first word includes: performing feature extraction on the first word based on a feature extraction sub-model in the first medical record classification model to obtain the word feature of the first word. The feature extraction sub-model is used to extract the word feature of the word, for example, the feature extraction sub-model is BERT (Bidirectional Encoder Representations from Transformers, a language representation model) or other network models.
[0129] Optionally, the process of obtaining the first similarity between the word feature and the weight feature of each first category based on the first medical record classification model includes: based on the first medical record classification model, regularizing the word feature of the first word and the weight feature of each first category respectively, obtaining the processed word feature and the processed weight feature of each first category, and obtaining the first similarity between the processed word feature and the processed weight feature of each first category.
[0130] When regularizing the weight features or word features, a bi-norm or other methods are used. By regularizing the word features and weight features, overfitting of the subsequent training of the first medical record classification model is avoided, thereby ensuring the accuracy of the subsequent training of the first medical record classification model.
[0131] Optionally, based on the cosine regularized linear layer in the first medical record classification model, regularization processing is performed on the word features of the first word and the weight features of each first category to obtain the processed word features and the processed weight features of each first category.
[0132] For example, the word feature is a word feature vector, the vector length of the word feature vector is obtained, the ratio of each feature value in the word feature vector to the vector length is determined, and the obtained multiple ratios constitute the word feature vector after the word feature vector is regularized.
[0133] Optionally, the process of regularizing the word feature vector satisfies the following relationship:
[0134]
[0135] Among them, v represents the word feature vector of any first word, Represents the word feature vector after regularization.
[0136] 304. The computer device normalizes each first similarity based on the first medical record classification model and the first similarity corresponding to each first category to obtain multiple first prediction labels for the first word, and uses the multiple first prediction labels to form first prediction information for the first word.
[0137] Among them, each first prediction label indicates the possibility that the corresponding first word belongs to a first category. After determining the first similarity between the first word and each first category, the first prediction information of the first word can be obtained by normalizing the first similarity corresponding to each first category. According to the above steps 303-304, the first prediction information of each first word in the third sample medical record can be obtained. By obtaining the first similarity between the first word and each first category and normalizing the first similarity between the first word and each first category, it is ensured that when the first medical record classification model is subsequently trained based on the first prediction label, the first medical record classification model region converges, thereby ensuring the accuracy of the trained first medical record classification model.
[0138] In one possible implementation, the first prediction label is a first prediction probability. Step 304 includes: based on the first medical record classification model, the first similarity corresponding to each first category and the adjustment parameter, normalizing each first similarity to obtain multiple first prediction probabilities of the first word.
[0139] The first medical record classification model includes the adjustment parameter, which is used to adjust the distribution of the plurality of first predicted probabilities. When obtaining the first predicted probability of the first word, the adjustment parameter is introduced to adjust the distribution of the first predicted probabilities to ensure the accuracy of the obtained first predicted probabilities.
[0140] In one possible implementation, the first predicted label is a first predicted probability, and step 304 satisfies the following relationship:
[0141]
[0142] Where x represents any first word in the third sample medical record, i represents the serial number of the first category included in the first medical record classification model, and p i(x) represents the first predicted label corresponding to the i-th first category, that is, the probability that the first word x belongs to the i-th first category, η represents the adjustment parameter, θ i represents the weighted feature of the first category of the i-th type; f(x) represents the word feature of the first word x; represents the weight feature of the first category after regularization processing of the i-th type, represents the word feature of the first word x after regularization processing, represents the first similarity between the first word x and the i-th first category; exp(·) identifies the power function of e; j is the sequence number of the first category included in the first medical record classification model.
[0143] It should be noted that the embodiment of the present application obtains the first prediction information of the first word based on the first similarity between the first word and each first category. In another embodiment, there is no need to execute steps 303-304, and other methods can be adopted to obtain the first prediction information of each first word in the third sample medical record based on the first medical record classification model.
[0144] 305. The computer device obtains, for each first word, a second similarity between the first word and each second category based on the second medical record classification model.
[0145] The second similarity between the first word and any second category is used to indicate the degree of similarity between the first word and the second category, and can reflect the possibility that the first word belongs to the second category. This step is similar to step 303 above and will not be described in detail here.
[0146] 306. The computer device trains the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the first difference information of each first word, where the first difference information includes multiple first similarity differences, and each first similarity difference includes the difference between the first similarity and the second similarity of the corresponding first word in a second category.
[0147] In an embodiment of the present application, for each first word in the third sample medical record, the first word corresponds to first prediction information and annotation information. The first prediction information is predicted based on the first medical record classification model, and the annotation information is obtained by annotation. The difference between the first prediction information of the first word and the annotation information of the first word can reflect the accuracy of the first medical record classification model.
[0148] For any first word in the third sample medical record, the first medical record classification model can determine the similarity between the first word and each first category. The first category includes the second category and the newly added category. Therefore, from the similarity between the first word and each first category, the first similarity between the first word and each second category can be determined. The second medical record classification model can determine the second similarity between the first word and each second category. In the same any second category, the difference between the first similarity and the second similarity corresponding to the first word is a first similarity difference. Based on the first similarity and second similarity of the first word in multiple second categories, multiple first similarity differences can be determined to constitute the first difference information of the first word. In this way, the first difference information of each first word in the third sample medical record can be obtained.
[0149] Since each first similarity difference can reflect the difference between the first medical record classification model and the second medical record classification model, the first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, as well as the first difference information of each first word. This not only takes into account the accuracy of the current first medical record classification model, but also ensures that the first medical record classification model retains the classification ability of the second medical record classification model, that is, retains the classification ability of the first medical record classification model in the second category, thereby improving the accuracy of the first medical record classification model.
[0150] In one possible implementation, step 306 includes training the first medical record classification model based on the difference between the sample label and the first predicted label of each first word in the same first category, and the first difference information of each first word.
[0151] In an embodiment of the present application, for any first word, the first prediction information of the first word includes multiple first prediction labels, and the annotation information of the first word includes multiple annotation labels. When determining the difference between the first prediction information and the annotation information of the first word, the difference between the sample label and the prediction label on the same first category in the first prediction information and the annotation information of the first word is determined, that is, multiple differences can be determined, and the first word corresponds to multiple differences. The multiple differences corresponding to each word and the first difference information of each first word are used to train the first medical record classification model to improve the accuracy of the first medical record classification model.
[0152] In one possible implementation, step 306 includes: determining a first loss value based on the difference between the first prediction information and the annotation information of each first word, determining a second loss value based on the first difference information of each first word, and training the first medical record classification model based on the sum of the first loss value and the second loss value.
[0153] In an embodiment of the present application, the first loss value is equivalent to the classification cross entropy loss, and the second loss value is equivalent to the model distillation loss. The first medical record classification model is trained based on the model distillation loss, so that the first medical record classification model can retain the classification ability of the second medical record classification model for the second category, that is, retain the knowledge of the original second medical record classification model, that is, achieve the effect of distilling the second medical record classification model.
[0154] Optionally, the first loss value, the second loss value, and the sum of the first loss value and the second loss value satisfy the following relationship:
[0155]
[0156] L1=L ce (x1)+L dis (x1)
[0157] Where x1 represents any first word in the third sample medical record; L ce (x1) represents the first loss value; i represents the serial number of the first category included in the first medical record classification model; y i represents the sample label of the first word x1 in the i-th category, p i represents the predicted label of the first word x1 in the i-th category, J represents the total number of first categories included in the first medical record classification model; L dis (x1) represents the second loss value; represents the weight feature after regularization of the weight feature of the jth second category in the first medical record classification model, f(x1) represents the word feature after regularization of the word feature of the first word x1 obtained based on the first medical record classification model, represents the first similarity between the first word x1 and the jth second category; represents the weight feature after regularization of the weight feature of the jth second category in the second medical record classification model, represents the word features of the first word x1 obtained based on the second medical record classification model after regularization processing, represents the second similarity between the first word x1 and the jth second category; J0 represents the total number of second categories included in the second medical record classification model; L1 represents the first loss value L ce (x1) and the second loss value L dis The sum of (x1).
[0158] like Figure 4As shown, during the training of the first medical record classification model based on the first sample medical record set, the second sample medical record in the first sample medical record set is retrieved from the target storage area. The retrieved second sample medical record is the memorized sample medical record. For the third sample medical record in the first sample medical record set, based on the second medical record classification model 401, feature extraction is performed on the first words in the third sample medical record to obtain the word features of each first word. Based on the cosine regularized linear layer in the second medical record classification model 401, the regularized word features and the weight features of each second category are obtained, and the second similarity corresponding to each second category is obtained. Based on the first medical record classification model 402, feature extraction is performed on the first words in the third sample medical record to obtain word features of each first word. Based on the cosine regularized linear layer in the first medical record classification model 402, the regularized word features and the weight features of each second category are obtained, the first similarity corresponding to each second category is obtained, and the first prediction information of each first word is obtained. The first prediction information includes multiple first prediction probabilities, and each first prediction probability indicates the possibility that the corresponding first word belongs to a first category; based on the first similarity and the second similarity of each first word in the same second category, the model distillation loss is determined, and based on the difference between the first prediction information and the annotation information of each first word, the classification cross entropy loss is determined. The first medical record classification model 402 is trained based on the model distillation loss and the classification cross entropy loss.
[0159] It should be noted that the embodiment of the present application is only illustrated by taking a third sample medical record in the first sample medical record set as an example. In another embodiment, the process of training the first medical record classification model based on a third sample medical record is one iteration, and the first medical record classification model is trained multiple times based on multiple third sample medical records in the first sample medical record set.
[0160] It should be noted that the embodiment of the present application uses the second medical record classification model to train the first medical record classification model. In another embodiment, there is no need to execute steps 305-306, and other methods can be adopted to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word.
[0161] It should be noted that the embodiment of the present application first obtains the first prediction information of each first word, and uses the second medical record classification model to obtain the first difference information of each first word, and trains the first medical record classification model based on the difference between the first prediction information and the annotation information of the first word, as well as the first difference information of each first word. In another embodiment, there is no need to execute steps 303-306, and other methods can be adopted to train the first medical record classification model based on the first sample medical record set.
[0162] 307. The computer device combines the target number of first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set to form a second sample medical record set.
[0163] In an embodiment of the present application, the second sample medical record set includes annotation information of each word contained in each sample medical record, and the annotation information includes multiple sample labels, each sample label indicating whether the corresponding word belongs to a first category.
[0164] In one possible implementation, after step 306, the method further includes: storing a target number of first sample medical records in the first sample medical record set in a target storage area, and forming a second sample medical record set from the sample medical records stored in the target storage area.
[0165] The number of sample medical records belonging to different categories stored in the target storage area is the target number. In this embodiment of the present application, before the first sample medical record is stored in the target storage area, the target storage area includes the second sample medical records belonging to each second category, and the target number of second sample medical records belonging to the same second category is the target number. After the target number of first sample medical records are stored in the target storage area, the target storage area includes the sample medical records of each first category, and the target number of sample medical records belonging to the same first category is the target number.
[0166] In one possible implementation, step 307 includes the following two methods:
[0167] The first method: the computer device randomly selects a target number of first sample medical records from the first sample medical record set, and the selected first sample medical records and the second sample medical records in the first sample medical record set constitute a second sample medical record set.
[0168] The second method includes the following steps 1-3:
[0169] Step 1: The computer device obtains the medical record features of each first sample medical record and the feature center of the first sample medical record in the first sample medical record set.
[0170] The medical record features are used to characterize the corresponding sample medical records. The feature center is the feature center of the medical record features of multiple first sample medical records in the first sample medical record set. The feature center can represent the common features of the sample medical records belonging to the newly added category. For example, if the medical record features of the first sample medical record are represented in the form of a feature vector, the feature center is also represented in the form of a feature vector.
[0171] In one possible implementation, step 1 includes: for each first sample medical record, based on the first medical record classification model, obtaining the word features of each third word in the first sample medical record, obtaining the medical record features of the first sample medical record based on the word features of each third word in the first sample medical record, and obtaining the feature center based on the medical record features of multiple first sample medical records in the first sample medical record set.
[0172] The word feature of the third word is used to characterize the third word. The method of obtaining the medical record feature of the first sample medical record includes: determining the sum of the word features of each third word in the first sample medical record as the medical record feature of the first sample medical record; or determining the average word feature of the word features of the third words in the first sample medical record as the medical record feature of the first sample medical record. The method of obtaining the feature center based on the medical record feature of each first sample medical record includes: determining the sum of the medical record features of each first sample medical record in the first sample medical record set as the feature center; or determining the average medical record feature of the medical record features of multiple first sample medical records in the first sample medical record set as the medical record feature of the first sample medical record.
[0173] Step 2: The computer device selects a target number of first sample medical records from the first sample medical record set based on the distance between the medical record feature and the feature center of each first sample medical record, wherein the distance corresponding to the selected first sample medical records is smaller than the distance corresponding to the remaining first sample medical records in the first sample medical record set.
[0174] The distance between the medical record feature and the feature center of any first sample medical record can reflect the degree to which the first sample medical record conforms to the newly added category. A larger distance indicates that the corresponding first sample medical record conforms less to the newly added category, and a smaller distance indicates that the corresponding first sample medical record conforms more to the newly added category. After determining the feature centers corresponding to all first sample medical records in the first sample medical record set, the sample medical record that best represents the newly added category can be selected from the first sample medical record set based on the distance between the medical record feature and the feature center of each first sample medical record.
[0175] Step 3: The computer device combines the selected first sample medical records and the second sample medical records in the first sample medical record set to form a second sample medical record set.
[0176] By selecting the target number of first sample medical records that best represent the newly added categories from the first sample medical record set, the selected first sample medical records and the second sample medical records in the first sample medical record set are combined to form a second sample medical record set, so that each sample medical record in the second sample medical record set can represent the corresponding category, thereby ensuring the accuracy of the sample medical records in the second sample medical record set.
[0177] 308. The computer device obtains, for each second word in the fourth sample medical record, a third similarity between the second word and each first category based on the first medical record classification model.
[0178] The fourth sample medical record is any sample medical record in the second sample medical record set, and the fourth sample medical record includes at least one second word. The first similarity between the second word and any one of the first categories indicates the degree of similarity between the second word and the first category, thereby indicating the likelihood that the second word belongs to the first category. This step is similar to step 303 above and is not further described here.
[0179] 309. The computer device normalizes each third similarity based on the first medical record classification model and the third similarity corresponding to each first category to obtain multiple second prediction labels for the second word, and uses the multiple second prediction labels to form second prediction information for the second word.
[0180] Each second predicted label indicates the possibility that the corresponding second word belongs to a first category. This step is similar to the above step 304 and will not be repeated here.
[0181] It should be noted that the embodiment of the present application is explained by obtaining the second prediction information of the second word based on the third similarity between the second word and each first category. In another embodiment, there is no need to execute steps 308-309, and other methods can be adopted to obtain the second prediction information of each second word in the fourth sample medical record based on the first medical record classification model.
[0182] 310. The computer device creates a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set.
[0183] In an embodiment of the present application, the third medical record classification model created is the same as the first medical record classification model trained based on the first sample medical record set. Afterwards, the first medical record classification model is trained again using the third medical record model, and during the training of the first medical record classification model, the third medical record classification model remains unchanged.
[0184] 311. The computer device obtains, for each second word, a fourth similarity between the second word and each first category based on the third medical record classification model.
[0185] This step is similar to the above step 303 and will not be described again here.
[0186] 312. The computer device trains the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the second difference information of each second word, where the second difference information includes multiple second similarity differences, and each second similarity difference includes the difference between the third similarity and the fourth similarity of the corresponding second word in a first category.
[0187] In one possible implementation, step 312 includes: determining a third loss value based on the difference between the second prediction information and the annotation information of each second word, determining a fourth loss value based on the second difference information of each second word, and training the first medical record classification model based on the sum of the third loss value and the fourth loss value.
[0188] Optionally, the third loss value, the fourth loss value, and the sum of the third loss value and the fourth loss value satisfy the following relationship:
[0189]
[0190] L2=L ce (x2)+L dis (x2)
[0191] Where x2 represents any second word in the fourth sample medical record; L ce (x2) represents the third loss value; i represents the serial number of the first category included in the first medical record classification model; y i represents the sample label of the second word x2 in the i-th category, p i represents the predicted label of the second word x2 in the i-th category, J represents the total number of the first category included in the first medical record classification model; L dis (x2) represents the fourth loss value; represents the weight feature after regularization of the weight feature of the first category of the jth type in the first medical record classification model. represents the word features of the second word x2 obtained based on the first medical record classification model after regularization processing, represents the third similarity between the second word x2 and the jth first category; represents the weight feature after regularization of the weight feature of the first category of the jth type in the third medical record classification model. represents the word features after regularization processing of the word features of the second word x2 obtained based on the third medical record classification model; represents the fourth similarity between the second word x2 and the jth first category; L2 represents the third loss value L ce (x2) and the fourth loss value L dis The sum of (x2).
[0192] In the embodiment of the present application, the process of training the first medical record classification model based on the first medical record set is the first training phase, and the process of retraining the first medical record classification model based on the second sample medical record set is the second training phase. To reduce the impact of the second training phase on the first medical record classification model trained in the first training phase, a smaller learning rate is used when determining the third loss value and the fourth loss value during the second training phase to achieve the effect of fine-tuning the first pathology classification model, thereby ensuring the accuracy of the first medical record classification model.
[0193] It should be noted that the embodiment of the present application is only illustrated by taking a fourth sample medical record in the second sample medical record set as an example. In another embodiment, the process of training the first medical record classification model based on a fourth sample medical record is one iteration, and the first medical record classification model is trained multiple times based on multiple fourth sample medical records in the second sample medical record set.
[0194] It should be noted that the embodiment of the present application uses the third medical record classification model to train the first medical record classification model. In another embodiment, there is no need to execute steps 310-312, and other methods can be adopted to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word.
[0195] It should be noted that the embodiment of the present application first obtains the second prediction information of each second word, and uses the third medical record classification model to obtain the second difference information of each second word, and trains the first medical record classification model based on the difference between the second prediction information and the annotation information of the second word, as well as the second difference information of each second word. In another embodiment, there is no need to execute steps 308-312, and other methods can be adopted to train the first medical record classification model again based on the second sample medical record set.
[0196] The method provided in the embodiment of the present application provides a training method for a medical record classification model, and classifies and manages medical records based on the trained medical record classification model. In addition, in the process of training the first medical record classification model, only the target number of second sample medical records from the sample medical records used to train the second medical record classification model is obtained. The first medical record classification model is trained based on the obtained second sample medical records and the first sample medical records belonging to the newly added category to improve the accuracy of the first medical record classification model. There is no need to train the first medical record classification model based on all the sample medical records used to train the second medical record classification model, thereby saving computational complexity. Afterwards, the first sample medical record set is adjusted to balance the number of sample medical records belonging to different first categories in the obtained second sample medical record set. The first medical record classification model is trained again based on the second sample medical record set, thereby reducing the impact of the excessive difference in the number of sample medical records belonging to different first categories and further improving the accuracy of the first medical record classification model.
[0197] Moreover, in the first training phase of the first medical record classification model, it is not necessary to use all the sample medical records used to train the second medical record classification model to train the first medical record classification model, thereby reducing the amount of computation in the process of training the first medical record classification model.
[0198] Moreover, in the two training stages of the first medical record classification model, the second medical record classification model or the third medical record classification model is used to determine the similarity difference, and the first medical record classification model is trained based on the determined similarity difference. That is, the model distillation loss is taken into account during the training of the first medical record classification model, so that the trained first medical record classification model retains the classification ability already possessed by the previous model, and on this basis, the accuracy of the first medical record classification model is further improved.
[0199] Furthermore, in an embodiment of the present application, after training a new medical record classification model, the sample medical records used to train the new medical record classification model are stored in the target storage area, so that other medical record classification models can be subsequently trained based on the sample medical records stored in the target storage area. Moreover, the number of sample medical records belonging to each category stored in the target storage area is the target number, and there is no need to store all the sample medical records of the previously trained medical record classification model, thereby saving storage resources.
[0200] exist Figure 2 On the basis of the shown embodiment, in the two training stages of the first medical record classification model, the differences between the word features of the words are obtained respectively, and training is performed in combination with the differences between the prediction information and the annotation information of the words, as well as the differences between the word features of the words. The training process is detailed in the following embodiment.
[0201] Figure 5This is a flowchart of a method for processing a medical record classification model provided by an embodiment of the present application, which is executed by a computer device, such as Figure 5 As shown, the method includes:
[0202] 501. The computer device obtains a first medical record classification model, which is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category.
[0203] This step is similar to the above step 301 and will not be described again here.
[0204] 502. The computer device forms a first sample medical record set with first sample medical records belonging to the newly added category and second sample medical records belonging to the second category. The second sample medical records are selected from the sample medical records for training the second medical record classification model, and the number of second sample medical records belonging to the same second category is the target number.
[0205] This step is similar to the above step 302 and will not be described again here.
[0206] 503. The computer device obtains, for each first word in the third sample medical record based on the first medical record classification model, a first word feature of the first word, and obtains a first similarity between the first word feature and the weight feature of each first category.
[0207] The first word feature is used to represent the corresponding first word. The first medical record classification model includes a weight feature of the first category. The first word feature and the first weight feature can be represented in any form, for example, the weight feature is represented in the form of a vector. The first similarity between the first word feature and any weight feature of the first category is used to represent the degree of similarity between the first word and the first category, which can reflect the possibility that the first word belongs to the first category.
[0208] In one possible implementation, the process of obtaining the first word feature includes: extracting features of the first word based on a feature extraction sub-model in the first medical record classification model to obtain the word feature of the first word.
[0209] In one possible implementation, the process of obtaining the first similarity between the word feature and the weight feature of each first category based on the first medical record classification model includes: based on the first medical record classification model, regularizing the word feature of the first word and the weight feature of each first category, respectively, to obtain the processed word feature and the processed weight feature of each first category, and obtaining the first similarity between the processed word feature and the processed weight feature of each first category.
[0210] It should be noted that the embodiment of the present application obtains the first similarity based on the word features of the first word and the weight features of the first category. In another embodiment, there is no need to execute step 503, and other methods can be adopted to obtain the first similarity between the first word and each first category based on the first medical record classification model.
[0211] 504. The computer device normalizes each first similarity based on the first medical record classification model and the first similarity corresponding to each first category to obtain multiple first prediction labels for the first word, and constitutes the first prediction information of the first word with the multiple first prediction labels.
[0212] The first prediction information includes multiple first prediction labels, each of which indicates the likelihood that the corresponding first word belongs to a first category, and the third sample medical record is any sample medical record in the first sample medical record set. This step is similar to step 304 above and will not be repeated here.
[0213] It should be noted that the embodiment of the present application obtains the first prediction information of the first word based on the first similarity between the first word and each first category. In another embodiment, there is no need to execute steps 503-504, and other methods can be adopted to obtain the first prediction information of each first word in the third sample medical record based on the first medical record classification model.
[0214] 505. The computer device obtains a second word feature of each first word based on the second medical record classification model.
[0215] The second word feature is used to characterize the first word. This step is similar to the above step 503 and will not be described in detail here.
[0216] 506. The computer device trains the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word.
[0217] Among them, for any first word, the difference between the first word feature and the second word feature of the first word can reflect the difference between the first medical record classification model and the second medical record classification model.
[0218] Based on the difference between the first prediction information and the annotation information of each first word, as well as the difference between the first word feature and the second word feature of each first word, the first medical record classification model is trained. This not only takes into account the accuracy of the current first medical record classification model, but also ensures that the first medical record classification model retains the classification ability of the second medical record classification model, that is, the classification ability of the first medical record classification model in the second category is retained, thereby improving the accuracy of the first medical record classification model.
[0219] It should be noted that the embodiment of the present application uses the second medical record classification model to train the first medical record classification model. In another embodiment, there is no need to execute steps 505-506, and other methods can be adopted to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word.
[0220] It should be noted that the embodiment of the present application is based on the first medical record classification model. The first word feature of each first word is first obtained, and then the first prediction information of each first word is obtained. The second medical record classification model is used to obtain the second word feature of each first word. Based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word, the first medical record classification model is trained. In another embodiment, there is no need to execute steps 503-506, and other methods can be adopted to train the first medical record classification model based on the first sample medical record set.
[0221] 507. The computer device combines the target number of first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set to form a second sample medical record set.
[0222] This step is similar to the above step 307 and will not be described again here.
[0223] 508. The computer device obtains, for each second word in the fourth sample medical record based on the first medical record classification model, a third word feature of the second word, and obtains a third similarity between the third word feature and the weight feature of each first category.
[0224] The fourth sample medical record is any sample medical record in the second sample medical record set, and the fourth sample medical record includes at least one second word. The first similarity between the second word and any one of the first categories is used to indicate the degree of similarity between the second word and the first category, thereby indicating the likelihood that the second word belongs to the first category. This step is similar to step 503 above and is not further described here.
[0225] It should be noted that the embodiment of the present application first obtains the third word feature of the second word and then obtains the third similarity. In another embodiment, there is no need to execute step 508, and other methods can be adopted to obtain the third similarity between the second word and each first category based on the first medical record classification model.
[0226] 509. The computer device normalizes each third similarity based on the first medical record classification model and the third similarity corresponding to each first category to obtain multiple second prediction labels for the second word, and uses the multiple second prediction labels to form second prediction information for the second word.
[0227] Each second predicted label indicates the possibility that the corresponding second word belongs to a first category. This step is similar to the above step 304 and will not be repeated here.
[0228] It should be noted that, in the embodiment of the present application, the second prediction information of the second word is obtained based on the third similarity between the second word and each first category. In another embodiment, there is no need to execute steps 508-509, and other methods can be adopted to obtain the second prediction information of each second word in the fourth sample medical record based on the first medical record classification model.
[0229] 510. The computer device creates a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set.
[0230] This step is similar to the above step 310 and will not be described again here.
[0231] 511. The computer device obtains a fourth word feature of each second word based on the third medical record classification model.
[0232] This step is similar to the above step 503 and will not be described again here.
[0233] 512. The computer device trains the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the difference between the third word feature and the fourth word feature of each second word.
[0234] This step is similar to the above step 506 and will not be described again here.
[0235] It should be noted that the embodiment of the present application uses the third medical record classification model to train the first medical record classification model. In another embodiment, there is no need to execute steps 510-512, and other methods can be adopted to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word.
[0236] It should be noted that the embodiment of the present application is based on the first medical record classification model. The third word feature of each second word is first obtained, and then the second prediction information of each second word is obtained. The third medical record classification model is used to obtain the fourth word feature of each second word. Based on the difference between the second prediction information and the annotation information of each second word, and the difference between the third word feature and the fourth word feature of each second word, the first medical record classification model is trained. In another embodiment, there is no need to execute steps 508-512, and other methods can be adopted to train the first medical record classification model again based on the second sample medical record set.
[0237] It should be noted that this application is only Figure 3 or Figure 5 The embodiment shown is used as an example to illustrate. Figure 3 or Figure 5 In the embodiment shown, the process of training the first medical record classification model is divided into two training stages. The first training stage is to train the first medical record classification model based on the first sample medical record set, and the second training stage is to train the first medical record classification model based on the second sample medical record set. In each training stage, Figure 3 and Figure 5 The training processes shown can be combined arbitrarily. For example, in the first training stage, the first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, the first difference information of each first word, and the difference between the first word feature and the second word feature of each first word; in the second training stage, the first medical record classification model is trained based on the difference between the second prediction information and the annotation information of each second word, the second difference information of each second word, and the difference between the third word feature and the fourth word feature of each second word.
[0238] The embodiment of the present application provides a method for incrementally training a medical record classification model. After each training of the medical record classification model, a target number of sample medical records belonging to different categories from the sample medical records used to train the medical record classification model are stored, thereby saving storage resources. The next time sample medical records of a newly added category are obtained, an incremental medical record classification model can be trained in combination with the stored sample medical records. In this way, it can be applied to the scenario of incremental learning, realizing a training method for a medical record classification model based on continuous learning. When a new category is added to the medical record classification model, an incremental medical record classification model can be trained based on the previously existing medical record classification model in accordance with the above method.
[0239] The method provided in the embodiment of the present application provides a training method for a medical record classification model, and classifies and manages medical records based on the trained medical record classification model. In addition, in the process of training the first medical record classification model, only the target number of second sample medical records from the sample medical records used to train the second medical record classification model is obtained. The first medical record classification model is trained based on the obtained second sample medical records and the first sample medical records belonging to the newly added category to improve the accuracy of the first medical record classification model. There is no need to train the first medical record classification model based on all the sample medical records used to train the second medical record classification model, thereby saving computational complexity. Afterwards, the first sample medical record set is adjusted to balance the number of sample medical records belonging to different first categories in the obtained second sample medical record set. The first medical record classification model is trained again based on the second sample medical record set, thereby reducing the impact of the excessive difference in the number of sample medical records belonging to different first categories and further improving the accuracy of the first medical record classification model.
[0240] Moreover, in the two training stages of the first medical record classification model, the second medical record classification model or the third medical record classification model is used to obtain the differences between the word features of the same word, and the first medical record classification model is trained based on the differences between the word features of the words. That is, the model distillation loss is taken into account in the process of training the first medical record classification model, so that the trained first medical record classification model retains the classification ability already possessed by the previous model, and on this basis, the accuracy of the first medical record classification model is further improved.
[0241] Based on the above Figure 2 、 Figure 3 and Figure 5 In the embodiment shown, in the medical scenario, as time goes by, the types of medical records will gradually increase. Figure 2 、 Figure 3 and Figure 5 The embodiment shown can realize the training of the medical record classification model in the medical scenario, and as the types of medical records increase, the medical record classification model can be continuously trained to ensure that the trained medical record classification model always matches the current medical record type. In the method provided in the embodiment of the present application, the medical record classification model is trained by combining memory replay and knowledge distillation. Figure 6 As shown, curve 601 is the accuracy of the medical record classification model trained based on the method provided in the embodiment of the present application, and curve 602 is the accuracy of the medical record classification model trained based on the method provided by the related technology. Compared with the model training method provided by the related technology, the method provided in the embodiment of the present application can weaken the impact of the imbalance in the number of samples of different categories and can improve the accuracy of the obtained medical record classification model.
[0242] In the above Figure 2 、 Figure 3 and Figure 5 Based on the embodiment shown, the embodiment of the present application also provides a process for classifying medical records, such as Figure 7 As shown, the process includes:
[0243] 701. The computer device obtains the first category to which each target word contained in the target medical record belongs based on the first medical record classification model.
[0244] In one possible implementation, step 701 includes: for each target word contained in the target medical record, based on the first medical record classification model, obtaining the probability that the target word belongs to each first category, and in response to the maximum probability among the multiple probabilities corresponding to the target word being greater than the target threshold, determining the first category corresponding to the maximum probability as the first category to which the target word belongs.
[0245] The target threshold is an arbitrary value, for example, the target threshold is 0.7 or 0.8.
[0246] Optionally, in response to a plurality of probabilities corresponding to the target word being smaller than a target threshold, it is determined that the first category to which the target word belongs is empty.
[0247] In an embodiment of the present application, if multiple probabilities corresponding to any target word are less than the target threshold, it is determined that the target word is irrelevant to multiple first categories, and the target word is a word in the medical record that is irrelevant to the category to which the medical record belongs.
[0248] 702. The computer device determines the first category to which the target medical record belongs based on the first category to which each target word belongs.
[0249] In one possible implementation, step 702 includes the following two methods:
[0250] The first way: in response to the target medical record containing a target word belonging to at least one first category, determining that the target medical record belongs to the at least one first category.
[0251] In an embodiment of the present application, the target medical record can belong to multiple first categories. For example, if the target medical record includes words belonging to the respiratory category and words belonging to the gastroenterology category, it is determined that the target medical record belongs to the respiratory disease type and also to the gastroenterology disease type.
[0252] The second method is to determine the number of target words belonging to each first category in the target medical record, and determine the first category corresponding to the largest number as the first category to which the target medical record belongs.
[0253] For example, the target medical record contains 10 words, and the first medical record classification model contains 3 first categories. In the target medical record, the number of target words belonging to the first first category is 3, the number of target words belonging to the second first category is 3, and the number of target words belonging to the third first category is 4. It can be determined that the target medical record belongs to the third first category.
[0254] Figure 8 This is a structural diagram of a processing device for a medical record classification model provided in an embodiment of the present application. Figure 8 As shown, the device includes:
[0255] An acquisition module 801 is configured to acquire a first medical record classification model, where the first medical record classification model is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category;
[0256] A training module 802 is configured to form a first sample medical record set from first sample medical records belonging to the newly added category and second sample medical records belonging to the second category, where the second sample medical records are selected from sample medical records used to train the second medical record classification model, and the number of second sample medical records belonging to the same second category is a target number; and to train the first medical record classification model based on the first sample medical record set;
[0257] A forming module 803 is configured to form a second sample medical record set by combining a target number of first sample medical records in the first sample medical record set and second sample medical records in the first sample medical record set;
[0258] The training module 802 is further configured to retrain the first medical record classification model based on the second sample medical record set.
[0259] In one possible implementation, the first sample medical record set includes annotation information of each word contained in each sample medical record, and the annotation information includes a plurality of sample labels, each sample label indicating whether the corresponding word belongs to a first category; Figure 9 As shown, the training module 802 includes:
[0260] an acquiring unit 8021 configured to acquire, based on the first medical record classification model, first prediction information for each first word in a third sample medical record, the first prediction information including a plurality of first prediction labels, each first prediction label indicating a likelihood that the corresponding first word belongs to a first category, the third sample medical record being any sample medical record in the first sample medical record set;
[0261] The training unit 8022 is configured to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word.
[0262] In another possible implementation, the acquisition unit 8021 is used to obtain, for each first word, a first similarity between the first word and each first category based on the first medical record classification model; based on the first medical record classification model and the first similarity corresponding to each first category, normalize each first similarity separately to obtain multiple first prediction labels for the first word, and constitute the first prediction information of the first word with the multiple first prediction labels.
[0263] In another possible implementation, the training unit 8022 is configured to train the first medical record classification model based on the difference between the sample label and the first predicted label of each first word in the same first category.
[0264] In another possible implementation, the apparatus further includes:
[0265] The acquisition module 801 is further configured to acquire, for each first word, a second similarity between the first word and each second category based on the second medical record classification model;
[0266] The training unit 8022 is used to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the first difference information of each first word, where the first difference information includes multiple first similarity differences, and each first similarity difference includes the difference between the first similarity and the second similarity of the corresponding first word in a second category.
[0267] In another possible implementation, the first medical record classification model includes a weight feature of the first category; the acquisition unit 8021 is used to obtain the first word feature of the first word based on the first medical record classification model, and obtain the first similarity between the first word feature and the weight feature of each first category.
[0268] In another possible implementation, the apparatus further includes:
[0269] The acquisition module 801 is further configured to acquire, for each first word, a second word feature of the first word based on the second medical record classification model;
[0270] The training unit 8022 is used to train the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word.
[0271] In another possible implementation, the second sample medical record set includes annotation information of each word contained in each sample medical record, and the annotation information includes multiple sample labels, each sample label indicating whether the corresponding word belongs to a first category; Figure 9As shown, the training module 802 includes:
[0272] an acquiring unit 8021 for acquiring, based on the first medical record classification model, second prediction information for each second word in a fourth sample medical record, the second prediction information comprising a plurality of second prediction labels, each second prediction label indicating a likelihood that the corresponding second word belongs to a first category, the fourth sample medical record being any sample medical record in the second sample medical record set;
[0273] The training unit 8022 is configured to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word.
[0274] In another possible implementation, the acquisition unit 8021 is used to obtain, for each second word, the third similarity between the second word and each first category based on the first medical record classification model; based on the first medical record classification model and the third similarity corresponding to each first category, normalize each third similarity separately to obtain multiple second prediction labels for the second word, and constitute the second prediction information of the second word with the multiple second prediction labels.
[0275] In another possible implementation, Figure 9 As shown, the device also includes:
[0276] A creation module 804 is configured to create a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0277] The acquisition module 801 is further configured to acquire, for each second word, a fourth similarity between the second word and each first category based on the third medical record classification model;
[0278] Training unit 8022 is used to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the second difference information of each second word, the second difference information including multiple second similarity differences, each second similarity difference including the difference between the third similarity and the fourth similarity of the corresponding second word in a first category.
[0279] In another possible implementation, the first medical record classification model includes weight features of the first category; the acquisition unit 8021 is used to obtain the third word features of the second word based on the first medical record classification model, and obtain the third similarity between the third word features and the weight features of each first category.
[0280] In another possible implementation, Figure 9 As shown, the device also includes:
[0281] A creation module 804 is configured to create a third medical record classification model that is identical to the first medical record classification model trained based on the first sample medical record set;
[0282] The acquisition module 801 is further configured to acquire, for each second word, a fourth word feature of the second word based on the third medical record classification model;
[0283] The training unit 8022 is used to train the first medical record classification model based on the difference between the second prediction information and the annotation information of each second word, and the difference between the third word feature and the fourth word feature of each second word.
[0284] In another possible implementation, module 803 is configured to obtain the medical record features of each first sample medical record and the feature center of the first sample medical record in the first sample medical record set; based on the distance between the medical record features and the feature center of each first sample medical record, a target number of first sample medical records are selected from the first sample medical record set, wherein the distance corresponding to the selected first sample medical records is smaller than the distance corresponding to the remaining first sample medical records in the first sample medical record set; and the selected first sample medical records and the second sample medical records in the first sample medical record set constitute a second sample medical record set.
[0285] In another possible implementation, module 803 is configured to obtain, for each first sample medical record, the word features of each third word in the first sample medical record based on the first medical record classification model; obtain the medical record features of the first sample medical record based on the word features of each third word in the first sample medical record; and obtain the feature center based on the medical record features of each first sample medical record.
[0286] In another possible implementation, Figure 9 As shown, the device also includes:
[0287] The acquisition module 801 is further configured to acquire the category to which each target word contained in the target medical record belongs based on the first medical record classification model;
[0288] The determination module 805 is configured to determine the category to which the target medical record belongs based on the category to which each target word belongs.
[0289] It should be noted that the processing device for the medical record classification model provided in the above embodiment is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the processing device for the medical record classification model provided in the above embodiment and the embodiment of the method for processing the medical record classification model are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0290] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the processing method of the medical record classification model of the above embodiment.
[0291] Optionally, the computer device is provided as a terminal. Figure 10 FIG1 shows a block diagram of a terminal 1000 provided by an exemplary embodiment of the present application. The terminal 1000 includes a processor 1001 and a memory 1002 .
[0292] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0293] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one computer program, which is used to be executed by the processor 1001 to implement the processing method of the medical record classification model provided in the method embodiment of the present application.
[0294] In some embodiments, terminal 1000 may optionally include a peripheral device interface 1003 and at least one peripheral device. Processor 1001, memory 1002, and peripheral device interface 1003 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1003 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1009.
[0295] The peripheral device interface 1003 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0296] The RF circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 1004 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 1004 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to Near Field Communication (NFC), which is not limited in this application.
[0297] The display screen 1005 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1005 is a touch screen display, the display screen 1005 also has the ability to collect touch signals on the surface or above the surface of the display screen 1005. The touch signal can be input as a control signal to the processor 1001 for processing. At this time, the display screen 1005 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 1005, which is set on the front panel of the terminal 1000; in other embodiments, there can be at least two display screens 1005, which are respectively set on different surfaces of the terminal 1000 or in a folding design; in other embodiments, the display screen 1005 can be a flexible display screen, which is set on the curved surface or folding surface of the terminal 1000. Even more, the display screen 1005 can be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0298] The camera assembly 1006 is used to capture images or videos. Optionally, the camera assembly 1006 includes a front camera and a rear camera. The front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0299] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1001 for processing, or input into the RF circuit 1004 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1000. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1001 or the RF circuit 1004 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 1007 may also include a headphone jack.
[0300] Power supply 1009 is used to power various components in terminal 1000. Power supply 1009 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1009 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0301] Those skilled in the art will understand that Figure 10 The structure shown in the figure does not constitute a limitation on the terminal 1000, and the terminal 1000 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0302] Optionally, the computer device is provided as a server. Figure 11 1 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1100 may vary significantly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1101 and one or more memories 1102, wherein the memories 1102 store at least one computer program, which is loaded and executed by the processor 1101 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be described in detail here.
[0303] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program. The at least one computer program is loaded and executed by a processor to implement the operations performed in the processing method of the medical record classification model of the above embodiment.
[0304] The present application also provides a computer program product or computer program, which includes computer program code stored in a computer-readable storage medium. A processor of a computer device reads the computer program code from the computer-readable storage medium and executes the computer program code, causing the computer device to implement the operations performed in the method for processing a medical record classification model in the above-described embodiment.
[0305] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0306] The above description is merely an optional embodiment of the embodiments of the present application and is not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for processing a medical record classification model, characterized in that: The method comprises: Obtaining a first medical record classification model, where the first medical record classification model is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category; A first sample medical record set is formed by first sample medical records belonging to the newly added category and second sample medical records belonging to the second category, where the second sample medical records are selected from sample medical records used to train the second medical record classification model, and the number of second sample medical records belonging to the same second category is a target number; For each first word in the third sample medical record, based on the first medical record classification model, obtain a first similarity between the first word and each of the first categories, perform normalization processing on each of the first similarities to obtain multiple first prediction labels for the first word, and use the multiple first prediction labels to form first prediction information for the first word, where each first prediction label indicates a likelihood that the corresponding first word belongs to one of the first categories, and the third sample medical record is any sample medical record in the first sample medical record set; training the first medical record classification model based on the first prediction information of each of the first words; The target number of the first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set form a second sample medical record set; Based on the second sample medical record set, the first medical record classification model is trained again.
2. The method according to claim 1, characterized in that The first sample medical record set includes annotation information of each word contained in each sample medical record, and the annotation information includes a plurality of sample labels, each sample label indicating whether the corresponding word belongs to one of the first categories; The training of the first medical record classification model based on the first prediction information of each first word includes: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word.
3. The method according to claim 2, characterized in that The method further comprises: For each of the first words, obtaining a second similarity between the first word and each of the second categories based on the second medical record classification model; The training of the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word includes: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, and the first difference information of each first word, wherein the first difference information includes multiple first similarity differences, and each first similarity difference includes the difference between the first similarity and the second similarity of the corresponding first word in a second category.
4. The method according to claim 2, characterized in that The first medical record classification model includes weight features of the first category; The obtaining, based on the first medical record classification model, a first similarity between the first word and each of the first categories includes: Based on the first medical record classification model, a first word feature of the first word is obtained, and a first similarity between the first word feature and the weight feature of each of the first categories is obtained.
5. The method according to claim 4, characterized in that The method further comprises: For each of the first words, obtaining a second word feature of the first word based on the second medical record classification model; The training of the first medical record classification model based on the difference between the first prediction information and the annotation information of each first word includes: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word.
6. The method according to any one of claims 1 to 5, characterized in that The second sample medical record set includes labeling information of each word contained in each sample medical record, and the labeling information includes a plurality of sample labels, each of the sample labels indicating whether the corresponding word belongs to one of the first categories; The retraining of the first medical record classification model based on the second sample medical record set includes: obtaining, based on the first medical record classification model, second prediction information for each second word in a fourth sample medical record, the second prediction information including a plurality of second prediction labels, each second prediction label indicating a likelihood that the corresponding second word belongs to one of the first categories, the fourth sample medical record being any sample medical record in the second sample medical record set; The first medical record classification model is trained based on the difference between the second prediction information and the annotation information of each second word.
7. A processing device for a medical record classification model, characterized in that: The device comprises: an acquisition module, configured to acquire a first medical record classification model, wherein the first medical record classification model is obtained by adjusting the second medical record classification model, wherein the first medical record classification model includes a first category, the second medical record classification model includes a second category, and the first category includes the second category and a newly added category; a training module, configured to form a first sample medical record set from first sample medical records belonging to the newly added category and second sample medical records belonging to the second category, wherein the second sample medical records are selected from sample medical records for training the second medical record classification model, and the number of second sample medical records belonging to the same second category is a target number; The training module is further configured to obtain, for each first word in a third sample medical record, a first similarity between the first word and each of the first categories based on the first medical record classification model, perform normalization processing on each of the first similarities to obtain a plurality of first prediction labels for the first word, and form first prediction information for the first word from the plurality of first prediction labels, wherein each first prediction label indicates a likelihood that the corresponding first word belongs to one of the first categories, and the third sample medical record is any sample medical record in the first sample medical record set; The training module is further configured to train the first medical record classification model based on the first prediction information of each first word; a forming module, configured to form a second sample medical record set by combining the target number of the first sample medical records in the first sample medical record set and the second sample medical records in the first sample medical record set; The training module is further used to retrain the first medical record classification model based on the second sample medical record set.
8. The device according to claim 7, characterized in that The first sample medical record set includes labeling information of each word contained in each sample medical record, and the labeling information includes multiple sample labels, each sample label indicating whether the corresponding word belongs to one of the first categories; the training module is used to: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word.
9. The device according to claim 8, characterized in that The device further comprises: The acquisition module is further configured to acquire, for each of the first words, a second similarity between the first word and each of the second categories based on the second medical record classification model; The training module is used to: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, and the first difference information of each first word, wherein the first difference information includes multiple first similarity differences, and each first similarity difference includes the difference between the first similarity and the second similarity of the corresponding first word in a second category.
10. The device according to claim 8, characterized in that The first medical record classification model includes weight features of the first category; the training module is further used to: Based on the first medical record classification model, a first word feature of the first word is obtained, and a first similarity between the first word feature and the weight feature of each of the first categories is obtained.
11. The device according to claim 10, characterized in that The device further comprises: The acquisition module is further configured to acquire, for each of the first words, a second word feature of the first word based on the second medical record classification model; The training module is used to: The first medical record classification model is trained based on the difference between the first prediction information and the annotation information of each first word, and the difference between the first word feature and the second word feature of each first word.
12. The device according to any one of claims 7 to 11, characterized in that: The second sample medical record set includes labeling information of each word contained in each sample medical record, and the labeling information includes multiple sample labels, each of which indicates whether the corresponding word belongs to one of the first categories; the training module is used to: obtaining, based on the first medical record classification model, second prediction information for each second word in a fourth sample medical record, the second prediction information including a plurality of second prediction labels, each second prediction label indicating a likelihood that the corresponding second word belongs to one of the first categories, the fourth sample medical record being any sample medical record in the second sample medical record set; The first medical record classification model is trained based on the difference between the second prediction information and the annotation information of each second word.
13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the processing method of the medical record classification model according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the operations performed in the method for processing a medical record classification model according to any one of claims 1 to 6.
15. A computer program product, characterized in that The computer program product includes a computer program code, which is stored in a computer-readable storage medium. The processor of a computer device reads the computer program code from the computer-readable storage medium, and the processor executes the computer program code, so that the computer device implements the operations performed in the processing method of the medical record classification model according to any one of claims 1 to 6.
Citation Information
Patent Citations
Model training method and device based on electronic medical record, equipment and storage medium
CN112614562A