A method for constructing a semantic tag determination model and a medical record parsing method
By constructing a semantic label determination model, the semantic analysis problem caused by irregular writing in electronic medical records is solved, and a more accurate semantic analysis effect is achieved.
Patent Information
- Application Number
- CN202111152503.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-29
AI Technical Summary
There is a problem of irregular writing in electronic medical records, which leads to difficult semantic analysis process and poor analysis results.
By constructing a semantic label determination model, using sample medical record text and actual semantic label information for model training, improve the semantic label determination performance, thereby improving the semantic analysis effect of medical record text.
It improves the semantic analysis effect of medical record data, can more accurately describe the semantic information in medical record text, and enhances the understanding and analysis ability of medical record text data.
Smart Images

Figure CN113903420B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent medical technology, and particularly to a method for constructing a semantic tag determination model and a medical record parsing method. Background Art
[0002] An electronic medical record refers to a carrier of various medical texts and images during a patient's hospitalization process, so that the electronic medical record can be used to record the medical activities of medical staff in the processes of examining, diagnosing, treating, etc. for the occurrence, development, and outcome of the patient's disease; and the electronic medical record usually includes outpatient (emergency) electronic medical records, inpatient electronic medical records, and other electronic medical records, etc.
[0003] However, due to the defect of non-standard writing in a large number of electronic medical records, the semantic parsing process for these electronic medical records is relatively difficult, resulting in relatively poor semantic parsing results for these electronic medical records. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to provide a method for constructing a semantic tag determination model and a medical record parsing method, which can improve the semantic parsing effect for medical record data.
[0005] The embodiments of this application provide a method for constructing a semantic tag determination model, and the method includes:
[0006] Obtain a sample medical record text and the actual semantic tag information of the sample medical record text;
[0007] Determine the predicted semantic tag information of the sample medical record text according to the sample medical record text and the model to be trained;
[0008] Update the model to be trained according to the predicted semantic tag information and the actual semantic tag information, and continue to execute the step of obtaining the predicted semantic tag information of the sample medical record text according to the sample medical record text and the model to be trained, until when a preset stop condition is reached, determine a semantic tag determination model according to the model to be trained.
[0009] The embodiments of this application also provide a medical record parsing method, and the method includes:
[0010] Obtain a medical record text to be processed;
[0011] Determine the predicted semantic tag information of the medical record text to be processed according to the medical record text to be processed and the semantic tag determination model; wherein, the semantic tag determination model is constructed by using any implementation manner of the method for constructing a semantic tag determination model provided by the embodiments of this application;
[0012] Determine the semantic parsing result of the medical record text to be processed according to the predicted semantic tag information of the medical record text to be processed.
[0013] The embodiment of the present application also provides a device for constructing a semantic tag determination model, including:
[0014] A first acquisition unit, configured to acquire a sample medical record text and the actual semantic tag information of the sample medical record text;
[0015] A first determination unit, configured to determine the predicted semantic tag information of the sample medical record text according to the sample medical record text and the model to be trained;
[0016] A model update unit, configured to update the model to be trained according to the predicted semantic tag information and the actual semantic tag information, and return to the first determination unit to continue to execute the step of obtaining the predicted semantic tag information of the sample medical record text according to the sample medical record text and the model to be trained, until when a preset stop condition is reached, determine a semantic tag determination model according to the model to be trained.
[0017] The embodiment of the present application also provides a medical record parsing device, including:
[0018] A second acquisition unit, configured to acquire a medical record text to be processed;
[0019] A second determination unit, configured to determine the predicted semantic tag information of the medical record text to be processed according to the medical record text to be processed and the semantic tag determination model; wherein, the semantic tag determination model is constructed by any implementation manner of the method for constructing a semantic tag determination model provided by the embodiment of the present application;
[0020] A third determination unit, configured to determine the semantic parsing result of the medical record text to be processed according to the predicted semantic tag information of the medical record text to be processed.
[0021] The embodiment of the present application also provides a device, the device includes: a processor, a memory, and a system bus;
[0022] The processor and the memory are connected through the system bus;
[0023] The memory is used to store one or more programs, the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes any implementation manner of the method for constructing a semantic tag determination model provided by the embodiment of the present application, or executes any implementation manner of the medical record parsing method provided by the embodiment of the present application.
[0024] The embodiments of the present application also provide a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a terminal device, the terminal device is caused to execute any implementation manner of the method for constructing a semantic tag determination model provided by the embodiments of the present application, or execute any implementation manner of the medical record parsing method provided by the embodiments of the present application.
[0025] The embodiments of the present application also provide a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute any implementation manner of the method for constructing a semantic tag determination model provided by the embodiments of the present application, or execute any implementation manner of the medical record parsing method provided by the embodiments of the present application.
[0026] Based on the above technical solutions, the present application has the following beneficial effects:
[0027] In the technical solution provided by the present application, first, a semantic tag determination model is constructed by using sample medical record texts and the actual semantic tag information of the sample medical record texts, so that the constructed semantic tag determination model has good semantic tag determination performance, so that the semantic tag determination model can perform relatively accurate semantic tag determination processing on a medical record text data (for example, a to-be-processed medical record text). Furthermore, the predicted semantic tag information determined by using the semantic tag determination model can relatively accurately describe the field information of at least one string in the medical record text data (for example, field names such as "symptom manifestation", "drug history", "trauma history", "family health status", "family infectious disease history", "family genetic disease history", etc.). In this way, the semantic parsing result of the medical record text data determined based on the predicted semantic tag information can relatively accurately describe the semantic information carried in the medical record text data (for example, "symptom manifestation" is... etc.). In this way, it is beneficial to improve the semantic parsing effect for medical record data. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0029] Figure 1 It is a schematic diagram of a medical record text data provided by the embodiments of the present application;
[0030] Figure 2 It is a flowchart of a method for constructing a semantic tag determination model provided by the embodiments of the present application;
[0031] Figure 3 A structural schematic diagram of a model to be trained provided by an embodiment of the present application;
[0032] Figure 4 A schematic diagram of a model to be trained provided by an embodiment of the present application;
[0033] Figure 5 A structural schematic diagram of another model to be trained provided by an embodiment of the present application;
[0034] Figure 6 A flowchart of a medical record parsing method provided by an embodiment of the present application;
[0035] Figure 7 A structural schematic diagram of a device for constructing a semantic label determination model provided by an embodiment of the present application;
[0036] Figure 8 A structural schematic diagram of a medical record parsing device provided by an embodiment of the present application. Detailed implementation manners
[0037] The inventors found in the research on electronic medical records that a large number of electronic medical records have the defect of non-standard writing. For example, some electronic medical records may have the phenomenon of missing field names (such as "symptom manifestation", "drug history", "symptom history", "trauma history", "family health status", "family infectious disease history", "family genetic disease history", etc.). Another example is that some electronic medical records may have the chaotic phenomenon of interspersing semantic description contents under different field names (as shown in Figure 1 shown, the description content of "drug history" and the description content of "negative symptom history" are interspersed in the description content of "symptom history", etc.). It can be seen that due to the defect of non-standard writing of these electronic medical records, the semantic parsing process for these electronic medical records is relatively difficult, resulting in relatively poor semantic parsing results for these electronic medical records.
[0038] Based on the above findings, in order to overcome the technical problems in the background art, the embodiments of the present application provide a method for constructing a semantic tag determination model and a medical record parsing method, which specifically include: first, using the sample medical record text and the actual semantic tag information of the sample medical record text to construct a semantic tag determination model, so that the constructed semantic tag determination model has better semantic tag determination performance, so that the semantic tag determination model can perform relatively accurate semantic tag determination processing on a medical record text data (for example, the medical record text to be processed), and further, the predicted semantic tag information determined by using the semantic tag determination model can relatively accurately describe the field information of at least one string in the medical record text data (for example, field names such as "symptom manifestation", "drug history", "trauma history", "family health status", "family infectious disease history", "family genetic disease history", etc.). In this way, the semantic parsing result of the medical record text data determined based on the predicted semantic tag information can relatively accurately describe the semantic information carried in the medical record text data (for example, "symptom manifestation" is... etc.), which is beneficial to improving the semantic parsing effect of medical record data.
[0039] In addition, the embodiments of the present application do not limit the execution entity of the method for constructing the semantic tag determination model. For example, the method for constructing the semantic tag determination model provided by the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (Personal Digital Assitant, PDA), or a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server.
[0040] In addition, the embodiments of the present application do not limit the execution entity of the medical record parsing method. For example, the medical record parsing method provided by the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (Personal Digital Assitant, PDA), or a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server.
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0042] To facilitate the understanding of the technical solution provided in this application, the related content of the "method for constructing a semantic tag determination model" and the related content of the "medical record parsing method" will be introduced successively below.
[0043] Method Embodiment 1
[0044] See Figure 2 , which is a flowchart of a method for constructing a semantic tag determination model provided in an embodiment of this application.
[0045] The method for constructing a semantic tag determination model provided in an embodiment of this application includes S201 - S205:
[0046] S201: Obtain a sample medical record text and the actual semantic tag information of the sample medical record text.
[0047] The above "sample medical record text" refers to a section of medical record content extracted from an electronic medical record; and the embodiment of this application does not limit the "sample medical record text". For example, it can be Figure 1 the medical record text data shown.
[0048] In addition, the embodiment of this application does not limit the number of the above "sample medical record texts". For example, it can be G. Where G is a positive integer (for example, G = 9600).
[0049] The above "actual semantic tag information of the sample medical record text" is used to represent the actual field information of each string in the sample medical record text.
[0050] It should be noted that the embodiment of this application does not limit the "field information". For example, it can include a field name and / or the correspondence between the field name and the string. Among them, the "field name" is used to describe the medical record connotation information of a string; and the embodiment of this application does not limit the "field name". For example, it can specifically be "symptom manifestation", "drug history", "trauma history", "family health status", "family infectious disease history", "family genetic disease history", or "others", etc. The "correspondence between the field name and the string" refers to the correspondence between a string in the medical record text data and the medical record connotation information of the string; and the embodiment of this application does not limit the "correspondence between the field name and the string". For example, as Figure 1 shown, it can include: the correspondence between the string "male, 38 years old" and the field name "others", the correspondence between the string "recurrent epigastric pain for 10 years, recurrence for 1 week, melena for 2 days" and the field name "symptom manifestation",...
[0051] In addition, the embodiment of this application does not limit the acquisition process of the above "actual semantic tag information of the sample medical record text". For example, it can specifically include Step 11 - Step 12:
[0052] Step 11: Extract K candidate semantic tags from the preset medical record specification data. Here, K is a positive integer.
[0053] The above-mentioned "preset medical record specification data" refers to the specification writing rules for describing an electronic medical record; and the embodiments of the present application do not limit the "preset medical record specification data". For example, it may include the "Medical Record Writing Specification 2014 Edition".
[0054] The above-mentioned "candidate semantic tags" refer to the field names that a standardized electronic medical record may involve, so that the "candidate semantic tags" can be used to label the medical record connotation information representing a string.
[0055] In addition, the embodiments of the present application do not limit the value of the above-mentioned "K". For example, when the above-mentioned "preset medical record specification data" is the "Medical Record Writing Specification 2014 Edition", K can be 61, and these 61 candidate semantic tags involve 9 field names in the aspect of medical record writing. Among them, the above-mentioned "9 aspects of medical record writing" include symptoms, treatment, examination, physical examination, patient conditions, other past medical histories, differential diagnosis, surgery, and others.
[0056] Step 12: Send the sample medical record text and K candidate semantic tags to the technical personnel, so that the technical personnel label the actual semantic tag information of the sample medical record text according to the K candidate semantic tags, so that the "actual semantic tag information of the sample medical record text" includes at least one of the K candidate semantic tags.
[0057] Based on the relevant content of the above Step 11 to Step 12, it can be seen that in order to improve the standardization of medical record semantic parsing, some standardized candidate semantic tags can be extracted from the preset medical record specification data first; then, based on these candidate semantic tags, the technical personnel label the actual semantic tag information of the sample medical record text, so that the actual semantic tag information can accurately and standardly represent the actual field information of each string in the sample medical record text.
[0058] S202: Determine the predicted semantic tag information of the sample medical record text according to the sample medical record text and the model to be trained.
[0059] The above-mentioned "model to be trained" is used to perform semantic tag determination processing on the input data of the model to be trained; and the embodiments of the present application do not limit the model to be trained. For example, it can be any machine learning model. Another example is that it can be implemented in any implementation manner of the model to be trained shown below. Method Embodiment 2 shown below.
[0060] The above "predicted semantic tag information of the sample medical record text" is used to represent the field information predicted for at least one string in the sample medical record text.
[0061] In addition, the embodiments of the present application do not limit the implementation manner of S202. For example, it can be implemented by any of the implementation manners shown in steps 21 - 24 below. Also, in a possible implementation manner, when the model to be trained does not include the "text segmentation layer" below (for example, Figure 3 the model 300 to be trained shown), S202 may specifically include S2021 - S2022:
[0062] S2021: Perform a preset partitioning process on the sample medical record text to obtain at least one sample segment.
[0063] The above "preset partitioning process" can be set in advance; moreover, the embodiments of the present application do not limit this "preset partitioning process". For example, it may specifically include: performing a partitioning process according to a preset partitioning rule.
[0064] The above "preset partitioning rule" refers to the rule required for partitioning a medical record text data; moreover, the embodiments of the present application do not limit this "preset partitioning rule". For example, it may specifically include: partitioning according to the positions of preset punctuation marks (such as commas, periods, semicolons, etc.). Among them, the "preset punctuation marks" can be set in advance or obtained by analyzing a large amount of medical record data (such as the above "G sample medical record texts", etc.) through big data analysis means.
[0065] In addition, to improve the text partitioning effect, it can be implemented by means of a machine learning model. Based on this, it can be known that S2021 may specifically include: inputting the sample medical record text into a pre-constructed text partitioning model to obtain at least one sample segment output by the text partitioning model.
[0066] It should be noted that the above "text partitioning model" is used to perform partitioning processing on the input data of this text partitioning model; moreover, this "text partitioning model" can be constructed according to the medical record text to be used and the actual partitioning result of this medical record text to be used. Among them, the "actual partitioning result of the medical record text to be used" refers to the result obtained by technicians manually partitioning this medical record text to be used, so that the "actual partitioning result of the medical record text to be used" is used to represent each text segment actually partitioned in this medical record text to be used. In addition, the embodiment of the present application does not limit the construction process of the "text partitioning model", and any existing or future model construction method can be used for implementation. In addition, the embodiment of the present application does not limit the relationship between the above "medical record text to be used" and the above "G sample medical record texts". For example, in order to save data acquisition resources, the above "medical record text to be used" can come from the above "G sample medical record texts".
[0067] Based on the relevant content of S2021 above, after obtaining the sample medical record text, the sample medical record text can be first partitioned to obtain at least one sample segment of the sample medical record text, so that these sample segments can accurately represent the character information carried by the sample medical record text, so as to be able to implement the semantic label determination process for the sample medical record text based on these sample segments subsequently.
[0068] S2022: Input at least one sample segment into the model to be trained, and obtain the predicted semantic label information of the sample medical record text output by the model to be trained, so that the predicted semantic label information includes the segment semantic label information of the at least one sample segment.
[0069] Among them, the segment semantic label information of the f-th sample segment is used to represent the field information predicted for the f-th sample segment in the sample medical record text. Among them, f is a positive integer, f ≤ F, F is a positive integer, and F represents the number of sample segments in the sample medical record text.
[0070] Based on the relevant content of S2022 above, after obtaining at least one sample segment (for example, Figure 4After multiple segments such as "headache for three days", "accompanied by cough and expectoration", and "ineffective with Ganmao tablets" are shown, these sample segments can be input into the model to be trained, so that the model to be trained performs semantic label determination processing on these sample segments, obtains the segment semantic label information of these sample segments, and outputs the segment semantic label information of these sample segments as the predicted semantic label information of the sample medical record text, so that the "predicted semantic label information of the sample medical record text" can represent the field information of each sample segment in the sample medical record text, including the possibility of each candidate semantic label, so as to subsequently measure the semantic label determination performance of the model to be trained with the help of the "predicted semantic label information of the sample medical record text".
[0071] Based on the relevant content of S202 above, after obtaining the sample medical record text, the model to be trained can be used to perform semantic label determination processing on each string in the sample medical record text (for example, the above "sample segment" or the above "text segment"), and obtain the predicted semantic label information of the sample medical record text, so that the predicted semantic label information can represent the predicted field information for each string in the sample medical record text, so as to subsequently measure the semantic label determination performance of the model to be trained with the help of the "predicted semantic label information of the sample medical record text".
[0072] S203: Determine whether the preset stop condition is reached. If so, execute S205; if not, execute S204.
[0073] Among them, the "preset stop condition" can be set in advance; and the embodiments of the present application do not limit the "preset stop condition". For example, it can specifically be that the model prediction loss value of the model to be trained is lower than the preset loss threshold, or the change rate of the model prediction loss value of the model to be trained is lower than the preset change rate threshold, or the number of updates of the model to be trained reaches the preset number threshold.
[0074] The above "model prediction loss value of the model to be trained" is used to represent the semantic label determination performance of the model to be trained; and the embodiments of the present application do not limit the determination process of the "model prediction loss value of the model to be trained". For example, any existing or future model loss value calculation method (such as the loss value calculation method based on cross-entropy) can be used for implementation. Another example is that Method Embodiment 3 Any possible implementation manner of determining the "model prediction loss value of the model to be trained" shown can be used for implementation.
[0075] S204: Update the model to be trained according to the predicted semantic label information of the sample medical record text and the actual semantic label information of the sample medical record text, and return to execute S202.
[0076] In an embodiment of the present application, if it is determined that the model to be trained in the current round does not meet the preset stop condition, it can be determined that the semantic label determination performance of the model to be trained is still relatively poor. Therefore, the model to be trained can be updated according to the gap (such as the Euclidean distance, etc.) between the predicted semantic label information of the sample medical record text and the actual semantic label information of the sample medical record text, so that the updated model to be trained has better semantic label determination performance, and based on the updated model to be trained, S202 and its subsequent steps are continued to be executed.
[0077] It should be noted that the embodiment of the present application does not limit the update process of the model to be trained. For example, any existing or future model update method can be used for implementation. Another example is that any model update method shown in Method Embodiment 4 can be used for implementation.
[0078] S205: Determine a semantic label determination model according to the model to be trained.
[0079] In an embodiment of the present application, if it is determined that the model to be trained in the current round meets the preset stop condition, it can be determined that the model to be trained has good semantic label determination performance. Therefore, a semantic label determination model can be determined according to the model to be trained (for example, the model to be trained can be directly determined as the semantic label determination model), so that the semantic label determination model also has good semantic label determination performance, so that subsequent semantic label determination processing can be based on the semantic label determination model (such as Method Embodiment 5 shown).
[0080] Based on the relevant content of S201 to S205 above, for the method for constructing a semantic label determination model provided in the embodiment of the present application, after obtaining the sample medical record text and the actual semantic label information of the sample medical record text, first determine the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained; then update the model to be trained according to the predicted semantic label information and the actual semantic label information, and continue to execute the above step of "obtaining the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained" until when the preset stop condition is reached, determine the semantic label determination model according to the model to be trained.
[0081] It can be seen that since the model to be trained is trained according to the sample medical record text and its actual semantic label information, the model to be trained can learn the semantic label determination rule from the sample medical record text and its actual semantic label information, so that the trained model to be trained has good semantic label determination performance. Furthermore, the semantic label determination model determined based on the trained model to be trained also has good semantic label determination performance. In this way, the semantic label determination model can perform relatively accurate semantic label determination processing on a medical record text data (for example, the medical record text to be processed).
[0082] Method Embodiment 2
[0083] To further improve the semantic label determination performance of the model to be trained, an exemplary implementation of the model to be trained is provided in an embodiment of the present application. As Figure 3 shown, the model to be trained 300 may include: a text encoding layer 301, an expert encoding layer 302, an expert weight determination layer 303, and a decision layer 304. Among them, the expert encoding layer 302 includes M expert encoding networks, and the input data of each expert encoding network includes the output data of the text encoding layer 301; the input data of the expert weight determination layer 303 includes the output data of the text encoding layer 301; the input data of the decision layer 304 includes the output data of each expert encoding network and the output data of the expert weight determination layer 303.
[0084] To facilitate understanding of the working principle of the model to be trained 300, the determination process of the above-mentioned "predicted semantic label information of the sample medical record text" will be used as an example for illustration below.
[0085] As an example, the process of using the model to be trained 300 to determine the above-mentioned "predicted semantic label information of the sample medical record text" may specifically include steps 21-24:
[0086] Step 21: Determine the text encoding result to be used according to the sample medical record text and the text encoding layer 301.
[0087] Among them, the "text encoding layer 301" is used to perform encoding processing on the input data of the text encoding layer 301; and the embodiment of the present application does not limit the "text encoding layer 301", and any encoding network can be used for implementation.
[0088] The above-mentioned "text encoding result to be used" is used to represent the semantic information carried by the above-mentioned "sample medical record text".
[0089] In addition, the embodiments of the present application do not limit the implementation manner of step 21. For example, the sample medical record text can be directly input into the text encoding layer 301, so that the text encoding layer 301 performs encoding processing on the sample medical record text to obtain and output the text encoding result to be used.
[0090] Actually, in some medical record text data, there may be a chaotic phenomenon where semantic description contents under different field names are interspersed (as Figure 1 shown, in the description content of "symptom history", the description content of "drug history" and the description content of "negative symptom history" are interspersed, etc.). Therefore, in order to improve the semantic parsing effect, the medical record text data can be divided into multiple text segments, so that subsequent semantic parsing processing can be performed on these text segments respectively (as Figure 4 shown). Based on this, the embodiments of the present application provide two possible implementation manners of step 21, and the two possible implementation manners are introduced below respectively.
[0091] In the first possible implementation manner, step 21 may specifically include steps 31 - 32:
[0092] Step 31: Perform a preset partitioning process on the sample medical record text to obtain at least one sample segment.
[0093] It should be noted that for the relevant content of step 31, please refer to the relevant content of S2021 above.
[0094] Step 32: Input at least one sample segment into the text encoding layer 301 to obtain the text encoding result to be used output by the text encoding layer 301, so that the text encoding result to be used includes the text encoding results of each sample segment.
[0095] Among them, the text encoding result of the f-th sample segment is used to represent the semantic information carried by the f-th sample segment in the sample medical record text. f is a positive integer, f ≤ F, F is a positive integer, and F represents the number of sample segments in the sample medical record text.
[0096] In addition, the embodiments of the present application do not limit the implementation manner of step 32. For example, when the text encoding layer 301 includes a first encoding network and a second encoding network, step 32 may specifically include steps 321 - 322:
[0097] Step 321: Input the f-th sample segment into the first encoding network to obtain the preliminary encoding result of the f-th sample segment output by the first encoding network. Among them, f is a positive integer, f ≤ F.
[0098] Among them, the "first encoding network" is used to perform encoding processing on the input data of the first encoding network; moreover, the embodiments of the present application do not limit the "first encoding network". For example, it can be implemented using any encoding network. For another example, it can be implemented by combining a bidirectional long short-term memory network (BiLSTM) and an attention mechanism (Attention).
[0099] The above "preliminary encoding result of the f-th sample segment" is used to represent the semantic information carried by the f-th sample segment.
[0100] Step 322: Input the preliminary encoding results of the F sample segments into the second encoding network to obtain the text encoding result to be used output by the second encoding network, so that the text encoding result to be used includes the text encoding results of the F sample segments.
[0101] Among them, the "second encoding network" is used to perform encoding processing on the input data of the second encoding network; moreover, the embodiments of the present application do not limit the working principle of the "second encoding network". For example, it may specifically include: referring to the preliminary encoding results of at least one other sample segment except the preliminary encoding result of the f-th sample segment in the preliminary encoding results of the F sample segments, performing secondary encoding processing on the preliminary encoding result of the f-th sample segment to obtain the text encoding result of the f-th sample segment, so that the "text encoding result of the f-th sample segment" can more accurately represent the semantic information carried by the f-th sample segment.
[0102] In addition, the embodiments of the present application do not limit the "second encoding network". For example, it can be implemented using any encoding network. For another example, it can be implemented by combining BiLSTM and an attention mechanism (Attention).
[0103] Based on the relevant content of the above steps 31 to 32, for the first possible implementation manner of step 21, the sample medical record text can be first subjected to a preset partitioning process to obtain at least one sample segment; then, the text encoding layer 301 is used to perform encoding processing on these sample segments to obtain the text encoding results of these sample segments, so that the text encoding results of these sample segments can represent the semantic information carried by the sample medical record text.
[0104] In the second possible implementation manner, as Figure 5 shown, when the model to be trained 300 further includes a text slicing layer 305 and the input data of the text encoding layer 301 includes the output data of the text slicing layer 305, step 21 may specifically include steps 41 - 42:
[0105] Step 41: Input the sample medical record text into the text segmentation layer 305 to obtain at least one text segment output by the text segmentation layer 305.
[0106] Among them, the "text segmentation layer 305" is used to perform segmentation processing on the input data of the text segmentation layer 305; and the embodiment of the present application does not limit the network structure of the "text segmentation layer 305".
[0107] The above "at least one text segment" is used to represent the character information carried by the sample medical record text.
[0108] Step 42: Input at least one text segment into the text encoding layer 301 to obtain the text encoding result to be used output by the text encoding layer 301, so that the text encoding result to be used includes the text encoding results of each text segment.
[0109] It should be noted that the relevant content of step 42 is similar to the relevant content of step 32 above, and only need to replace "sample segment" with "text segment" in the relevant content of step 32 above.
[0110] Based on the relevant content of the above steps 41 to 42, for the second possible implementation manner of step 21, the sample medical record text can be first segmented by the text segmentation layer 305 to obtain at least one text segment; then the text encoding layer 301 is used to perform encoding processing on these text segments to obtain the text encoding results of these text segments, so that the text encoding results of these text segments can represent the semantic information carried by the sample medical record text.
[0111] Based on the relevant content of the above step 21, after obtaining the sample medical record text, the text encoding layer 301 in the model 300 to be trained can be used to perform encoding processing on the character information in the sample medical record text to obtain the text encoding result to be used, so that the text encoding result to be used can represent the semantic information carried by the text encoding result to be used.
[0112] Step 22: Input the text encoding result to be used into the m-th expert encoding network to obtain the m-th expert encoding result output by the m-th expert encoding network. Where m is a positive integer and m ≤ M.
[0113] Among them, the "m-th expert encoding network" is used to perform expert encoding processing on the input data of the m-th expert encoding network according to the m-th expert encoding processing performance; and the embodiment of the present application does not limit the implementation manner of the "m-th expert encoding network". For example, it can be implemented by a multi-layer perceptron (MLP).
[0114] It should be noted that different expert coding networks have different expert coding processing performances when performing expert coding processing, so that different expert coding networks are good at processing different semantic parsing directions when performing expert coding processing. For example, as Figure 4 shown, "private network 1" refers to the first expert coding network; and this "private network 1" is more proficient in processing medical record text data with relatively high semantic parsing difficulty (that is, relatively high semantic label determination difficulty); "private network 2" refers to the second expert coding network; and this "private network 2" is more proficient in processing medical record text data with relatively low semantic parsing difficulty (that is, relatively low semantic label determination difficulty) and relatively rare field names (that is, relatively low field name occurrence frequency); "private network 3" refers to the third expert coding network; and this "private network 3" is more proficient in processing medical record text data with relatively low semantic parsing difficulty and relatively common field names (that is, relatively high field name occurrence frequency).
[0115] Step 23: Input the text coding result to be used into the expert weight determination layer 303 to obtain the predicted expert weight values corresponding to the M expert coding networks output by the expert weight determination layer 303.
[0116] Among them, the "expert weight determination layer 303" is used to determine the decision influence degree of the input data of the expert weight determination layer 303; and the embodiments of the present application do not limit the "expert weight determination layer 303". For example, it can be implemented by means of a fully connected network.
[0117] The predicted expert weight value corresponding to the m-th expert coding network is used to represent the influence degree of the above "m-th expert coding result" on the determination process of the above "predicted semantic label information of the sample medical record text", so that the "predicted expert weight value corresponding to the m-th expert coding network" can represent the influence degree of the m-th expert coding network on the determination process of the "predicted semantic label information of the sample medical record text". Among them, m is a positive integer, m ≤ M.
[0118] Step 24: Input the M expert coding results and the predicted expert weight values corresponding to the M expert coding networks into the decision layer 304 to obtain the predicted semantic label information of the sample medical record text output by the decision layer 304.
[0119] Among them, the "decision layer 304" is used to perform semantic label comprehensive decision processing on the input data of the decision layer 304.
[0120] In addition, the embodiments of the present application do not limit the implementation manner of the "decision layer 304". For example, the decision layer 304 includes an expert decision network and a decision fusion network, and the input data of the decision fusion network includes the output data of the expert decision network.
[0121] To facilitate the understanding of the working principle of the decision-making layer 304, the following takes the determination process of the above-mentioned "predicted semantic label information of the sample medical record text" as an example for illustration.
[0122] As an example, when the decision-making layer 304 includes an expert decision-making network and a decision fusion network, the determination process of the above-mentioned "predicted semantic label information of the sample medical record text" may specifically include steps 51 - step 52:
[0123] Step 51: Input the m-th expert coding result into the expert decision-making network to obtain the m-th expert decision result output by the expert decision-making network. Here, m is a positive integer, and m ≤ M.
[0124] Among them, the "expert decision-making network" is used to perform semantic label decision processing on the input data of the expert decision-making network; moreover, the embodiments of the present application do not limit the implementation manner of the "expert decision-making network". For example, it can be implemented by means of any decoding network (such as a Conditional Random Field (CRF)).
[0125] The above-mentioned "m-th expert decision result" refers to the decoding result obtained by decoding the above-mentioned "m-th expert coding result", so that the "m-th expert decision result" is used to represent the semantic label information predicted for the sample medical record text by the m-th expert network.
[0126] In addition, the embodiments of the present application do not limit the representation manner of the above-mentioned "m-th expert decision result". For example, if the above-mentioned "actual semantic label information of the sample medical record text" is determined from the above-mentioned "K candidate semantic labels", the "m-th expert decision result" can be expressed as Among them, represents the possibility that the field information of the y-th string (for example, the above-mentioned "sample segment" or the above-mentioned "text segment") in the sample medical record text predicted by the m-th expert network includes the k-th candidate semantic label. y is a positive integer, y ≤ Y, Y is a positive integer, and Y represents the number of strings in the sample medical record text; k is a positive integer, k ≤ K, K is a positive integer, and K represents the number of candidate semantic labels.
[0127] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "m-th expert decision result". For example, in order to further improve the determination efficiency of the above-mentioned "m-th expert decision result", the embodiments of the present application also provide another possible implementation manner for determining the above-mentioned "m-th expert decision result", which may specifically include steps 61 - step 62:
[0128] Step 61: Perform a dot product operation on the character feature vector of the k-th candidate semantic label and the coding result of the m-th expert to obtain the k-th label attribution probability information of the coding result of the m-th expert. Here, k is a positive integer, and k ≤ K.
[0129] Among them, the "dot product operation" is used to calculate the dot product between two vector data.
[0130] The above-mentioned "character feature vector of the k-th candidate semantic label" is used to represent the character information carried by the k-th candidate semantic label; moreover, the embodiment of the present application does not limit the determination process of the "character feature vector of the k-th candidate semantic label". For example, a preset vector extraction network can be used to perform feature vector extraction processing on the k-th candidate semantic label to obtain the character feature vector of the k-th candidate semantic label. Among them, the "preset vector extraction network" can be implemented by using any existing or future feature vector extraction network (such as word2vec, BiLSTM, etc.).
[0131] The above-mentioned "k-th label attribution probability information of the coding result of the m-th expert" is used to represent the possibility that the field information of at least one string in the sample medical record text predicted by the m-th expert network includes the k-th candidate semantic label; moreover, the embodiment of the present application does not limit the calculation process of the "k-th label attribution probability information of the coding result of the m-th expert". For example, it can be implemented using formula (1).
[0132]
[0133] In the formula, L m,k represents the above-mentioned "k-th label attribution probability information of the coding result of the m-th expert"; represents the possibility that the field information of the y-th string (such as the above "sample segment" or the above "text segment") in the sample medical record text predicted by the m-th expert network includes the k-th candidate semantic label; represents the coding result of the m-th expert; represents the coding result of the m-th expert network for the y-th string (such as the above "sample segment" or the above "text segment") in the sample medical record text, and is a 1×U vector data; F K represents the above-mentioned "character feature vector of the k-th candidate semantic label", and F K is a U×1 vector data; U is a positive integer.
[0134] Step 62: Perform a set operation on the 1st label attribution probability information to the K-th label attribution probability information of the coding result of the m-th expert to obtain the decision result of the m-th expert.
[0135] In an embodiment of the present application, after obtaining the first label attribution probability information to the Kth label attribution probability information of the mth expert coding result, the K label attribution probability information can be subjected to set processing to obtain the mth expert decision result, so that the mth expert decision result can represent the possibility that the field information of each string (for example, each sample segment or each text segment) in the sample medical record text predicted by the mth expert network includes each candidate semantic label.
[0136] Based on the relevant content of step 51 above, for the to-be-trained model 300, after the mth expert coding network in the to-be-trained model 300 determines the mth expert coding result for the sample medical record text, the expert decision network in the to-be-trained model 300 can perform decoding processing on the mth expert coding result to obtain the mth expert decision result, so that the mth expert decision result can represent the semantic label information predicted by the mth expert network for the sample medical record text.
[0137] Step 52: Input the M expert decision results and the predicted expert weight values corresponding to the M expert coding networks into the decision fusion network to obtain the predicted semantic label information of the sample medical record text output by the decision fusion network.
[0138] Among them, the "decision fusion network" is used to perform fusion processing on the input data of the decision fusion network; and the embodiment of the present application does not limit the implementation manner of the "decision fusion network". For example, it can be implemented by means of formulas (2)-(3).
[0139]
[0140] P = {p 1 , p 2 , …, p K} (3)
[0141] In the formula, P represents the predicted semantic label information of the sample medical record text; p k represents the possibility that the field information of at least one string in the sample medical record text includes the kth candidate semantic label; t m represents the predicted expert weight value corresponding to the mth expert coding network; L m,k represents the "kth label attribution probability information of the mth expert coding result" above; m is a positive integer, m ≤ M, and M is a positive integer.
[0142] Based on the relevant content of the above steps 51 to 52, for the decision-making layer 304, it can perform comprehensive decision-making processing on the input data of the decision-making layer 304 by means of the expert decision-making network and the decision-making fusion network, which is conducive to improving the decision-making effect.
[0143] Based on the relevant content of the above steps 21 to 24, for Figure 3 the to-be-trained model 300 shown, the to-be-trained model 300 can first perform separate decision-making for each expert aspect by means of each expert encoding network; then refer to all the separate decision-making results for comprehensive decision-making, so that the comprehensive decision-making result can better represent the field information of each string (for example, each sample segment or each text segment) in the input data of the to-be-trained model 300, including the possibility of each candidate semantic label, so as to be able to determine the semantic label determination performance of the to-be-trained model 300 with the help of the comprehensive decision-making result subsequently.
[0144] Method Embodiment 3
[0145] In addition, in order to further improve the training effect of the to-be-trained model, the embodiments of the present application also provide three possible implementation manners for determining the "model prediction loss value of the to-be-trained model", which are introduced below in combination with three cases respectively.
[0146] Case 1. In order to improve the accuracy of the above "model prediction loss value of the to-be-trained model", the "model prediction loss value of the to-be-trained model" can be determined based on the accuracy of the decision-making results determined by each expert encoding network. Based on this, the embodiments of the present application provide the first possible implementation manner for determining the "model prediction loss value of the to-be-trained model", which may specifically include steps 71 to 73:
[0147] Step 71: Determine the m-th expert decision loss value according to the m-th expert decision result and the actual semantic label information of the sample medical record text. Where m is a positive integer and m ≤ M.
[0148] Among them, the "m-th expert decision loss value" is used to represent the gap between the semantic label information predicted for the sample medical record text by means of the m-th expert encoding network and the actual semantic label information of the sample medical record text; moreover, the embodiments of the present application do not limit the determination process of the "m-th expert decision loss value". For example, any existing or future method capable of determining the distance between two pieces of information (such as a calculation method based on similarity, etc.) can be used for implementation.
[0149] In addition, in order to ensure that different expert encoding networks in the model to be trained have different encoding performances, different loss functions can be adopted to measure the prediction performance of the decision results determined by different expert encoding networks. Based on this, an embodiment of the present application also provides another possible implementation manner for determining the "m-th expert decision loss value", which may specifically include: determining the m-th expert decision loss value according to the m-th expert decision result, the actual semantic label information of the sample medical record text, and the network loss function corresponding to the m-th expert encoding network.
[0150] The above-mentioned "network loss function corresponding to the m-th expert encoding network" is used to guide the m-th expert encoding network to perform expert encoding processing on the input data of the m-th expert encoding network according to the m-th expert encoding processing performance; and the "network loss function corresponding to the m-th expert encoding network" can be preset.
[0151] In addition, in order to make the expert encoding processing performances of different expert encoding networks different when performing expert encoding processing, different penalty factors can be used for guidance. Based on this, it can be known that for the above-mentioned "network loss function corresponding to the m-th expert encoding network", the penalty factor in the "network loss function corresponding to the m-th expert encoding network" is different from the penalty factor in the network loss function corresponding to any other expert encoding network except the m-th expert encoding network among the M expert encoding networks. For the convenience of understanding, the following takes Figure 4 the three expert encoding networks shown as an example for illustration.
[0152] As an example, for Figure 4 the three expert encoding networks shown, in order to ensure that "private network 1" is better at processing medical record text data with relatively high semantic parsing difficulty, as shown in formulas (4)-(6), the network loss function corresponding to "private network 1" may include a semantic parsing difficulty penalty factor; in order to ensure that "private network 2" is better at processing medical record text data with relatively low semantic parsing difficulty and relatively rare field names, as shown in formula (7), the network loss function corresponding to "private network 2" may include a label frequency penalty factor; in order to ensure that "private network 3" is better at processing medical record text data with relatively low semantic parsing difficulty and relatively common field names, as shown in formula (9), the network loss function corresponding to "private network 3" may not include any penalty factors.
[0153]
[0154]
[0155]
[0156] In the formula, loss hard represents the expert decision loss value of "private network 1" for the sample medical record text; z 1 represents the serial number of "private network 1", that is, "private network 1" is the z 1 th expert coding network; Y represents the number of strings in the sample medical record text (for example, the number of sample segments or text segments); T g represents the serial number of the semantic label actually included in the field information of the yth string in the sample medical record text (that is, the yth string in the sample medical record text actually includes the T g th candidate semantic label); represents the possibility that the field information of the yth string in the sample medical record text predicted by "private network 1" includes the T g th candidate semantic label; represents the possibility that the field information of the yth string in the sample medical record text predicted by "private network 1" includes the kth candidate semantic label; k is a positive integer, k ≤ K, K is a positive integer, and K represents the number of candidate semantic labels.
[0157] It should be noted that the above and are the semantic parsing difficulty penalty factors; and if the semantic parsing difficulty of a string is higher, then this will be smaller, so that b y,k will also be smaller, thus making larger, and then making smaller, so that in occupies a smaller proportion, resulting in loss hard being larger.
[0158]
[0159]
[0160] In the formula, loss cls represents the expert decision loss value of "private network 2" for the sample medical record text; z 2 represents the serial number of "private network 2", that is, "private network 2" is the z 2 th expert coding network; Y represents the number of strings in the sample medical record text (for example, the number of sample segments or text segments); T g represents the serial number of the semantic label actually included in the field information of the yth string in the sample medical record text; represents the possibility that the field information of the yth string in the sample medical record text predicted by "private network 2" includes the T g th candidate semantic label; Indicates the possibility that the field information of the y-th string in the sample medical record text predicted by means of the "private network 2" includes the k-th candidate semantic label; S k Indicates the frequency characterization data of the k-th candidate semantic label; k is a positive integer, k ≤ K, K is a positive integer, and K represents the number of candidate semantic labels.
[0161] It should be noted that the above and are the label frequency penalty factors; moreover, if the T g -th candidate semantic label is rarer, then the frequency characterization data of the T g -th candidate semantic label is smaller, so that is larger, thereby making smaller, so that loss cls is larger.
[0162] It should also be noted that the determination process of the above "frequency characterization data of the k-th candidate semantic label" is similar to the determination process of the "frequency characterization data of the q-th semantic label" shown below.
[0163]
[0164] In the formula, loss ce represents the expert decision loss value of the "private network 3" for the sample medical record text; z 3 represents the serial number of the "private network 3", that is, the "private network 3" is the z 3 -th expert coding network; Y represents the number of strings in the sample medical record text (for example, the number of sample segments or text segments); T g represents the serial number of the semantic label actually included in the field information of the y-th string in the sample medical record text; represents the possibility that the field information of the y-th string in the sample medical record text predicted by means of the "private network 3" includes the T g -th candidate semantic label; represents the possibility that the field information of the y-th string in the sample medical record text predicted by means of the "private network 3" includes the k-th candidate semantic label; k is a positive integer, k ≤ K, K is a positive integer, and K represents the number of candidate semantic labels.
[0165] Based on the relevant content of step 71 above, after obtaining the m-th expert decision result, the m-th expert decision loss value can be determined according to the difference between the m-th expert decision result and the actual semantic label information of the sample medical record text, so that the m-th expert decision loss value can represent the gap between the semantic label information predicted for the sample medical record text by the m-th expert encoding network and the actual semantic label information of the sample medical record text, thereby enabling the m-th expert decision loss value to represent the expert encoding processing performance of the m-th expert encoding network. Wherein, m is a positive integer and m ≤ M.
[0166] Step 72: Determine the semantic prediction loss value of the sample medical record text according to the M expert decision loss values and the predicted expert weight values corresponding to the M expert encoding networks.
[0167] Among them, the "semantic prediction loss value of the sample medical record text" is used to represent the semantic label determination performance reflected by the model decision information of the sample medical record text.
[0168] In addition, the embodiments of the present application do not limit the implementation manner of step 72. For example, step 72 may specifically include: performing weighted summation on the M expert decision loss values according to the predicted expert weight values corresponding to the M expert encoding networks to obtain the semantic prediction loss value of the sample medical record text. For the convenience of understanding, the following is combined with Figure 4 for illustration.
[0169] As an example, for Figure 4 the three expert encoding networks shown, it can use formula (10) to determine the semantic prediction loss value of the sample medical record text.
[0170] L all =w hard ×loss hard +w cls ×loss cls +w ce ×loss ce (10)
[0171] In the formula, L all represents the semantic prediction loss value of the sample medical record text; loss hard represents the expert decision loss value of "private network 1" for the sample medical record text; w hard represents the predicted expert weight value corresponding to "private network 1"; loss cls represents the expert decision loss value of "private network 2" for the sample medical record text; w cls represents the predicted expert weight value corresponding to "private network 2"; loss ce represents the expert decision loss value of "private network 3" for the sample medical record text; w ceRepresents the predicted expert weight value corresponding to "Private Network 3".
[0172] Step 73: Determine the model prediction loss value of the model to be trained according to the semantic prediction loss value of the sample medical record text.
[0173] It should be noted that the embodiment of the present application does not limit the implementation manner of step 73. For example, when the number of sample medical record texts is G, step 73 may specifically include: directly adding the semantic prediction loss values of G sample medical record texts to obtain the model prediction loss value.
[0174] Based on the relevant content of the above steps 71 to 73, it can be seen that in some cases, the model prediction loss value of the model to be trained can be determined by means of the gap between all expert decision results and the actual semantic label information of the sample medical record text, so that the model prediction loss value can better represent the semantic label determination performance of the model to be trained.
[0175] Case 2. In order to improve the accuracy of the "model prediction loss value of the model to be trained", the appropriate degree of the predicted expert weight value can be referred to to determine the "model prediction loss value of the model to be trained". Based on this, the embodiment of the present application provides a second possible implementation manner for determining the "model prediction loss value of the model to be trained", which may specifically include steps 81 - 83:
[0176] Step 81: Obtain the prior expert weight values corresponding to M expert coding networks.
[0177] Among them, the prior expert weight value corresponding to the m-th expert coding network refers to the actual expert weight value corresponding to the m-th expert coding network. m is a positive integer, m ≤ M, and M is a positive integer.
[0178] In addition, the embodiment of the present application does not limit the determination process of the "prior expert weight value corresponding to the m-th expert coding network" above. For example, it can be set by technical personnel in advance.
[0179] Actually, because the expert coding processing performances of different expert coding networks are different when performing expert coding processing, the decision-making influence degrees of different expert coding networks on different sample medical record texts are different. Based on this, the embodiment of the present application also provides another possible implementation manner for determining the "prior expert weight value corresponding to the m-th expert coding network", which may specifically include steps 91 - 92:
[0180] Step 91: Determine the sample type information of the sample medical record text.
[0181] Among them, the above "sample type information of the sample medical record text" is used to describe the characteristics presented by the semantic tags of the sample medical record text (for example, the difficulty of determining semantic tags, the commonness of field names involved in semantic tags, etc.).
[0182] The embodiments of the present application do not limit the above "sample type information of the sample medical record text". For example, it includes the characterization data of the label determination difficulty of the sample medical record text and / or the characterization data of the label frequency of the sample medical record text.
[0183] The above "characterization data of the label determination difficulty of the sample medical record text" is used to characterize the difficulty of determining the field information of at least one string in the sample medical record text; and the embodiments of the present application do not limit the implementation manner of the "characterization data of the label determination difficulty of the sample medical record text". For example, it may specifically include steps 101 - 104:
[0184] Step 101: Perform a preset partitioning process on the sample medical record text to obtain at least one sample segment.
[0185] It should be noted that for the relevant content of step 101, please refer to the relevant content of step 31 above.
[0186] Step 102: Input each sample segment into a pre - constructed semantic parsing model to obtain the semantic label parsing information of each sample segment output by the semantic parsing model.
[0187] Among them, the "semantic parsing model" is used to perform semantic label determination processing on the input data of the semantic parsing model; and the embodiments of the present application do not limit the "semantic parsing model". For example, it can be implemented using a BiLSTM + CRF model. In addition, the embodiments of the present application do not limit the construction process of the "semantic parsing model", and any existing or future model construction method can be used for implementation.
[0188] The "semantic label parsing information of the sample segment" is used to represent the predicted field information for the sample segment.
[0189] Step 103: Determine the model parsing loss value of the sample medical record text according to the semantic label parsing information of each sample segment and the actual semantic label information of the sample medical record text.
[0190] Among them, the above "model parsing loss value of the sample medical record text" is used to represent the semantic label determination performance of the semantic parsing model for the sample medical record text.
[0191] In addition, the embodiments of the present application do not limit the determination process of the "model parsing loss value of the sample medical record text". For example, when the number of sample segments is F, the determination process of the "model parsing loss value of the sample medical record text" may specifically include: first, determining the model parsing loss value of the f-th sample segment according to the semantic label parsing information of the f-th sample segment and the actual semantic label information of the f-th sample segment; then, determining the average value among the model parsing loss values of the F sample segments as the model parsing loss value of the sample medical record text.
[0192] It should be noted that the embodiments of the present application do not limit the implementation manner of the above "model parsing loss value of the f-th sample segment", and any existing or future method capable of determining the distance between two pieces of information can be used for implementation.
[0193] Step 104: Determine the label determination difficulty characterization data of the sample medical record text according to the model parsing loss value of the sample medical record text.
[0194] It should be noted that the embodiments of the present application do not limit the implementation manner of Step 104. For example, specifically, it may be: determining the model parsing loss value of the sample medical record text as the label determination difficulty characterization data of the sample medical record text.
[0195] Based on the relevant content of the above Steps 101 to 104, in some cases, the label determination difficulty characterization data of a sample medical record text can be predicted by means of a pre-constructed semantic parsing model with a semantic label determination function, so that the label determination difficulty characterization data can represent the semantic label determination difficulty of the sample medical record text.
[0196] The above "label frequency characterization data of the sample medical record text" is used to characterize the common degree of the actual semantic labels of at least one string in the sample medical record text (that is, the occurrence frequency of the actual semantic labels of at least one string in the sample medical record text); moreover, the embodiments of the present application do not limit the implementation manner of the "label frequency characterization data of the sample medical record text". For example, specifically, it may include: performing an occurrence frequency statistical process on the actual semantic label information of the sample medical record text based on the actual semantic label information of at least one reference medical record text to obtain the label frequency characterization data of the sample medical record text. For the sake of easy understanding, the following is illustrated with examples.
[0197] As an example, when the actual semantic tag information of the sample medical record text includes the first semantic tag to the Qth semantic tag, the determination process of the above-mentioned "tag frequency characterization data of the sample medical record text" may include: first, count the occurrence frequency of the qth semantic tag in the actual semantic tag information of at least one reference medical record text to obtain the frequency characterization data of the qth semantic tag; then, determine the average value between the frequency characterization data of the first semantic tag and the frequency characterization data of the Qth semantic tag as the tag frequency characterization data of the sample medical record text.
[0198] It should be noted that the above-mentioned "at least one reference medical record text" can be preset; moreover, the embodiments of the present application do not limit the acquisition method of the "at least one reference medical record text". For example, in order to improve the usage efficiency of medical record texts, G sample medical record texts can be directly determined as the "at least one reference medical record text", so that the tag frequency characterization data of each sample medical record text can be determined based on the G sample medical record texts subsequently.
[0199] Based on the relevant content of step 91 above, after obtaining the sample medical record text, the sample type information of the sample medical record text can be determined, so that the sample type information can represent the characteristics presented by the semantic tags of the sample medical record text (for example, the determination difficulty of semantic tags, the commonness of field names involved in semantic tags, etc.), so that the prior expert weight values corresponding to each expert coding network can be determined based on the sample type information subsequently.
[0200] Step 92: Determine the prior expert weight values corresponding to M expert coding networks according to the sample type information of the sample medical record text and the preset mapping relationship. Among them, the preset mapping relationship includes the corresponding relationship between the sample type information and the prior expert weight values corresponding to M expert coding networks.
[0201] The above-mentioned "preset mapping relationship" is used to record the prior expert weight values corresponding to M expert coding networks corresponding to each sample type information; moreover, the "preset mapping relationship" can be preset.
[0202] In addition, the embodiments of the present application do not limit the above-mentioned "preset mapping relationship". For example, for Figure 4 the three expert coding networks shown, the "preset mapping relationship" may include the corresponding relationship between the first candidate type information and the first prior weight set, the corresponding relationship between the second candidate type information and the second prior weight set, the corresponding relationship between the third candidate type information and the third prior weight set, and the corresponding relationship between the fourth candidate type information and the fourth prior weight set.
[0203] The above-mentioned "first candidate type information" includes first label determination difficulty characterization data and first label frequency characterization data. The first label determination difficulty characterization data reaches a difficulty threshold, and the first label frequency characterization data is higher than a frequency threshold, so that the "first candidate type information" is used to represent the medical record text data of "difficult to classify - high frequency". Among them, both the "frequency threshold" and the "frequency threshold" can be preset in advance.
[0204] In the above-mentioned "first prior weight set", the prior expert weight value corresponding to "private network 1" is higher than the prior expert weight values corresponding to the other two expert coding networks; and the embodiments of the present application do not limit the "first prior weight set". For example, it can be [0.8, 0.1, 0.1]; and "0.8" refers to the prior expert weight value corresponding to "private network 1"; "0.1" refers to the prior expert weight value corresponding to "private network 2"; "0.1" refers to the prior expert weight value corresponding to "private network 3". It can be seen that for the medical record text data of "difficult to classify - high frequency", the decision-making is mainly carried out by means of "private network 1".
[0205] The above-mentioned "second candidate type information" includes second label determination difficulty characterization data and second label frequency characterization data. The second label determination difficulty characterization data reaches a difficulty threshold, and the second label frequency characterization data is not higher than a frequency threshold, so that the "second candidate type information" is used to represent the medical record text data of "difficult to classify - low frequency".
[0206] In the above-mentioned "second prior weight set", the prior expert weight values corresponding to "private network 1" and "private network 2" are both higher than the prior expert weight value corresponding to "private network 3"; and the embodiments of the present application do not limit the "second prior weight set". For example, it can be [0.45, 0.45, 0.1]. Among them, "0.45" refers to the prior expert weight value corresponding to "private network 1"; "0.45" refers to the prior expert weight value corresponding to "private network 2"; "0.1" refers to the prior expert weight value corresponding to "private network 3". It can be seen that for the medical record text data of "difficult to classify - low frequency", the decision-making is mainly carried out jointly by means of "private network 1" and "private network 2".
[0207] The above-mentioned "third candidate type information" includes third label determination difficulty characterization data and third label frequency characterization data. The third label determination difficulty characterization data is lower than a difficulty threshold, and the third label frequency characterization data is higher than a frequency threshold, so that the "third candidate type information" is used to represent the medical record text data of "easy to classify - high frequency".
[0208] Among the above-mentioned "third set of prior weights", the prior expert weight value corresponding to "private network 3" is higher than the prior expert weight values corresponding to the other two expert coding networks; moreover, the embodiments of the present application do not limit this "third set of prior weights". For example, it can be [0.2, 0.2, 0.6]. Among them, "0.2" refers to the prior expert weight value corresponding to "private network 1"; "0.2" refers to the prior expert weight value corresponding to "private network 2"; "0.6" refers to the prior expert weight value corresponding to "private network 3". It can be seen that for the medical record text data of "easy classification - high frequency", the decision-making is mainly carried out by means of "private network 3".
[0209] The above-mentioned "fourth candidate type information" includes fourth label determination difficulty characterization data and fourth label frequency characterization data. The fourth label determination difficulty characterization data is lower than the difficulty threshold, and the fourth label frequency characterization data is not higher than the frequency threshold, so that the "fourth candidate type information" is used to represent the medical record text data of "easy classification - low frequency".
[0210] Among the above-mentioned "fourth set of prior weights", the prior expert weight value corresponding to "private network 2" is higher than the prior expert weight values corresponding to the other two expert coding networks; moreover, the embodiments of the present application do not limit this "fourth set of prior weights". For example, it can be [0.1, 0.8, 0.1]. Among them, "0.1" refers to the prior expert weight value corresponding to "private network 1"; "0.8" refers to the prior expert weight value corresponding to "private network 2"; "0.1" refers to the prior expert weight value corresponding to "private network 3". It can be seen that for the medical record text data of "difficult classification - low frequency", the decision-making is mainly carried out by means of "private network 2".
[0211] Based on the relevant content of step 92 above, after obtaining the sample type information of the sample medical record text, the prior weight set corresponding to the sample type information can be found from the pre-constructed preset mapping relationship, and the prior expert weight values corresponding to the M expert coding networks can be obtained, so that these prior expert weight values can better guide the training and updating process of the model to be trained.
[0212] Based on the relevant content of step 81 above, after obtaining the sample medical record text, the prior knowledge of the expert weights that need to be learned when the model to be trained performs semantic label determination processing on the sample medical record text can be determined according to the sample medical record text, so that the model to be trained can be better trained and updated based on this prior knowledge of the expert weights.
[0213] Step 82: Determine the semantic prediction loss value of the sample medical record text according to the predicted semantic label information and the actual semantic label information of the sample medical record text.
[0214] Among them, the "semantic prediction loss value of the sample medical record text" is used to represent the semantic label determination performance of the model to be trained reflected by the model decision information of the sample medical record text.
[0215] In addition, the embodiments of the present application do not limit the implementation manner of step 82. For example, any existing or future method capable of determining the prediction loss by referring to the prediction information and the actual information can be adopted for implementation.
[0216] Step 83: Determine the weight prediction loss value of the sample medical record text according to the prior expert weight values corresponding to the M expert coding networks and the prediction expert weight values corresponding to the M expert coding networks.
[0217] Among them, the "weight prediction loss value" is used to represent the semantic label determination performance reflected by the weight prediction information of the model to be trained.
[0218] In addition, the embodiments of the present application do not limit the implementation manner of step 83. For example, the difference between the prior expert weight values corresponding to the M expert coding networks and the prediction expert weight values corresponding to the M expert coding networks can be determined as the weight prediction loss value of the sample medical record text. It should be noted that the embodiments of the present application do not limit the determination process of the "difference". For example, the Euclidean distance, etc. can be adopted for implementation.
[0219] Step 84: Determine the model prediction loss value of the model to be trained according to the semantic prediction loss value of the sample medical record text and the weight prediction loss value of the sample medical record text.
[0220] It should be noted that the embodiments of the present application do not limit the implementation manner of step 84. For example, when the number of sample medical record texts is G, step 84 may specifically include: first, determining the sum value between the semantic prediction loss value of the g-th sample medical record text and the weight prediction loss value of the g-th sample medical record text as the sample prediction loss value of the g-th sample medical record text; then, performing statistical analysis processing on the sample prediction loss values of the G sample medical record texts to obtain the model prediction loss value of the model to be trained. Among them, the "statistical analysis processing" can be pre-processed. For example, it can be taking the average value, taking the sum value, etc.
[0221] Based on the relevant content of the above steps 81 to 84, it can be seen that in some cases, the model prediction loss value of the model to be trained can be determined by means of the difference between the prior expert weight values corresponding to all expert coding networks and the weight prediction loss value, so that the model prediction loss value can better represent the semantic label determination performance of the model to be trained.
[0222] In Case 3, to improve the accuracy of the "model prediction loss value of the model to be trained", the accuracy of each decision result determined by the expert coding network and the appropriateness of the predicted expert weight value can be referred to simultaneously to determine the "model prediction loss value of the model to be trained". Based on this, the embodiments of the present application provide a third possible implementation manner for determining the "model prediction loss value of the model to be trained", which may specifically include Step 111 - Step 113:
[0223] Step 111: Determine the m-th expert decision loss value according to the m-th expert decision result and the actual semantic label information of the sample medical record text. Here, m is a positive integer, and m ≤ M.
[0224] Step 112: Determine the semantic prediction loss value of the sample medical record text according to the M expert decision loss values and the predicted expert weight values corresponding to the M expert coding networks.
[0225] Step 113: Determine the sample type information of the sample medical record text.
[0226] Step 114: Determine the weight prediction loss value of the sample medical record text according to the prior expert weight values corresponding to the M expert coding networks and the predicted expert weight values corresponding to the M expert coding networks.
[0227] Step 115: Determine the model prediction loss value of the model to be trained according to the semantic prediction loss value of the sample medical record text and the weight prediction loss value of the sample medical record text.
[0228] It should be noted that for the relevant content of Step 111 - Step 115, please refer to Step 71, Step 72, Step 81, Step 83, and Step 84 above respectively.
[0229] Based on the relevant content of the above Step 111 to Step 115, it can be seen that in some cases, the model prediction loss value of the model to be trained can be determined by means of the gap between all expert decision results and the actual semantic label information of the sample medical record text, and the gap between the prior expert weight values corresponding to all expert coding networks and the weight prediction loss value, so that the model prediction loss value can better represent the semantic label determination performance of the model to be trained.
[0230] Method Embodiment 4
[0231] To further improve the training effect of the model to be trained, the embodiments of the present application also provide three possible implementation manners for updating the model to be trained (that is, S203), which are introduced below in combination with three cases respectively.
[0232] In Case 1, to improve the learning effect of the model to be trained, the differences between the decision results determined by each expert coding network and the actual semantic label information can be referred to update the model to be trained. Based on this, the first possible implementation manner for updating the model to be trained provided by the embodiments of the present application may specifically include: updating the model to be trained according to M expert decision results, the predicted expert weight values corresponding to the M expert coding networks, and the actual semantic label information of the sample medical record text.
[0233] It can be seen that in some cases, the semantic prediction loss value of the sample medical record text can be determined first according to M expert decision results, the predicted expert weight values corresponding to the M expert coding networks, and the actual semantic label information of the sample medical record text (such as the relevant content shown in the above steps 71 - step 72); then, based on the semantic prediction loss value of the sample medical record text, the model to be trained can be updated.
[0234] It should be noted that the above "updating the model to be trained according to the semantic prediction loss value of the sample medical record text" may specifically include: first determining the model prediction loss value of the model to be trained according to the semantic prediction loss value of the sample medical record text (such as the relevant content shown in the above step 73); then using the model prediction loss value of the model to be trained to perform backpropagation update processing on the model to be trained to obtain the updated model to be trained.
[0235] In Case 2, to improve the learning effect of the model to be trained, the differences between the predicted expert weight values and the prior expert weight values can be further referred to update the model to be trained. Based on this, the second possible implementation manner for updating the model to be trained provided by the embodiments of the present application may specifically include: updating the model to be trained according to the predicted semantic label information of the sample medical record text, the actual semantic label information of the sample medical record text, the prior expert weight values corresponding to the M expert coding networks, and the predicted expert weight values corresponding to the M expert coding networks.
[0236] It can be seen that in some cases, the model prediction loss value of the model to be trained can be determined first according to the predicted semantic label information of the sample medical record text, the actual semantic label information of the sample medical record text, the prior expert weight values corresponding to the M expert coding networks, and the predicted expert weight values corresponding to the M expert coding networks (such as the relevant content shown in the above steps 82 - 84); then using the model prediction loss value of the model to be trained to perform backpropagation update processing on the model to be trained to obtain the updated model to be trained, so that the updated model to be trained has better semantic label determination performance.
[0237] In Case 3, to improve the learning effect of the model to be trained, the differences between the decision results determined by each expert coding network and the actual semantic label information, as well as the differences between the predicted expert weight values and the prior expert weight values, can be referred to simultaneously to update the model to be trained. Based on this, the third possible implementation manner for updating the model to be trained is provided in an embodiment of the present application, which may specifically include: updating the model to be trained according to the M expert decision results of the sample medical record text, the actual semantic label information of the sample medical record text, the prior expert weight values corresponding to the M expert coding networks, and the predicted expert weight values corresponding to the M expert coding networks.
[0238] It can be seen that in some cases, the model prediction loss value of the model to be trained can be determined first according to the M expert decision results of the sample medical record text, the actual semantic label information of the sample medical record text, the prior expert weight values corresponding to the M expert coding networks, and the predicted expert weight values corresponding to the M expert coding networks (such as the relevant content shown in the above steps 111, 112, 114, and 115); then, the model prediction loss value of the model to be trained is used to perform backpropagation update processing on the model to be trained to obtain the updated model to be trained, so that the updated model to be trained has better semantic label determination performance.
[0239] In addition, based on the relevant content of the semantic label determination model constructed in the above method embodiment, an embodiment of the present application also provides a medical record parsing method, which will be described below with reference to the accompanying drawings.
[0240] Method Embodiment 5
[0241] See Figure 6 , which is a flowchart of a medical record parsing method provided in an embodiment of the present application.
[0242] The medical record parsing method provided in an embodiment of the present application includes S601 - S603:
[0243] S601: Obtain the medical record text to be processed.
[0244] Among them, the "medical record text to be processed" refers to the medical record text data that needs to be parsed; and the embodiment of the present application does not limit the "medical record text to be processed". For example, it may include a section of medical record content extracted from an electronic medical record.
[0245] S602: Determine the predicted semantic label information of the medical record text to be processed according to the medical record text to be processed and the semantic label determination model.
[0246] Among them, the "semantic label determination model" is constructed by any implementation manner of the method for constructing the semantic label determination model provided in the embodiments of the present application; for the relevant content of the "semantic label determination model", please refer to the relevant content of the "semantic label determination model" above.
[0247] The above "predicted semantic label information of the medical record text to be processed" is used to represent the field information predicted for at least one string in the medical record text to be processed; and the determination process of the "predicted semantic label information of the medical record text to be processed" is similar to the determination process of the "predicted semantic label information of the sample medical record text" above, and only needs to replace the "sample medical record text" in any implementation manner of the determination process of the "predicted semantic label information of the sample medical record text" above with the "medical record text to be processed", and the "model to be trained" with the "semantic label determination model".
[0248] S603: Determine the semantic parsing result of the medical record text to be processed according to the predicted semantic label information of the medical record text to be processed.
[0249] In the embodiments of the present application, after obtaining the predicted semantic label information of the medical record text to be processed, the semantic parsing result of the medical record text to be processed can be determined according to each field information carried by the predicted semantic label information and the string corresponding to each field information in the medical record text to be processed, so that the semantic parsing result can represent the semantic information carried by the medical record text to be processed (for example, semantic information such as symptom manifestation is..., drug history is...).
[0250] Actually, since the determination process of the above "predicted semantic label information of the medical record text to be processed" involves text division processing, there may be a phenomenon that at least two adjacent strings in the medical record text to be processed have the same field information. Therefore, in order to improve the semantic parsing effect, another possible implementation manner of S603 is provided in the embodiments of the present application. In this implementation manner, when the above "predicted semantic label information of the medical record text to be processed" includes the fragment semantic label information of at least one fragment to be processed in the medical record text to be processed, S603 may specifically include step 121-step 122:
[0251] Step 121: Integrate the fragment semantic label information of at least one fragment to be processed according to a preset integration rule to obtain the fragment semantic label information of at least one fragment to be used.
[0252] Among them, the "fragment to be processed" refers to the fragment obtained by dividing the medical record text to be processed during the determination process of the above "predicted semantic label information of the medical record text to be processed"; and the embodiments of the present application do not limit the number of fragments to be processed. For example, it may specifically be R.
[0253] The segment semantic tag information of the r-th segment to be processed is used to represent the field information predicted for the r-th segment to be processed.
[0254] The above-mentioned "preset integration rule" refers to a rule for merging semantic tag information that is preset in advance; moreover, the embodiments of the present application do not limit this "preset integration rule". For example, specifically, it may include: for any two adjacent segments to be processed, if the segment semantic tag information of these two segments to be processed is the same, then these two segments to be processed can be first merged to obtain a merged segment corresponding to these two segments to be processed; then, the segment semantic tag information common to these two segments to be processed is determined as the segment semantic tag information of the merged segment corresponding to these two segments to be processed.
[0255] Step 122: Determine the semantic parsing result of the medical record text to be processed according to the segment semantic tag information of at least one segment to be used.
[0256] In the embodiments of the present application, after obtaining the segment semantic tag information of at least one segment to be used, the semantic parsing result of the medical record text to be processed can be determined according to each segment to be used and the semantic tag information of each segment to be used, so that the semantic parsing result can represent the semantic information carried by the medical record text to be processed in a more concise manner, which is beneficial to improving the semantic parsing effect for medical record data.
[0257] Based on the relevant content of the above steps 121 to 122, after obtaining the predicted semantic tag information of the medical record text to be processed, the predicted semantic tag information can be integrated to obtain the integrated predicted semantic tag information, so that the integrated predicted semantic tag information can accurately represent the field information of each string in the medical record text to be processed with as few semantic tags as possible, so that the semantic parsing result of the medical record text to be processed determined based on the integrated predicted semantic tag information can represent the semantic information carried by the medical record text to be processed in a more concise manner, which is beneficial to improving the semantic parsing effect for medical record data.
[0258] In some cases, it is more likely that multiple consecutive strings in a medical record text to be processed have the same character information. Therefore, in order to improve the semantic parsing effect, another possible implementation manner of S603 is also provided in the embodiments of the present application, which may specifically include steps 131 - 132:
[0259] Step 131: Correct the semantic tag information to be corrected according to the preset error correction rule to obtain the corrected semantic tag.
[0260] Among them, the "semantic label information to be corrected" refers to the semantic label information that needs to be corrected; moreover, the embodiments of the present application do not limit the "semantic label information to be corrected". For example, it can be the "predicted semantic label information of the medical record text to be processed" mentioned above, or the "fragment semantic label information of at least one fragment to be used" mentioned above.
[0261] The "preset error correction rule" can be preset in advance; moreover, the embodiments of the present application do not limit the "preset error correction rule". For example, when the above-mentioned "semantic label information to be corrected" includes the fragment semantic label information of at least one string in the medical record text to be processed, the "preset error correction rule" can specifically include: for any number of consecutively located strings, if the fragment semantic label information of a non-edge string among these strings is different from the fragment semantic label information of the other strings except this non-edge string among these strings, and the fragment semantic label information of all the other strings except this non-edge string among these strings is the same, then it can be determined that the fragment semantic label information of this non-edge string is incorrect. Therefore, the fragment semantic label information of this non-edge string can be modified to "the fragment semantic label information of all the other strings except this non-edge string among these strings". Among them, the "non-edge string" refers to any string other than the string with the earliest position and the string with the latest position among these strings.
[0262] Step 132: Determine the semantic parsing result of the medical record text to be processed according to the corrected semantic label.
[0263] To facilitate the understanding of Step 132, the following will be described with three examples.
[0264] Example 1, when the above-mentioned "semantic label information to be corrected" is the "predicted semantic label information of the medical record text to be processed" mentioned above, Step 132 can specifically include: directly determining the corrected semantic label as the semantic parsing result of the medical record text to be processed.
[0265] Example 2, when the above-mentioned "semantic label information to be corrected" is the "predicted semantic label information of the medical record text to be processed" mentioned above, and the corrected semantic label includes the semantic label correction information of at least one fragment to be processed in the medical record text to be processed, Step 132 can specifically include: first, according to the preset integration rule, perform integration processing on the semantic label correction information of at least one fragment to be processed to obtain the fragment semantic label information of at least one fragment to be used; then, according to the fragment semantic label information of at least one fragment to be used, determine the semantic parsing result of the medical record text to be processed.
[0266] Example 3. When the above "semantic label information to be corrected" is the "fragment semantic label information of at least one fragment to be used" in the above text, step 132 may specifically include: directly determining the corrected semantic label as the semantic parsing result of the medical record text to be processed.
[0267] Based on the relevant content of the above steps 131 to 132, it can be seen that after obtaining the predicted semantic label information of the medical record text to be processed, integration processing and error correction processing can be performed on the predicted semantic label information to obtain the processed predicted semantic label information, so that the processed predicted semantic label information can accurately represent the field information of each string in the medical record text to be processed with as few semantic labels as possible. Thus, the semantic parsing result of the medical record text to be processed determined based on the processed predicted semantic label information can more accurately represent the semantic information carried by the medical record text to be processed, which is beneficial to improving the semantic parsing effect of medical record data.
[0268] Based on the relevant content of the above S601 to S603, it can be seen that after obtaining the medical record text to be processed, the pre-constructed semantic label determination model can be used to perform semantic label determination processing on the medical record text to be processed to obtain the predicted semantic label information of the medical record text to be processed; then, referring to the predicted semantic label information, the semantic parsing result of the medical record text to be processed can be determined, so that the semantic parsing result can accurately represent the semantic information carried by the medical record text to be processed, which is beneficial to improving the semantic parsing effect of medical record data.
[0269] In addition, in some application scenarios (for example, scenarios such as determining whether an electronic medical record is written in a standard manner), it is necessary to determine whether there is a field missing phenomenon in the above medical record text to be processed. Based on this, another possible implementation manner of the medical record parsing method is provided in the embodiments of the present application. In this implementation manner, in addition to including the above S601 - S603, the medical record parsing method may further include S604:
[0270] S604: Determine whether there is a field missing phenomenon in the medical record text to be processed according to the semantic parsing result of the medical record text to be processed and the preset complete condition of medical record fields.
[0271] Among them, the "preset complete condition of medical record fields" refers to a judgment condition preset for determining whether there is a field missing phenomenon in a medical record content; and the embodiments of the present application do not limit the "preset complete condition of medical record fields". For example, it may specifically include: covering at least one specific field name. It should be noted that the above "at least one specific field name" can be determined according to the above "preset medical record specification data" (for example, "Medical Record Writing Specification 2014 Edition").
[0272] It can be seen that when the "preset medical record field completeness condition" includes covering at least one specific field name, if it is determined that the semantic parsing result of the medical record text to be processed completely covers the above "at least one specific field name", it can be determined that the semantic parsing result of the medical record text to be processed meets the preset medical record field completeness condition, so it can be determined that there is no field missing phenomenon in the medical record text to be processed; if it is determined that the semantic parsing result of the medical record text to be processed fails to cover the above "at least one specific field name", it can be determined that the semantic parsing result of the medical record text to be processed does not meet the preset medical record field completeness condition, so it can be determined that there is a field missing phenomenon in the medical record text to be processed.
[0273] Based on the relevant content of S604 above, it can be known that after obtaining the semantic parsing result of the medical record text to be processed, it can be determined whether the semantic parsing result of the medical record text to be processed meets the preset medical record field completeness condition. If it meets, it can be determined that the medical record text to be processed meets the field completeness requirement; if it does not meet, it can be determined that there is a field missing phenomenon in the medical record text to be processed, so that the identification and processing of field missing in medical record text data can be realized.
[0274] Based on the method for constructing a semantic label determination model provided in the above method embodiments, an embodiment of the present application further provides a device for constructing a semantic label determination model, which will be explained and described below with reference to the accompanying drawings.
[0275] Device Embodiment 1
[0276] Embodiment 1 of the device introduces the device for constructing a semantic label determination model. For relevant content, please refer to the above method embodiments.
[0277] See Figure 7 , which is a schematic structural diagram of a device for constructing a semantic label determination model provided in an embodiment of the present application.
[0278] The device 700 for constructing a semantic label determination model provided in an embodiment of the present application includes:
[0279] A first acquisition unit 701, configured to acquire a sample medical record text and the actual semantic label information of the sample medical record text;
[0280] A first determination unit 702, configured to determine the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained;
[0281] A model update unit 703, configured to update the model to be trained according to the predicted semantic label information and the actual semantic label information, and return to the first determination unit 702 to continue to execute the step of obtaining the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained, until when a preset stop condition is reached, a semantic label determination model is determined according to the model to be trained.
[0282] In a possible implementation manner, the model to be trained includes a text encoding layer, an expert encoding layer, an expert weight determination layer, and a decision-making layer; wherein, the expert encoding layer includes M expert encoding networks; M is a positive integer;
[0283] The first determination unit 702 includes:
[0284] A first determination subunit, configured to obtain a text encoding result to be used according to the sample medical record text and the text encoding layer;
[0285] A second determination subunit, configured to input the text encoding result to be used into the m-th expert encoding network to obtain the m-th expert encoding result output by the m-th expert encoding network; wherein, m is a positive integer, and m ≤ M;
[0286] A third determination subunit, configured to input the text encoding result to be used into the expert weight determination layer to obtain the predicted expert weight values corresponding to the M expert encoding networks output by the expert weight determination layer;
[0287] A fourth determination subunit, configured to input the M expert encoding results and the predicted expert weight values corresponding to the M expert encoding networks into the decision-making layer to obtain the predicted semantic label information of the sample medical record text output by the decision-making layer.
[0288] In a possible implementation manner, the apparatus 700 for constructing the semantic label determination model further includes:
[0289] A third acquisition unit, configured to acquire the prior expert weight values corresponding to the M expert encoding networks;
[0290] The model update unit 703 includes:
[0291] A first update subunit, configured to update the model to be trained according to the predicted semantic label information, the actual semantic label information, the prior expert weight values corresponding to the M expert encoding networks, and the predicted expert weight values corresponding to the M expert encoding networks.
[0292] In a possible implementation manner, the third acquisition unit includes:
[0293] A fifth determination subunit, configured to determine the sample type information of the sample medical record text;
[0294] A sixth determination subunit, configured to determine the prior expert weight values corresponding to the M expert coding networks according to the sample type information and a preset mapping relationship; wherein, the preset mapping relationship includes the corresponding relationship between the sample type information and the prior expert weight values corresponding to the M expert coding networks.
[0295] In a possible implementation manner, the sample type information includes label determination difficulty characterization data, and the third acquisition unit includes: a seventh determination subunit;
[0296] Or,
[0297] The sample type information includes label frequency characterization data; the third acquisition unit includes: an eighth determination subunit;
[0298] Or,
[0299] The sample type information includes label determination difficulty characterization data and label frequency characterization data, and the third acquisition unit includes: a seventh determination subunit and an eighth determination subunit;
[0300] The seventh determination subunit is configured to perform a preset partitioning process on the sample medical record text to obtain at least one sample segment; input each of the sample segments into a pre-constructed semantic parsing model to obtain semantic label parsing information of each of the sample segments output by the semantic parsing model; determine a model parsing loss value of the sample medical record text according to the semantic label parsing information of each of the sample segments and the actual semantic label information; determine the label determination difficulty characterization data of the sample medical record text according to the model parsing loss value;
[0301] The eighth determination subunit is configured to perform a frequency statistics process on the actual semantic label information based on the actual semantic label information of at least one reference medical record text to obtain the label frequency characterization data of the sample medical record text.
[0302] In a possible implementation manner, the first update subunit is specifically configured to: determine a semantic prediction loss value of the sample medical record text according to the predicted semantic label information and the actual semantic label information; determine a weight prediction loss value of the sample medical record text according to the prior expert weight values corresponding to the M expert coding networks and the predicted expert weight values corresponding to the M expert coding networks; determine a model prediction loss value of the to-be-trained model according to the semantic prediction loss value of the sample medical record text and the weight prediction loss value of the sample medical record text; update the to-be-trained model according to the model prediction loss value.
[0303] In a possible implementation manner, the decision-making layer includes an expert decision-making network and a decision-making fusion network; the fourth determination subunit includes:
[0304] The ninth determination subunit is configured to input the m-th expert encoding result into the expert decision-making network to obtain the m-th expert decision-making result output by the expert decision-making network; where m is a positive integer and m ≤ M;
[0305] The tenth determination subunit is configured to input M expert decision-making results and the predicted expert weight values corresponding to the M expert encoding networks into the decision-making fusion network to obtain the predicted semantic label information of the sample medical record text output by the decision-making fusion network.
[0306] In a possible implementation manner, the model updating unit 703 includes:
[0307] The second updating subunit is configured to update the model to be trained according to the M expert decision-making results, the predicted expert weight values corresponding to the M expert encoding networks, and the actual semantic label information.
[0308] In a possible implementation manner, the second updating unit includes:
[0309] The eleventh determination subunit is configured to determine the m-th expert decision-making loss value according to the m-th expert decision-making result and the actual semantic label information; where m is a positive integer and m ≤ M;
[0310] The twelfth determination subunit is configured to determine the semantic prediction loss value of the sample medical record text according to the M expert decision-making loss values and the predicted expert weight values corresponding to the M expert encoding networks;
[0311] The third updating subunit is configured to update the model to be trained according to the semantic prediction loss value of the sample medical record text.
[0312] In a possible implementation manner, the eleventh determination subunit is specifically configured to: determine the m-th expert decision-making loss value according to the m-th expert decision-making result, the actual semantic label information, and the network loss function corresponding to the m-th expert encoding network; where the penalty factor in the network loss function corresponding to the m-th expert encoding network is different from the penalty factors in the network loss functions corresponding to any other expert encoding network among the M expert encoding networks except the m-th expert encoding network.
[0313] In a possible implementation manner, the actual semantic label information of the sample medical record text is determined from K candidate semantic labels; where K is a positive integer;
[0314] The ninth determination subunit is specifically configured to: perform a vector dot product operation on the m-th expert encoding result and the character feature vector of the k-th candidate semantic label to obtain the k-th label attribution probability information of the m-th expert encoding result; where k is a positive integer and k ≤ K; perform a set process on the 1st to K-th label attribution probability information of the m-th expert encoding result to obtain the m-th expert decision result.
[0315] In a possible implementation manner, the model to be trained further includes a text segmentation layer;
[0316] The first determination subunit includes:
[0317] The thirteenth determination subunit is configured to input the sample medical record text into the text segmentation layer to obtain at least one text segment output by the text segmentation layer;
[0318] The fourteenth determination subunit is configured to input the at least one text segment into the text encoding layer to obtain the text encoding result to be used output by the text encoding layer; where the text encoding result to be used includes the text encoding results of the at least one text segment.
[0319] In a possible implementation manner, the text encoding layer includes a first encoding network and a second encoding network; the fourteenth determination subunit is specifically configured to: input each of the text segments into the first encoding network to obtain the preliminary encoding results of each of the text segments output by the first encoding network; input the preliminary encoding results of the at least one text segment into the second encoding network to obtain the text encoding results of the at least one text segment output by the second encoding network.
[0320] In a possible implementation manner, the first determination unit 702 is specifically configured to: perform a preset partitioning process on the sample medical record text to obtain at least one sample segment; input the at least one sample segment into the model to be trained to obtain the predicted semantic label information of the sample medical record text output by the model to be trained; where the predicted semantic label information includes the segment semantic label information of the at least one sample segment.
[0321] Based on the medical record parsing method provided in the above method embodiment, an embodiment of the present application further provides a medical record parsing device, which will be explained and described below with reference to the accompanying drawings.
[0322] Device Embodiment 2
[0323] Embodiment 2 of the device introduces the medical record parsing device, and for related content, please refer to the above method embodiment.
[0324] See Figure 8 , which is a schematic structural diagram of a medical record parsing device provided by an embodiment of the present application.
[0325] The medical record parsing device 800 provided by the embodiment of the present application includes:
[0326] A second acquisition unit 801, configured to acquire a medical record text to be processed;
[0327] A second determination unit 802, configured to determine predicted semantic label information of the medical record text to be processed according to the medical record text to be processed and a semantic label determination model, where the semantic label determination model is constructed by using the construction method of the semantic label determination model according to any one of claims 1 to 14;
[0328] A third determination unit 803, configured to determine a semantic parsing result of the medical record text to be processed according to the predicted semantic label information of the medical record text to be processed.
[0329] In a possible implementation manner, the predicted semantic label information of the medical record text to be processed includes fragment semantic label information of at least one to-be-processed fragment in the medical record text to be processed;
[0330] The third determination unit 803 is specifically configured to: perform an integration process on the fragment semantic label information of the at least one to-be-processed fragment according to a preset integration rule to obtain fragment semantic label information of at least one to-be-used fragment;
[0331] Determine a semantic parsing result of the medical record text to be processed according to the fragment semantic label information of the at least one to-be-used fragment.
[0332] In a possible implementation manner, the process of determining the semantic parsing result includes: performing an error correction process on the semantic label information to be corrected according to a preset error correction rule to obtain an error-corrected semantic label; determining a semantic parsing result of the medical record text to be processed according to the error-corrected semantic label.
[0333] In a possible implementation manner, the medical record parsing device 800 further includes:
[0334] A specification determination unit, configured to determine whether there is a field missing phenomenon in the medical record text to be processed according to the semantic parsing result of the medical record text to be processed and a preset complete condition of medical record fields.
[0335] Furthermore, an embodiment of the present application further provides a device, including: a processor, a memory, and a system bus;
[0336] The processor and the memory are connected through the system bus;
[0337] The memory is used to store one or more programs, and the one or more programs include instructions which, when executed by the processor, cause the processor to execute any implementation manner of the method for constructing a semantic tag determination model provided in the embodiments of the present application, or execute any implementation manner of the medical record parsing method provided in the embodiments of the present application.
[0338] Further, an embodiment of the present application further provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions run on a terminal device, the terminal device is caused to execute any implementation manner of the method for constructing a semantic tag determination model provided in the embodiments of the present application, or execute any implementation manner of the medical record parsing method provided in the embodiments of the present application.
[0339] Further, an embodiment of the present application further provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute any implementation manner of the method for constructing a semantic tag determination model provided in the embodiments of the present application, or execute any implementation manner of the medical record parsing method provided in the embodiments of the present application.
[0340] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0341] It should be noted that the various embodiments in this specification are described in a progressive manner, and the key point of each embodiment is to describe the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0342] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0343] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for constructing a semantic label determination model, characterized in that, the method includes: Obtaining sample medical record texts and the actual semantic label information of the sample medical record texts; Determining the predicted semantic label information of the sample medical record texts according to the sample medical record texts and the model to be trained; the model to be trained includes an expert coding layer and an expert weight determination layer, the expert coding layer includes a plurality of expert coding networks, and the output data of the expert weight determination layer includes the predicted expert weight values corresponding to the plurality of expert coding networks; Updating the model to be trained according to the predicted semantic label information and the actual semantic label information, and continuing to execute the step of obtaining the predicted semantic label information of the sample medical record texts according to the sample medical record texts and the model to be trained, until when a preset stop condition is reached, determining a semantic label determination model according to the model to be trained, and different expert coding networks in the semantic label determination model are good at processing different semantic parsing directions during expert coding processing; The updating the model to be trained according to the predicted semantic label information and the actual semantic label information includes: Updating the model to be trained according to the predicted semantic label information, the actual semantic label information, the prior expert weight values corresponding to the plurality of expert coding networks, and the predicted expert weight values corresponding to the plurality of expert coding networks; wherein, the prior expert weight values are obtained by looking up the prior weight set corresponding to the sample type information of the sample medical record texts from a preset mapping relationship, and the sample type information includes label determination difficulty characterization data and label frequency characterization data; the preset mapping relationship includes the corresponding relationship between candidate type information and the prior weight set, the label determination difficulty characterization data in the first candidate type information reaches the difficulty threshold, the label frequency characterization data in the first candidate type information is higher than the frequency threshold, the label determination difficulty characterization data in the second candidate type information reaches the difficulty threshold, the label frequency characterization data in the second candidate type information is not higher than the frequency threshold, the label determination difficulty characterization data in the third candidate type information is lower than the difficulty threshold, the label frequency characterization data in the third candidate type information is higher than the frequency threshold, and the label determination difficulty characterization data in the fourth candidate type information is lower than the difficulty threshold, and the label frequency characterization data in the fourth candidate type information is not higher than the frequency threshold.
2. The method according to claim 1, characterized in that, the model to be trained further includes a text coding layer and a decision layer; the plurality of expert coding networks includes M expert coding networks; M is a positive integer; The process of determining the predicted semantic label information includes: Determining the text coding result to be used according to the sample medical record texts and the text coding layer; Inputting the text coding result to be used into the mth expert coding network to obtain the mth expert coding result output by the mth expert coding network; wherein, m is a positive integer and m≤M; Input the encoded result of the text to be used into the expert weight determination layer to obtain the predicted expert weight values corresponding to the M expert encoding networks output by the expert weight determination layer; Input the M expert encoding results and the predicted expert weight values corresponding to the M expert encoding networks into the decision layer to obtain the predicted semantic label information of the sample medical record text output by the decision layer.
3. The method according to claim 1, characterized in that, the process of determining the label determination difficulty characterization data includes: Performing a preset partitioning process on the sample medical record text to obtain at least one sample segment; inputting each sample segment into a pre-constructed semantic parsing model to obtain the semantic label parsing information of each sample segment output by the semantic parsing model; determining the model parsing loss value of the sample medical record text according to the semantic label parsing information of each sample segment and the actual semantic label information; determining the label determination difficulty characterization data of the sample medical record text according to the model parsing loss value; the process of determining the label frequency characterization data includes: Based on the actual semantic label information of at least one reference medical record text, performing an occurrence frequency statistical process on the actual semantic label information to obtain the label frequency characterization data of the sample medical record text.
4. The method according to claim 1, characterized in that, updating the model to be trained according to the predicted semantic label information, the actual semantic label information, the prior expert weight values corresponding to the multiple expert encoding networks, and the predicted expert weight values corresponding to the multiple expert encoding networks includes: Determining the semantic prediction loss value of the sample medical record text according to the predicted semantic label information and the actual semantic label information; Determining the weight prediction loss value of the sample medical record text according to the prior expert weight values corresponding to the multiple expert encoding networks and the predicted expert weight values corresponding to the multiple expert encoding networks; Determining the model prediction loss value of the model to be trained according to the semantic prediction loss value of the sample medical record text and the weight prediction loss value of the sample medical record text; Updating the model to be trained according to the model prediction loss value.
5. The method according to claim 2, characterized in that, the decision layer includes an expert decision network and a decision fusion network; the process of determining the predicted semantic label information includes: Inputting the m-th expert encoding result into the expert decision network to obtain the m-th expert decision result output by the expert decision network; where m is a positive integer and m ≤ M; Inputting the M expert decision results and the predicted expert weight values corresponding to the M expert encoding networks into the decision fusion network to obtain the predicted semantic label information of the sample medical record text output by the decision fusion network.
6. The method according to claim 5, characterized in that, updating the model to be trained according to the predicted semantic label information and the actual semantic label information includes: Update the model to be trained according to the M expert decision results, the predicted expert weight values corresponding to the M expert encoding networks, and the actual semantic label information.
7. The method according to claim 6, wherein, the updating of the model to be trained according to the M expert decision results, the predicted expert weight values corresponding to the M expert encoding networks, and the actual semantic label information includes: determine the decision loss value of the m-th expert according to the decision result of the m-th expert and the actual semantic label information; where m is a positive integer and m ≤ M; determine the semantic prediction loss value of the sample medical record text according to the decision loss values of the M experts and the predicted expert weight values corresponding to the M expert encoding networks; update the model to be trained according to the semantic prediction loss value of the sample medical record text.
8. The method according to claim 7, wherein, the process of determining the decision loss value of the m-th expert includes: determine the decision loss value of the m-th expert according to the decision result of the m-th expert, the actual semantic label information, and the network loss function corresponding to the m-th expert encoding network; wherein, the penalty factor in the network loss function corresponding to the m-th expert encoding network is different from the penalty factor in the network loss function corresponding to any other expert encoding network among the M expert encoding networks except the m-th expert encoding network.
9. The method according to claim 5, wherein, the actual semantic label information of the sample medical record text is determined from K candidate semantic labels; where K is a positive integer; the process of determining the decision result of the m-th expert includes: perform a vector dot product operation on the encoding result of the m-th expert and the character feature vector of the k-th candidate semantic label to obtain the k-th label attribution probability information of the encoding result of the m-th expert; where k is a positive integer and k ≤ K; perform a set process on the 1st label attribution probability information to the K-th label attribution probability information of the encoding result of the m-th expert to obtain the decision result of the m-th expert.
10. The method according to claim 2, wherein, the model to be trained further includes a text slicing layer; the process of determining the text encoding result to be used includes: input the sample medical record text into the text slicing layer to obtain at least one text segment output by the text slicing layer; input the at least one text segment into the text encoding layer to obtain the text encoding result to be used output by the text encoding layer; wherein, the text encoding result to be used includes the text encoding results of the at least one text segment.
11. The method according to claim 10, wherein, the text encoding layer includes a first encoding network and a second encoding network; the process of determining the text encoding results of the at least one text segment includes: input each of the text segments into the first encoding network to obtain the preliminary encoding results of each of the text segments output by the first encoding network; Input the preliminary encoding result of the at least one text segment into the second encoding network to obtain the text encoding result of the at least one text segment output by the second encoding network.
12. The method according to any one of claims 1-9, wherein, the obtaining the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained includes: performing a preset partitioning process on the sample medical record text to obtain at least one sample segment; inputting the at least one sample segment into the model to be trained to obtain the predicted semantic label information of the sample medical record text output by the model to be trained; wherein, the predicted semantic label information includes the segment semantic label information of the at least one sample segment.
13. A medical record parsing method, wherein, the method includes: obtaining a medical record text to be processed; determining the predicted semantic label information of the medical record text to be processed according to the medical record text to be processed and the semantic label determination model; wherein, the semantic label determination model is constructed by using the construction method of the semantic label determination model according to any one of claims 1 to 12; determining the semantic parsing result of the medical record text to be processed according to the predicted semantic label information of the medical record text to be processed.
14. The method according to claim 13, wherein, the predicted semantic label information of the medical record text to be processed includes the segment semantic label information of at least one segment to be processed in the medical record text to be processed; the determining the semantic parsing result of the medical record text to be processed according to the predicted semantic label information of the medical record text to be processed includes: integrating the segment semantic label information of the at least one segment to be processed according to a preset integration rule to obtain the segment semantic label information of at least one segment to be used; determining the semantic parsing result of the medical record text to be processed according to the segment semantic label information of the at least one segment to be used.
15. The method according to claim 13 or 14, wherein, the process of determining the semantic parsing result includes: performing error correction on the semantic label information to be corrected according to a preset error correction rule to obtain the corrected semantic label; determining the semantic parsing result of the medical record text to be processed according to the corrected semantic label.
16. The method according to claim 13, wherein, the method further includes: determining whether there is a phenomenon of missing fields in the medical record text to be processed according to the semantic parsing result of the medical record text to be processed and a preset complete condition of medical record fields.
17. A device for constructing a semantic label determination model, wherein, includes: a first obtaining unit, configured to obtain a sample medical record text and the actual semantic label information of the sample medical record text; A first determination unit, configured to determine predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained; the model to be trained includes an expert encoding layer and an expert weight determination layer, the expert encoding layer includes a plurality of expert encoding networks, and the output data of the expert weight determination layer includes predicted expert weight values corresponding to the plurality of expert encoding networks; A model update unit, configured to update the model to be trained according to the predicted semantic label information and the actual semantic label information, and return to the first determination unit to continue executing the step of obtaining the predicted semantic label information of the sample medical record text according to the sample medical record text and the model to be trained, until when a preset stop condition is reached, a semantic label determination model is determined according to the model to be trained, and different expert encoding networks in the semantic label determination model are good at processing different semantic parsing directions during expert encoding processing; The model update unit is specifically configured to update the model to be trained according to the predicted semantic label information, the actual semantic label information, prior expert weight values corresponding to the plurality of expert encoding networks, and predicted expert weight values corresponding to the plurality of expert encoding networks; wherein, the prior expert weight values are obtained by looking up a prior weight set corresponding to the sample type information of the sample medical record text from a preset mapping relationship, and the sample type information includes label determination difficulty characterization data and label frequency characterization data; the preset mapping relationship includes a correspondence between candidate type information and a prior weight set, the label determination difficulty characterization data in the first candidate type information reaches a difficulty threshold, the label frequency characterization data in the first candidate type information is higher than a frequency threshold, the label determination difficulty characterization data in the second candidate type information reaches a difficulty threshold, the label frequency characterization data in the second candidate type information is not higher than a frequency threshold, the label determination difficulty characterization data in the third candidate type information is lower than a difficulty threshold, the label frequency characterization data in the third candidate type information is higher than a frequency threshold, and the label determination difficulty characterization data in the fourth candidate type information is lower than a difficulty threshold, and the label frequency characterization data in the fourth candidate type information is not higher than a frequency threshold.
18. A medical record parsing device Characterized in that It includes: A second acquisition unit, configured to acquire a medical record text to be processed; A second determination unit, configured to determine predicted semantic label information of the medical record text to be processed according to the medical record text to be processed and the semantic label determination model; wherein, the semantic label determination model is constructed by using the construction method of the semantic label determination model according to any one of claims 1 to 12; A third determination unit, configured to determine a semantic parsing result of the medical record text to be processed according to the predicted semantic label information of the medical record text to be processed.
19. A device Characterized in that The device includes: a processor, a memory, and a system bus; The processor and the memory are connected through the system bus; The memory is used to store one or more programs, and the one or more programs include instructions which, when executed by the processor, cause the processor to execute the method for constructing the semantic tag determination model according to any one of claims 1 to 12, or execute the medical record parsing method according to any one of claims 13 to 16.
20. A computer-readable storage medium, characterized in that the computer-readable storage medium stores instructions which, when running on a terminal device, cause the terminal device to execute the method for constructing the semantic tag determination model according to any one of claims 1 to 12, or execute the medical record parsing method according to any one of claims 13 to 16.
21. A computer program product, characterized in that when the computer program product runs on a terminal device, it causes the terminal device to execute the method for constructing the semantic tag determination model according to any one of claims 1 to 12, or execute the medical record parsing method according to any one of claims 13 to 16.
Citation Information
Patent Citations
Recommendation method and device based on artificial intelligence, electronic equipment and storage medium
CN111291266A
Electronic medical record standardized segmentation method and device
CN112732863A
Cited By
Medical data retrieval method and system and storage medium
CN122112235A