A corpus prediction method, model training method and related devices

By using a target corpus prediction model in the electronic medical record writing process, the model selects predicted corpus with high relevance and matching degree based on the preceding corpus as writing prompts, which solves the problem of inaccurate prediction of the following corpus in the existing technology and improves the writing efficiency of electronic medical records.

CN116301401BActive Publication Date: 2026-05-15BEIJING JIAHE HAISEN HEALTH TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING JIAHE HAISEN HEALTH TECH CO LTD
Filing Date
2023-03-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing input method software cannot provide effective predictive text prompts when writing electronic medical records, resulting in low efficiency in writing electronic medical records.

Method used

The target corpus prediction model selects target prediction corpora that meet the preset relevance conditions from the candidate prediction corpus set based on the preceding corpus, determines the matching degree between the target prediction corpus and the preceding corpus, and provides writing prompts after sorting.

Benefits of technology

It improves the efficiency of writing electronic medical records, ensures that prompts match the writing scenario of electronic medical records, and enhances the accuracy and efficiency of the input method software in predicting the following text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116301401B_ABST
    Figure CN116301401B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a corpus prediction method, a model training method and related devices. The method comprises: obtaining preceding corpus input in the process of writing an electronic medical record; then, selecting, by a target corpus prediction model, each target prediction corpus from a candidate prediction corpus set that satisfies a preset condition in terms of relevance to the preceding corpus; further, determining, for each target prediction corpus, a matching degree between the target prediction corpus and the preceding corpus as a matching degree corresponding to the target prediction corpus; finally, sorting each target prediction corpus according to the respective matching degrees of the target prediction corpora, and taking the sorted target prediction corpora as writing prompt information of the electronic medical record. The method can provide effective prompt information when writing an electronic medical record, thereby improving the writing efficiency of the electronic medical record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, specifically to a corpus prediction method, a model training method, and related apparatus. Background Technology

[0002] With the rapid development of electronic technology, the amount of data in all industries has exploded, ushering in the era of big data. Big data from transportation, meteorology, finance, commerce, and biomedicine permeates people's daily work and lives. In this era of big data, electronic medical records (EMRs) are gradually becoming more common in hospitals. Also known as computer-based patient records, EMRs are stored, managed, and transmitted using electronic devices, replacing handwritten paper medical records.

[0003] However, due to the unique nature of medical corpora, when using relevant input method software to write electronic medical records, the predicted context provided by the input method software to give writing prompts generally cannot provide effective prompts when writing electronic medical records. That is, the predicted context provided is not the context actually needed when writing electronic medical records, and therefore cannot effectively improve the writing efficiency of electronic medical records. Summary of the Invention

[0004] This application provides a corpus prediction method, a model training method, and related apparatus that can provide effective prompts when writing electronic medical records, thereby improving the efficiency of electronic medical record writing.

[0005] In view of this, the first aspect of this application provides a corpus prediction method, the method comprising:

[0006] Obtain the preceding text input during the writing of electronic medical records;

[0007] Using the target corpus prediction model, based on the above text corpus, select each target prediction corpus from the candidate prediction corpus set whose relevance to the above text corpus meets the preset conditions;

[0008] For each target prediction corpus, the matching degree between the target prediction corpus and the preceding corpus is determined, and this degree is used as the matching degree corresponding to the target prediction corpus.

[0009] Based on the matching degree of each target prediction corpus, the target prediction corpus is sorted, and the sorted target prediction corpus is used as writing prompt information for the electronic medical record.

[0010] Optionally, using a target corpus prediction model, based on the preceding text corpus, select from the candidate prediction corpus set each target prediction corpus whose relevance to the preceding text corpus meets preset conditions, including:

[0011] The correlation between each candidate predicted corpus in the candidate predicted corpus set and the preceding text corpus is determined using the target corpus prediction model, and is used as the correlation corresponding to the candidate predicted corpus.

[0012] In the candidate prediction corpus set, the n candidate prediction corpora with the highest relevance are selected as the target prediction corpus; where n is an integer greater than 1.

[0013] Optionally, determining the matching degree between the target prediction corpus and the preceding text corpus for each target prediction corpus, as the matching degree corresponding to the target prediction corpus, includes:

[0014] Extract the target entity from the above corpus, and determine the entity corpus features corresponding to the above corpus based on the target entity;

[0015] For each target prediction corpus, the similarity between the corpus features of the target prediction corpus and the entity corpus features is calculated, which is used as the matching degree corresponding to the target prediction corpus.

[0016] A second aspect of this application provides a model training method, the method comprising:

[0017] Acquire training electronic medical record corpus;

[0018] Based on the training electronic medical record corpus, multiple training samples are generated, including training context corpus and training context corpus; and based on the training electronic medical record corpus, a candidate prediction corpus set is constructed.

[0019] Based on the multiple training samples, an initial corpus prediction model is trained; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus.

[0020] Once the initial corpus prediction model obtained after training meets the training termination condition, the initial corpus prediction model is used as the target corpus prediction model; the target corpus prediction model is used to provide writing prompts in the electronic medical record writing scenario.

[0021] Optionally, the step of generating multiple training samples based on the training electronic medical record corpus includes:

[0022] The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus.

[0023] The training preceding text is formed by using a first number of adjacent word segments in the training electronic medical record corpus; the training following text is formed by using a second number of word segments in the training electronic medical record corpus that are located after the training preceding text.

[0024] Optionally, constructing a candidate prediction corpus set based on the training electronic medical record corpus includes:

[0025] The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus.

[0026] The candidate prediction corpus set is constructed using the various word segments included in the training electronic medical record corpus.

[0027] Optionally, the method further includes:

[0028] Based on each test sample included in the test sample set, the initial corpus prediction model is tested to obtain the test accuracy corresponding to the initial corpus prediction model; the test samples include the test context corpus and the test context corpus;

[0029] If the test accuracy of the initial corpus prediction model exceeds the preset accuracy threshold, then the initial corpus prediction model is determined to meet the training termination condition.

[0030] Optionally, the step of testing the initial corpus prediction model based on each test sample included in the test sample set to obtain the test accuracy corresponding to the initial corpus prediction model includes:

[0031] For each test sample, the initial corpus prediction model determines the predicted context corpus set corresponding to the test context corpus based on the test context corpus in the test sample and the candidate prediction corpus set; if the predicted context corpus set includes the test context corpus in the test sample, then the test sample is determined to be an accurate prediction sample.

[0032] The test accuracy of the initial corpus prediction model is determined based on the proportion of the accurately predicted samples in the test sample set.

[0033] A third aspect of this application provides a corpus prediction apparatus, the apparatus comprising:

[0034] The corpus acquisition module is used to acquire the above-mentioned corpus input during the writing of electronic medical records;

[0035] The corpus prediction module is used to select, from the candidate prediction corpus set, each target prediction corpus that meets the preset condition in terms of relevance to the above text corpus, based on the target corpus prediction model and the above text corpus.

[0036] The matching degree determination module is used to determine the matching degree between the target prediction corpus and the preceding text corpus for each target prediction corpus, and use it as the matching degree corresponding to the target prediction corpus;

[0037] The corpus sorting module is used to sort the target prediction corpora according to their respective matching degrees, and use the sorted target prediction corpora as writing prompts for the electronic medical record.

[0038] A fourth aspect of this application provides a model training apparatus, the apparatus comprising:

[0039] The training corpus acquisition module is used to acquire training electronic medical record corpora.

[0040] The training corpus processing module is used to generate multiple training samples based on the training electronic medical record corpus, the training samples including training context corpus and training context corpus; and to construct a candidate prediction corpus set based on the training electronic medical record corpus.

[0041] The model training module is used to train an initial corpus prediction model based on the multiple training samples; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus.

[0042] The model testing module is used to take the initial corpus prediction model as the target corpus prediction model after the training termination condition is met; the target corpus prediction model is used to provide writing prompts in the electronic medical record writing scenario.

[0043] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0044] This application provides a corpus prediction method, which includes: acquiring the preceding text input during the writing of an electronic medical record; then, using a target corpus prediction model, selecting target prediction corpora from a set of candidate prediction corpora whose relevance to the preceding text meets preset conditions based on the preceding text; further, for each target prediction corpus, determining the matching degree between the target prediction corpus and the preceding text, as the matching degree corresponding to the target prediction corpus; finally, sorting the target prediction corpora according to their respective matching degrees, and using the sorted target prediction corpora as writing prompts for the electronic medical record. This method innovatively provides a corresponding writing prompt information determination method for electronic medical record writing scenarios. The method pre-trains a target corpus prediction model to predict the preceding text during electronic medical record writing. This target corpus prediction model can determine the relevance between the preceding text input during electronic medical record writing and each candidate prediction corpus in the candidate prediction corpus set. Based on this, it determines the target prediction corpus with a high relevance to the preceding text, which serves as the writing prompt information subsequently displayed to the user. Since each candidate prediction corpus stored in the candidate prediction corpus set is based on... The candidate prediction corpus determined by training the electronic medical record corpus consists of corpora commonly used in electronic medical record writing scenarios. Therefore, it can be ensured that the target prediction corpus determined for the input context is also applicable to the electronic medical record writing scenario. Furthermore, the matching degree between each target prediction corpus and the context corpus is determined, and the target prediction corpus is ranked accordingly. The ranked results are provided to the user as writing prompts. This ensures that the target prediction corpus that matches the current input context corpus is recommended to the user first, providing better writing prompts and thus improving the efficiency of electronic medical record writing. Attached Figure Description

[0045] Figure 1 A flowchart illustrating a corpus prediction method provided in an embodiment of this application;

[0046] Figure 2 A flowchart illustrating another corpus prediction method provided in an embodiment of this application;

[0047] Figure 3 A schematic flowchart illustrating a model training method provided in an embodiment of this application;

[0048] Figure 4 This is a schematic diagram of the structure of a corpus prediction device provided in an embodiment of this application;

[0049] Figure 5 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application. Detailed Implementation

[0050] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0051] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0052] See Figure 1 , Figure 1 This is a flowchart illustrating the corpus prediction method provided in the embodiments of this application, as shown below. Figure 1 As shown, this corpus prediction method includes the following steps:

[0053] Step 101: Obtain the above text data entered during the writing of electronic medical records.

[0054] In this embodiment of the application, a user can write electronic medical records using input method software specifically designed for writing electronic medical records. During the process of the user writing the electronic medical record using this input method software, the software can acquire the preceding text input by the user.

[0055] The preceding text corpus here refers to the text currently input by the user. Specifically, it can be determined based on punctuation marks in the user's input. For example, all text from the last punctuation mark to the last character of the current input (including the last character) can be used as the preceding text corpus. Alternatively, it can be determined based on a preset number of word segments. For example, a preset number of word segments (including the last segment) can be counted backward from the last word segment in the user's input and used as the preceding text corpus. Of course, in practical applications, the preceding text corpus can also be determined in other ways, and this application embodiment does not impose any limitations on this.

[0056] Step 102: Using the target corpus prediction model, select from the candidate prediction corpus set each target prediction corpus whose relevance to the above text corpus meets the preset conditions based on the above text corpus.

[0057] After obtaining the preceding text corpus input by the user, preprocessing can be performed. Specifically, the corpus can be segmented into words to obtain individual words. Then, the encoding vector for each word is determined by looking up a table, and the position vector for each word is determined based on its position in the corpus. For each word, its encoding vector and position vector are concatenated to obtain its comprehensive vector. Finally, the comprehensive vectors of each word are concatenated according to their order in the corpus to obtain the model input data for that corpus. Additionally, word segmentation can be cleaned after segmentation to remove useless information and punctuation marks.

[0058] Then, using a pre-trained target corpus prediction model, based on the model input data corresponding to the preceding text corpus, several target prediction corpora with high relevance to the preceding text corpus are selected from the candidate prediction corpus set. This target corpus prediction model is used to predict the relevance between the preceding text corpus and each candidate prediction corpus in the candidate prediction corpus set. The training method of this target corpus prediction model will be described in detail below. The candidate prediction corpus set stores multiple candidate prediction corpora, all of which are determined based on the training electronic medical record corpus. The construction method of this candidate prediction corpus set will also be described in detail below.

[0059] In practice, the target corpus prediction model can first determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the preceding corpus based on the input preceding corpus, and use this relevance as the relevance of the candidate prediction corpus. Then, the candidate prediction corpus sets can be arranged in descending order of their relevance. Next, the top n (n is an integer greater than 1) candidate prediction corpus are selected as the target prediction corpus corresponding to the preceding corpus, that is, the n candidate prediction corpus with the highest relevance are selected as the target prediction corpus corresponding to the preceding corpus.

[0060] To facilitate understanding, the following example illustrates the working process of a target corpus prediction model, which consists of an encoder structure and a softmax structure. Figure 2 This is a schematic diagram illustrating the working principle of the target corpus prediction model.

[0061] Suppose the input corpus is "The patient experienced mild chest tightness 1 year ago". Before processing this corpus through the target corpus prediction model, it needs to be segmented into words, specifically into "patient", "1 year ago", "experienced", and "mild chest tightness". Then, each segment is encoded to determine its corresponding encoding vector (embedding). The embedding can be determined by looking up a table. Next, based on the embedding and position vector of each segment, the actual input data for the target corpus prediction model is determined. Here, the position vector is determined by the position of each segment in the input corpus. Specifically, for each segment, its embedding and position vector are concatenated to obtain a comprehensive vector. Then, based on the position of each segment in the input corpus, the comprehensive vectors are concatenated to obtain the actual input data for the target corpus prediction model.

[0062] After obtaining the actual input data, the actual input data is fed into the Encoder structure of the target corpus prediction model. Taking the Encoder structure in the transformer model as an example, the basic calculation formula of the Encoder structure is as follows:

[0063] H1=LayerNorm(X+MultiHeadAttention(X))

[0064] H2=LayerNorm(H1+FeedForward(H1))

[0065] Where X is the actual input data of the Encoder structure, H1 is the output of the first LayerNorm (normalization layer) in the Encoder structure, and H2 is the output of the second LayerNorm in the Encoder structure.

[0066] The specific formula for MultiHeadAttention is as follows:

[0067]

[0068] Where Q, K, and V are the actual input data to MultiHeadAttention, and d k Let Q, K, and V be the number of dimensions.

[0069] The specific formula for FeedForward is as follows:

[0070] max(0,XW1+b1)W2+b2

[0071] Among them, W1, W2, b1, and b2 are all model parameters of FeedForward.

[0072] The specific formula for LayerNorm is as follows:

[0073]

[0074] Where x is the input data of LayerNorm and y is the output data of LayerNorm. Specifically, x can be the input data of the first LayerNorm in the Encoder structure, that is, the model input data determined based on the above corpus, or x can be the input data of the second LayerNorm in the Encoder structure, that is, the output data of the first LayerNorm in the Encoder structure; y can be the output data of the first LayerNorm in the Encoder structure or the output data of the second LayerNorm in the Encoder structure, that is, the model processing result used to represent the semantic features of the above corpus; ∈, γ and β are all model parameters of LayerNorm.

[0075] After the Encoder structure of the target corpus prediction model processes the actual input data, the processing result can be further input into the Softmax structure. The Softmax structure determines the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus.

[0076] After obtaining the correlation between each candidate predicted corpus and the preceding text in the candidate predicted corpus set, the three candidate predicted corpuses with the highest correlation with the preceding text can be selected as the target predicted corpus corresponding to the preceding text. For example, for the preceding text "The patient experienced mild chest tightness 1 year ago", the target predicted corpus prediction model can determine "after rest", "without" and "visited at" in the candidate predicted corpus set as the target predicted corpus corresponding to the preceding text.

[0077] It should be understood that in practical applications, the target corpus prediction model can also have other structures. For example, the encoder structure in the target corpus prediction model can be not only the encoder structure in a transformer model, but also a fully connected model, a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, and so on. Furthermore, the target corpus prediction model can also be an N-gram model; this application embodiment does not impose any limitations on the structure of the target corpus prediction model.

[0078] Step 103: For each target prediction corpus, determine the matching degree between the target prediction corpus and the preceding text corpus, and use it as the matching degree corresponding to the target prediction corpus.

[0079] After determining the target prediction corpus corresponding to the input preceding text, we can further determine the matching degree between each target prediction corpus and the preceding text, and use this as the matching degree for the target prediction corpus, so that we can sort the target prediction corpus corresponding to the preceding text in the future.

[0080] In practice, the target entity can be extracted from the preceding text, and the entity text features corresponding to the preceding text can be determined based on the target entity. Then, for each target prediction text corresponding to the preceding text, the similarity between the text features of the target prediction text and the entity text features is calculated as the matching degree corresponding to the target prediction text.

[0081] More specifically, based on a pre-defined target entity type, entities belonging to that target entity type can be extracted from the preceding text corpus and used as the target entities for that preceding text corpus. The target entity type could be, for example, disease, symptom, medicine, surgery, examination, testing, etc. For each target entity in the preceding text corpus, the corresponding corpus vector can be determined by looking up a table. This corpus vector can then be used as the entity corpus feature for that preceding text corpus. It should be understood that when the preceding text corpus includes multiple target entities, the corpus vectors corresponding to each of the multiple target entities can be concatenated to obtain the entity corpus feature for that preceding text corpus. For each target prediction corpus corresponding to that preceding text corpus, the corresponding corpus vector can be determined by looking up a table and used as the corpus feature for that target prediction corpus.

[0082] Furthermore, for each target prediction corpus corresponding to the preceding text corpus, the matching degree between the target prediction corpus and the preceding text corpus can be determined by calculating the similarity between its corpus features and the entity corpus features corresponding to the preceding text corpus. The specific formula for calculating the similarity is as follows:

[0083]

[0084] Where A and B represent the entity corpus features corresponding to the preceding text corpus and the target prediction corpus corpus, respectively. i B represents the corpus vector corresponding to the i-th target entity in the preceding corpus, i.e., the i-th sub-feature in the entity corpus features, where n represents the number of target entities included in the preceding corpus. jLet represent the corpus features of the j-th target prediction corpus, and m represent the total number of target prediction corpora. and represent sub-features in the entity corpus features and sub-features in the target prediction corpus features, respectively.

[0085] Step 104: Sort the target prediction corpora according to their respective matching degrees, and use the sorted target prediction corpora as writing prompts for the electronic medical record.

[0086] After determining the matching degree of each target prediction corpus corresponding to the input preceding text, the target prediction corpora can be sorted in descending order of matching degree. The sorted target prediction corpora are then recommended to the user to provide relevant writing prompts during the electronic medical record writing process. Accordingly, the user can select from the recommended target prediction corpora. The selected target prediction corpus will be deployed as the following text corpus after the preceding text corpus mentioned in step 101 to further improve the currently being written electronic medical record. It should be understood that the following text corpus here refers to the corpus deployed after the preceding text corpus mentioned in step 101 in the electronic medical record. In this embodiment, the following text corpus can be selected from the recommended target prediction corpora. Of course, in practical applications, the user can also manually input the following text corpus. Figure 2 As shown, for the target prediction corpus "after rest", "without accompanying" and "at the clinic" corresponding to the above corpus "the patient experienced mild chest tightness 1 year ago", after determining the matching degree between each of them and the above corpus, they are sorted in order of matching degree from high to low, and the writing prompt information is determined to be "after rest", "without accompanying" and "at the clinic".

[0087] The corpus prediction method provided in this application innovatively offers a corresponding writing prompt information determination scheme for electronic medical record writing scenarios. This method pre-trains a target corpus prediction model for predicting the following text during electronic medical record writing. This target corpus prediction model can determine the relevance between the preceding text input during electronic medical record writing and each candidate prediction corpus in the candidate prediction corpus set. Based on this, it determines the target prediction corpus with a high relevance to the preceding text, which serves as the writing prompt information subsequently displayed to the user. Since the candidate prediction corpus set stores various candidate prediction corpus... The corpora are all determined based on the training corpora of electronic medical records. These candidate prediction corpora are commonly used in electronic medical record writing scenarios. Therefore, it can be ensured that the target prediction corpora determined for the input context are also applicable to the electronic medical record writing scenario. Furthermore, the matching degree between each target prediction corpus and the context is determined, and the target prediction corpora are ranked accordingly. The ranked results are provided to the user as writing prompts. This ensures that the target prediction corpus that matches the current input context is recommended to the user first, providing better writing prompts and thus improving the efficiency of electronic medical record writing.

[0088] See Figure 3 , Figure 3 This is a flowchart illustrating the model training method provided in the embodiments of this application, as shown below. Figure 3 As shown, the model training method includes the following steps:

[0089] Step 301: Obtain the training electronic medical record corpus.

[0090] In this embodiment of the application, a large amount of electronic medical record corpus can be obtained from relevant databases. These relevant databases may be, for example, databases used by hospitals or research institutions to store electronic medical records. Then, these electronic medical record corpus are divided into training electronic medical record corpus and test electronic medical record corpus according to a specific ratio. For example, the obtained electronic medical record corpus can be divided into training electronic medical record corpus and test electronic medical record corpus according to a ratio of 9:1 or 8:2. The training electronic medical record corpus is used to construct training samples, and the test electronic medical record corpus is used to construct test samples.

[0091] Step 302: Generate multiple training samples based on the training electronic medical record corpus, the training samples including the training context corpus and the training context corpus; and construct a candidate prediction corpus set based on the training electronic medical record corpus.

[0092] After obtaining the electronic medical record (EMR) corpus, multiple training samples can be generated based on the training EMR corpus, and multiple test samples can be generated based on the test EMR corpus. Furthermore, a candidate prediction corpus set can be constructed based on the training EMR corpus; this candidate prediction corpus set is the one described above. Figure 1 The candidate prediction corpus set in the illustrated embodiment.

[0093] In practice, training samples can be generated from the training electronic medical record corpus in the following way: the training electronic medical record corpus is segmented to obtain each segment included in the training electronic medical record corpus; then, a first number of adjacent segments in the training electronic medical record corpus are used to form a training context corpus; a second number of segments in the training electronic medical record corpus that are located after the training context corpus are used to form a training context corpus. The training context corpus and the training context corpus can then constitute a training sample.

[0094] For example, taking the training electronic medical record corpus as "the patient experienced mild chest tightness 1 year ago", the corpus is segmented into words, resulting in the words "patient", "1 year ago", "occurred", "mild", and "chest tightness". Two adjacent segments can be used to form the preceding training text, and the segment following the preceding text can be used as the following training text. Thus, the following training samples can be constructed: Training Sample 1 (where the preceding training text is "the patient experienced mild chest tightness 1 year ago" and the following training text is "occurred"), Training Sample 2 (where the preceding training text is "occurred 1 year ago" and the following training text is "mild"), and Training Sample 3 (where the preceding training text is "occurred mildly" and the following training text is "chest tightness").

[0095] It should be understood that the implementation method of generating test samples based on the test electronic medical record corpus is basically the same as the implementation method of generating training samples based on the training electronic medical record corpus described above, and will not be repeated here.

[0096] In practice, a candidate prediction corpus can be constructed based on the training electronic medical record corpus as follows: The training electronic medical record corpus is segmented to obtain the individual words included in the corpus; then, the individual words included in each training electronic medical record corpus are used to construct the candidate prediction corpus. In other words, the individual words included in each training electronic medical record corpus can be directly used as candidate prediction corpora to form the candidate prediction corpus.

[0097] Step 303: Based on the multiple training samples, train an initial corpus prediction model; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus.

[0098] After generating multiple training samples based on the training electronic medical record corpus, these training samples can be further used to train the initial corpus prediction model. Specifically, for each training sample, the training context corpus included can be preprocessed to obtain the embeddings corresponding to each word segment in the training context corpus. The embeddings and positional encoding vectors of each word segment in the training context corpus are then used to construct the input data for the initial corpus prediction model. This input data is then fed into the initial corpus prediction model, which analyzes and processes the input data, outputting the correlation between the training context corpus and each candidate prediction corpus in the candidate prediction corpus set. Furthermore, based on the training context corpus corresponding to the training context corpus and the correlation between the training context corpus and each candidate prediction corpus, a loss function can be constructed. The model parameters of the initial corpus prediction model are then adjusted based on this loss function, thus training the initial corpus prediction model.

[0099] It should be understood that the working principle of the initial corpus prediction model here is the same as that described above. Figure 1 The target corpus prediction models described in the embodiments have the same working principle. The only difference between the two is the model parameters. Therefore, the working principle of the initial corpus prediction model will not be described again here.

[0100] Step 304: After the initial corpus prediction model obtained from the training meets the training termination condition, the initial corpus prediction model is used as the target corpus prediction model; the target corpus prediction model is used to provide writing prompts in the electronic medical record writing scenario.

[0101] During the training of the initial corpus prediction model, previously constructed test samples can be used to test whether the trained initial corpus prediction model meets the training termination condition. Once it is determined that the initial corpus prediction model meets the training termination condition, it can be regarded as the target corpus prediction model and deployed in a real electronic medical record writing scenario to provide writing prompts to users.

[0102] When specifically testing the initial corpus prediction model, the model can be tested based on each test sample included in the test sample set to obtain the test accuracy corresponding to the initial corpus prediction model. The test samples here are similar to the training samples mentioned above, including the test context corpus and the test context corpus. If the test accuracy corresponding to the initial corpus prediction model exceeds the preset accuracy threshold, it can be determined that the initial corpus prediction model meets the training termination condition.

[0103] More specifically, for each test sample in the test sample set, the preceding context of that test sample can be processed to obtain corresponding input data. This input data is then fed into the currently trained initial corpus prediction model. This initial corpus prediction model can accordingly determine multiple predicted contexts that are closely related to the preceding context, and uses these predicted contexts to form a set of predicted contexts corresponding to the preceding context. If the set of predicted contexts includes the corresponding context of the preceding context in the test sample, then the test sample is determined to be an accurate prediction sample; conversely, if the set of predicted contexts does not include the corresponding context of the preceding context in the test sample, then the test sample is determined not to be an accurate prediction sample. Furthermore, the proportion of each test sample that is an accurate prediction sample in the test sample set can be calculated as the test accuracy corresponding to the initial corpus prediction model. If the test accuracy corresponding to the initial corpus prediction model exceeds a preset accuracy threshold, then the initial corpus prediction model is determined to meet the training termination condition.

[0104] The model training method provided in this application trains a target corpus prediction model specifically for predicting contextual corpus during electronic medical record writing. This target corpus prediction model can determine the relevance between the preceding text input during electronic medical record writing and each candidate prediction corpus in a candidate prediction corpus set. Based on this, it determines the target prediction corpus with a high relevance to the preceding text, which serves as writing prompts subsequently displayed to the user. Since each candidate prediction corpus stored in the candidate prediction corpus set is determined based on the training electronic medical record corpus, and these candidate prediction corpuses are commonly used in electronic medical record writing scenarios, it can be ensured that the target prediction corpus determined for the input preceding text is also applicable to the electronic medical record writing scenario. This provides effective writing prompts in the electronic medical record writing scenario.

[0105] See Figure 4 , Figure 4 This is a schematic diagram of the corpus prediction device provided in the embodiments of this application, as shown below. Figure 4 As shown, the corpus prediction device includes:

[0106] Corpus acquisition module 401 is used to acquire the above-mentioned corpus input during the writing of electronic medical records;

[0107] The corpus prediction module 402 is used to select, based on the preceding text corpus, each target prediction corpus that satisfies a preset condition in the candidate prediction corpus set through the target corpus prediction model;

[0108] The matching degree determination module 403 is used to determine the matching degree between the target prediction corpus and the preceding text corpus for each target prediction corpus, and use it as the matching degree corresponding to the target prediction corpus;

[0109] The corpus sorting module 404 is used to sort the target prediction corpora according to their respective matching degrees, and use the sorted target prediction corpora as writing prompts for the electronic medical record.

[0110] Optionally, the corpus prediction module 402 is specifically used for:

[0111] The correlation between each candidate predicted corpus in the candidate predicted corpus set and the preceding corpus is determined using the target corpus prediction model, and is used as the correlation corresponding to the candidate predicted corpus.

[0112] In the candidate prediction corpus set, the n candidate prediction corpora with the highest relevance are selected as the target prediction corpus; where n is an integer greater than 1.

[0113] Optionally, the matching degree determination module 403 is specifically used for:

[0114] Extract the target entity from the above corpus, and determine the entity corpus features corresponding to the above corpus based on the target entity;

[0115] For each target prediction corpus, the similarity between the corpus features of the target prediction corpus and the entity corpus features is calculated, which is used as the matching degree corresponding to the target prediction corpus.

[0116] See Figure 5 , Figure 5 This is a schematic diagram of the structure of the model training device provided in the embodiments of this application, as shown below. Figure 5 As shown, the model training device includes:

[0117] Training corpus acquisition module 501 is used to acquire training electronic medical record corpus;

[0118] The training corpus processing module 502 is used to generate multiple training samples based on the training electronic medical record corpus, wherein the training samples include training context corpus and training context corpus; and to construct a candidate prediction corpus set based on the training electronic medical record corpus.

[0119] The model training module 503 is used to train an initial corpus prediction model based on the multiple training samples; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus.

[0120] The model testing module 504 is used to take the initial corpus prediction model as the target corpus prediction model after the initial corpus prediction model obtained from training meets the training termination condition; the target corpus prediction model is used to provide writing prompts in the electronic medical record writing scenario.

[0121] Optionally, the training corpus processing module 502 is specifically used for:

[0122] The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus;

[0123] The training preceding text is formed by using a first number of adjacent word segments in the training electronic medical record corpus; the training following text is formed by using a second number of word segments in the training electronic medical record corpus that are located after the training preceding text.

[0124] Optionally, the training corpus processing module 502 is specifically used for:

[0125] The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus;

[0126] The candidate prediction corpus set is constructed using the various word segments included in the training electronic medical record corpus.

[0127] Optionally, the model testing module 504 is further configured to:

[0128] Based on each test sample included in the test sample set, the initial corpus prediction model is tested to obtain the test accuracy corresponding to the initial corpus prediction model; the test samples include the test context corpus and the test context corpus;

[0129] If the test accuracy of the initial corpus prediction model exceeds the preset accuracy threshold, then the initial corpus prediction model is determined to meet the training termination condition.

[0130] Optionally, the model testing module 504 is specifically used for:

[0131] For each test sample, the initial corpus prediction model determines the predicted context corpus set corresponding to the test context corpus based on the test context corpus in the test sample and the candidate prediction corpus set; if the predicted context corpus set includes the test context corpus in the test sample, then the test sample is determined to be an accurate prediction sample.

[0132] The test accuracy of the initial corpus prediction model is determined based on the proportion of the accurately predicted samples in the test sample set.

[0133] The corpus prediction device and model training device provided in this application embodiment innovatively offer a corresponding writing prompt information determination scheme for electronic medical record writing scenarios. This method pre-trains a target corpus prediction model for predicting the following text during electronic medical record writing. This target corpus prediction model can determine the correlation between the preceding text input during electronic medical record writing and each candidate prediction corpus in the candidate prediction corpus set. Based on this, it determines the target prediction corpus with a high correlation to the preceding text, which serves as the writing prompt information subsequently displayed to the user. Since the candidate prediction corpus set stores various... The candidate prediction corpora are all determined based on the training electronic medical record corpora. These candidate prediction corpora are commonly used in electronic medical record writing scenarios. Therefore, it can be ensured that the target prediction corpora determined for the input context are also applicable to electronic medical record writing scenarios. Furthermore, the matching degree between each target prediction corpus and the context is determined, and the target prediction corpora are ranked accordingly. The ranked results are provided to the user as writing prompts. This ensures that the target prediction corpus that matches the current input context is recommended to the user first, providing better writing prompts and thus improving the efficiency of electronic medical record writing.

[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0135] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0137] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0138] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing computer programs.

[0139] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0140] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A corpus prediction method, characterized in that, The method includes: Obtain the preceding text input during the writing of electronic medical records; Based on the training electronic medical record corpus, the training electronic medical record corpus is segmented into words or each segmented word included in each training electronic medical record corpus is directly used as a candidate prediction corpus to construct a candidate prediction corpus set; the training electronic medical record corpus is used to generate training samples to train the target corpus prediction model. Using the target corpus prediction model, based on the preceding text, select each target prediction corpus from the candidate prediction corpus set whose relevance to the preceding text meets preset conditions; Based on a pre-defined target entity type, multiple target entities belonging to the target entity type are extracted from the preceding text corpus. A corpus vector corresponding to each of the multiple target entities is determined. By concatenating the corpus vectors, the entity corpus features corresponding to the preceding text corpus are determined. The target entity type includes one or more of the following: disease, symptom, medicine, surgery, examination, and testing. For each of the target prediction corpora, based on the formula: ; The similarity between the corpus features of the target prediction corpus and the entity corpus features is calculated as the matching degree corresponding to the target prediction corpus; where A and B represent the entity corpus features corresponding to the above corpus and the corpus features of the target prediction corpus, respectively. Let represent the corpus vector corresponding to the i-th target entity in the aforementioned corpus, and n represent the number of target entities included in the aforementioned corpus. The corpus feature represents the j-th target prediction corpus, and m represents the total number of target prediction corpora; Based on the matching degree of each target prediction corpus, the target prediction corpus is sorted, and the sorted target prediction corpus is used as writing prompt information for the electronic medical record.

2. The method according to claim 1, characterized in that, The step of selecting target prediction corpora from the candidate prediction corpus set based on the preceding text corpus using a target corpus prediction model, where the relevance between the target prediction corpus and the preceding text corpus meets preset conditions, includes: The correlation between each candidate predicted corpus in the candidate predicted corpus set and the preceding corpus is determined using the target corpus prediction model, and is used as the correlation corresponding to the candidate predicted corpus. In the candidate prediction corpus set, the n candidate prediction corpora with the highest relevance are selected as the target prediction corpus; where n is an integer greater than 1.

3. A model training method, characterized in that, The method includes: Acquire training electronic medical record corpus; Based on the training electronic medical record corpus, multiple training samples are generated, including training context corpus and training context corpus; and based on the training electronic medical record corpus, word segmentation is performed on the training electronic medical record corpus or each word segmentation included in each training electronic medical record corpus is directly used as candidate prediction corpus to construct a candidate prediction corpus set; the training electronic medical record corpus is used to generate training samples to train the target corpus prediction model; Based on the multiple training samples, an initial corpus prediction model is trained; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus. Once the initial corpus prediction model obtained through training meets the training termination condition, the initial corpus prediction model is used as the target corpus prediction model; the target corpus prediction model is applied to the corpus prediction method of claim 1 to select the target prediction corpus.

4. The method according to claim 3, characterized in that, The step of generating multiple training samples based on the training electronic medical record corpus includes: The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus; The training preceding text is formed by using a first number of adjacent word segments in the training electronic medical record corpus; the training following text is formed by using a second number of word segments in the training electronic medical record corpus that are located after the training preceding text.

5. The method according to claim 3, characterized in that, The step of constructing a candidate prediction corpus set based on the training electronic medical record corpus includes: The training electronic medical record corpus is segmented to obtain the individual words included in the training electronic medical record corpus; The candidate prediction corpus set is constructed using the various word segments included in the training electronic medical record corpus.

6. The method according to claim 3, characterized in that, The method further includes: Based on each test sample included in the test sample set, the initial corpus prediction model is tested to obtain the test accuracy corresponding to the initial corpus prediction model; the test samples include the test context corpus and the test context corpus; If the test accuracy of the initial corpus prediction model exceeds the preset accuracy threshold, then the initial corpus prediction model is determined to meet the training termination condition.

7. The method according to claim 6, characterized in that, The process of testing the initial corpus prediction model based on each test sample included in the test sample set, and obtaining the test accuracy corresponding to the initial corpus prediction model, includes: For each test sample, the initial corpus prediction model determines the predicted context corpus set corresponding to the test context corpus based on the test context corpus in the test sample and the candidate prediction corpus set; if the predicted context corpus set includes the test context corpus in the test sample, then the test sample is determined to be an accurate prediction sample. The test accuracy of the initial corpus prediction model is determined based on the proportion of the accurately predicted samples in the test sample set.

8. A corpus prediction device, characterized in that, The device includes: The corpus acquisition module is used to acquire the preceding text corpus input during the writing of electronic medical records; based on the training electronic medical record corpus, the training electronic medical record corpus is processed by word segmentation or each word segmentation included in each training electronic medical record corpus is directly used as candidate prediction corpus to construct a candidate prediction corpus set; the training electronic medical record corpus is used to generate training samples to train the target corpus prediction model. The corpus prediction module is used to select, based on the preceding text, each target prediction corpus from the candidate prediction corpus set whose relevance to the preceding text meets a preset condition, using a target corpus prediction model; The matching degree determination module, based on a pre-defined target entity type, extracts multiple target entities belonging to the target entity type from the preceding text corpus, determines a corpus vector corresponding to each of the multiple target entities, and determines the entity corpus features corresponding to the preceding text corpus by concatenating the corpus vectors. The target entity type includes one or more of the following: disease, symptom, medicine, surgery, examination, and testing. For each target prediction corpus, based on the formula: ; The similarity between the corpus features of the target prediction corpus and the entity corpus features is calculated as the matching degree corresponding to the target prediction corpus; where A and B represent the entity corpus features corresponding to the above corpus and the corpus features of the target prediction corpus, respectively. Let represent the corpus vector corresponding to the i-th target entity in the aforementioned corpus, and n represent the number of target entities included in the aforementioned corpus. The corpus feature represents the j-th target prediction corpus, and m represents the total number of target prediction corpora; The corpus sorting module is used to sort the target prediction corpora according to their respective matching degrees, and use the sorted target prediction corpora as writing prompts for the electronic medical record.

9. A model training device, characterized in that, The device includes: The training corpus acquisition module is used to acquire training electronic medical record corpora. The training corpus processing module is used to generate multiple training samples based on the training electronic medical record corpus, wherein the training samples include training context corpus and training context corpus; and to perform word segmentation processing on the training electronic medical record corpus or directly use each word segmentation included in each training electronic medical record corpus as candidate prediction corpus to construct a candidate prediction corpus set; the training electronic medical record corpus is used to generate training samples to train the target corpus prediction model; The model training module is used to train an initial corpus prediction model based on the multiple training samples; the initial corpus prediction model is used to determine the relevance between each candidate prediction corpus in the candidate prediction corpus set and the input preceding corpus. The model testing module is used to take the initial corpus prediction model as the target corpus prediction model after the initial corpus prediction model obtained from the training meets the training termination condition; the target corpus prediction model is applied to the corpus prediction method of claim 1 to select the target prediction corpus.