Method and apparatus for generating a CRF form for a rare disease
By entering keywords in the user terminal, using the medical report database and dictionary library to construct matrix similarity and text word segmentation processing, the CRF form for rare diseases is automatically generated, which solves the problems of lack of flexibility and high manual workload of rare diseases CRF forms in the prior art, and achieves efficient and accurate form generation.
Patent Information
- Application Number
- CN202510163036.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In the prior art, CRF forms for rare diseases lack flexibility, resulting in manual modification when the research content or disease type changes, increasing manual workload.
Enter keywords through the user terminal, use the medical report database and dictionary library on the server to construct a matrix to calculate the similarity, automatically generate CRF forms related to keywords, including medical word similarity and text word segmentation processing, and build a form containing guiding items.
This reduces the manual workload and improves the efficiency and accuracy of CRF form generation. Users only need to enter keywords to obtain relevant CRF forms.
Smart Images

Figure CN119623441B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly, to a method and apparatus for generating a CRF form for rare diseases. Background Art
[0002] The configuration of medical CRF (Case Report Form) forms usually customizes the specific content of the CRF form according to the research needs of researchers such as clinicians, public health physicians, pharmacists, and health economists. For rare diseases, there are both homogeneities and heterogeneities between different disease types or different disease subtypes. Therefore, the data collection form content required for different disease types or different disease subtypes is also different. In this context, the CRF form in the customization mode lacks sufficient flexibility. When the research content or the research disease type changes, it is necessary to manually modify the CRF form in the customization mode, resulting in a large amount of manual work. Summary of the Invention
[0003] In view of this, embodiments of this application provide a method and apparatus for generating a CRF form for rare diseases to provide corresponding CRF forms according to different rare diseases, thereby reducing the amount of manual work.
[0004] In a first aspect, an embodiment of this application provides a method for generating a CRF form for rare diseases. The generating method is applied to a form generation system, the form generation system includes a user terminal and a server, the server is provided with a medical report database and a medical dictionary library, the generating method runs on the server, and the generating method includes:
[0005] Responding to a keyword input by a user on the user terminal, performing word frequency statistics on each medical report in the medical report database, and determining an alternative word with the highest word frequency in each medical report;
[0006] For each medical word and each word to be compared in the medical dictionary, a matrix of (m + 1)*(n + 1) is constructed according to the first string length of the medical word and the second string length of the word to be compared, where m is the first string length, n is the second string length, and the words to be compared include the keyword and the alternative word;
[0007] Initialize the matrix elements in the first row of the matrix to integers from 0 to n in sequence, initialize the matrix elements in the first column of the matrix to integers from 0 to m in sequence, and for other matrix elements Tij in the matrix, determine the other matrix elements Tij according to a preset rule; wherein, the preset rule includes: if the i-th character of the medical term is the same as the j-th character of the term to be compared, then the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical term is different from the j-th character of the term to be compared, then the matrix element Tij is equal to the sum of the smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left of the matrix cell where the matrix element Tij is located and the value 1;
[0008] After determining the matrix elements in each matrix cell of the matrix, calculate the similarity between the medical term and the term to be compared using the following formula:
[0009] S = 1 - B / A;
[0010] wherein, B is the matrix element in the matrix cell located at the lower right corner of the matrix, and A is the maximum value of m and n;
[0011] After obtaining the first similarity between each alternative term and each medical term and the second similarity between the keyword and each medical term, determine the medical term corresponding to the first similarity with the smallest difference from the target similarity among the first similarities of each alternative term and the target medical term, wherein, the target medical term is the medical term corresponding to the maximum first similarity of each alternative term, and the target similarity is the maximum second similarity;
[0012] Construct a CRF form according to the terms with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical term.
[0013] Optionally, the constructing a CRF form according to the terms with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical term includes:
[0014] After obtaining the titles at all levels in the target medical report, perform word segmentation on the titles at all levels to obtain candidate word segments included in the titles at all levels;
[0015] For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment;
[0016] Calculate the similarity between the candidate word segment and a preset medical field term using the following formula:
[0017] S = ;
[0018] Among them, K is the first word vector corresponding to the candidate word segmentation, L is the second word vector corresponding to the preset medical field word, is the norm of the first word vector, is the norm of the second word vector, and d is the weight;
[0019] For the text content corresponding to each level of title, perform text processing on the text content to obtain the dependency relationship between text segmentations and text segmentations;
[0020] According to the part-of-speech of each text segmentation, determine the candidate text segmentations that have a dependency relationship with the text segmentations whose part-of-speech is a numeral from each text segmentation;
[0021] Select the part-of-speech of the specified type from the candidate text segmentations as the target text segmentation;
[0022] Use the candidate word segmentations with a similarity higher than the preset threshold as the parent items. For each parent item, use the target text segmentations included in the text content under the title corresponding to the parent item as the child items to construct the CRF form.
[0023] Optionally, the generation method further includes:
[0024] Display the CRF form on the user terminal.
[0025] Optionally, the user terminal includes a mobile terminal and a fixed terminal.
[0026] Optionally, the generation method further includes:
[0027] In response to the user's addition operation, add the item corresponding to the addition operation to the position specified by the addition operation.
[0028] In a second aspect, an embodiment of the present application provides a device for generating a CRF form for rare diseases. The generating device is applied in a form generation system. The form generation system includes a user terminal and a server. The server is provided with a medical report database and a medical dictionary database. The generating device is configured on the server. The generating device includes:
[0029] A statistical unit, configured to respond to keywords input by the user on the user terminal, perform word frequency statistics on each medical report in the medical report database, and determine the alternative words with the highest word frequency in each medical report;
[0030] A construction unit, for each medical term in the medical dictionary and each term to be compared, constructs an (m + 1) * (n + 1) matrix according to the first string length of the medical term and the second string length of the term to be compared, where m is the first string length, n is the second string length, and the terms to be compared include the keyword and alternative terms;
[0031] A first determination unit, for sequentially initializing the matrix elements of the first row of the matrix to integers from 0 to n, sequentially initializing the matrix elements of the first column of the matrix to integers from 0 to m, and for other matrix elements Tij in the matrix, determining other matrix elements Tij according to a preset rule; where the preset rule includes: if the i-th character of the medical term is the same as the j-th character of the term to be compared, then the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical term is different from the j-th character of the term to be compared, then the matrix element Tij is equal to the sum of the minimum matrix element in the three matrix cells adjacent to the upper left, upper, and left sides of the matrix cell where the matrix element Tij is located and the value 1;
[0032] A calculation unit, after determining the matrix elements in each matrix cell of the matrix, calculates the similarity between the medical term and the term to be compared using the following formula:
[0033] S = 1 - B / A;
[0034] where B is the matrix element in the matrix cell located in the lower right corner of the matrix, and A is the maximum value of m and n;
[0035] A second determination unit, after obtaining the first similarity between each alternative term and each medical term and the second similarity between the keyword and each medical term, determines the medical term corresponding to the first similarity with the smallest difference from the target similarity among the first similarities between each alternative term and the target medical term, where the target medical term is the medical term corresponding to the maximum first similarity of each alternative term, and the target similarity is the maximum second similarity;
[0036] A generation unit, constructs a CRF form according to the headings at all levels in the target medical report corresponding to the target medical term and the terms with numerical values matched in the text content under the headings at all levels.
[0037] Optionally, when the generation unit constructs a CRF form according to the headings at all levels in the target medical report corresponding to the target medical term and the terms with numerical values matched in the text content under the headings at all levels, it includes:
[0038] After obtaining the headings at all levels in the target medical report, perform word segmentation on the headings at all levels to obtain candidate word segments included in the headings at all levels;
[0039] For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment;
[0040] Use the following formula to calculate the similarity between the candidate word segment and the preset medical field words:
[0041] S = ;
[0042] where K is the first word vector corresponding to the candidate word segment, L is the second word vector corresponding to the preset medical field word, is the norm of the first word vector, is the norm of the second word vector, and d is the weight;
[0043] For the text content corresponding to each level of heading, perform text processing on the text content to obtain the dependency relationship between the text word segments;
[0044] According to the part of speech of each text word segment, determine candidate text word segments that have a dependency relationship with the text word segments whose part of speech is a numeral from each text word segment;
[0045] Select a specified type of part of speech from the candidate text word segments as the target text word segment;
[0046] Use the candidate word segments with a similarity higher than the preset threshold as the parent items, and for each parent item, use the target text word segments included in the text content under the heading corresponding to the parent item as the child items to construct the CRF form.
[0047] Optionally, the generating device further includes:
[0048] A display unit, configured to display the CRF form on the user terminal.
[0049] Optionally, the user terminal includes a mobile terminal and a fixed terminal.
[0050] Optionally, the generating device further includes:
[0051] An adding unit, configured to respond to a user's adding operation and add the item corresponding to the adding operation to the position specified by the adding operation.
[0052] The technical solution provided by the embodiments of the present application may include the following beneficial effects:
[0053] In this application, when a CRF form for a rare disease is needed, the user can input keywords of the rare disease through the user terminal. There is a medical report database set in the server. Since the medical reports stored in the medical report database contain items that need to be filled in the CRF form corresponding to the relevant diseases, such as: basic patient information, past medical history, personal physical examination results, medical test results, imaging test results, doctor's opinions and suggestions, etc., various matters capable of establishing a CRF form can be obtained from the medical reports corresponding to the rare disease. By comparing with the medical terms included in the medical dictionary, medical reports with a relatively high relevance to the keywords can be determined, and then it is convenient to construct a CRF form related to the current rare disease. Through the above method, the user can obtain the corresponding CRF form only by inputting keywords, which helps to reduce the manual workload.
[0054] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a schematic flowchart of a method for generating a CRF form for a rare disease provided by an embodiment of the present application;
[0057] Figure 2 It is a schematic flowchart of another method for generating a CRF form for a rare disease provided by an embodiment of the present application;
[0058] Figure 3 It is a schematic structural diagram of a device for generating a CRF form for a rare disease provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some, rather than all, of the embodiments of this application. Components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of this application that is claimed, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0060] It should be noted in advance that the medical reports involved in this application are guiding medical documents that can be used to guide doctors or medical workers in treating and researching related rare diseases. For example, guiding medical documents on how to diagnose related diseases, how to treat them, related precautions, and related disease indicators. The specific medical reports stored in the medical report database can be set according to actual needs and will not be specifically limited here.
[0061] Figure 1 The figure is a schematic flowchart of a method for generating a CRF form for a rare disease provided by an embodiment of this application. The generation method is applied to a form generation system. The form generation system includes a user terminal and a server. The server is provided with a medical report database and a medical dictionary library. The generation method runs on the server. As Figure 1 shown, the generation method includes the following steps:
[0062] Step 101: In response to a keyword input by a user on the user terminal, perform a word frequency statistics on each medical report in the medical report database to determine the alternative word with the highest word frequency in each medical report.
[0063] Step 102: For each medical term and each term to be compared in the medical dictionary, construct an (m + 1) * (n + 1) matrix according to the first string length of the medical term and the second string length of the term to be compared, where m is the first string length, n is the second string length, and the terms to be compared include the keyword and the alternative words.
[0064] Step 103: Initialize the elements of the first row of the matrix to integers from 0 to n in sequence, initialize the elements of the first column of the matrix to integers from 0 to m in sequence, and for other matrix elements Tij in the matrix, determine the other matrix elements Tij according to a preset rule; wherein, the preset rule includes: if the i-th character of the medical term is the same as the j-th character of the term to be compared, then the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical term is different from the j-th character of the term to be compared, then the matrix element Tij is equal to the sum of the smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left of the matrix cell where the matrix element Tij is located and the value 1.
[0065] Step 104: After determining the matrix elements in each matrix cell of the matrix, calculate the similarity between the medical term and the term to be compared using the following Formula 1:
[0066] S = 1 - B / A; (Formula 1)
[0067] wherein, B is the matrix element in the matrix cell located at the lower right corner of the matrix, and A is the maximum value of m and n.
[0068] Step 105: After obtaining the first similarity between each alternative term and each medical term and the second similarity between the keyword and each medical term, determine the medical term corresponding to the first similarity with the smallest difference from the target similarity among the first similarities of each alternative term and the target medical term, wherein, the target medical term is the medical term corresponding to the maximum first similarity of each alternative term, and the target similarity is the maximum second similarity.
[0069] Step 106: Construct a CRF form according to the words with values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical term.
[0070] Specifically, when a user (including but not limited to: doctors, medical workers, disease researchers, etc.) needs to use a CRF form for a certain rare disease, they can enter keywords related to the rare disease in the user terminal. For example: the name of the rare disease. After the server obtains this keyword, it will perform a word frequency statistics on each medical report stored in the medical report database to determine the alternative words with the highest word frequency in each medical report. By selecting the alternative words with the highest word frequency in the medical report, the specific disease type targeted by the corresponding medical report can be determined. For a disease, there may be differences in the description methods. In order to determine the standard terms corresponding to each alternative word and the keyword, steps 102 - 104 can be used for determination. Taking a certain medical word in the medical dictionary as "kitten" and the word to be compared as "sitting" as an example, the string length of "kitten" is 6, and the string length of "sitting" is 7. Then, a 7*8 matrix is constructed. Then, the matrix elements of the first row of the 7*8 matrix are initialized to 0 - 7 in sequence, and the matrix elements of the first column are initialized to 0 - 6 in sequence. The obtained matrix is as follows:
[0071]
[0072] For other matrix elements Tij (the matrix composed of the blank matrix cells in the above matrix), they are determined in sequence from left to right and from top to bottom. The preset rule used for determination is: if the i-th character of this medical word is the same as the j-th character of this word to be compared, then the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of this medical word is not the same as the j-th character of this word to be compared, then the matrix element Tij is equal to the sum of the smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left of the matrix cell where the matrix element Tij is located and the value 1.
[0073] Taking the matrix element in the second row and second column of the above array as an example, Tij is T11, that is: the first character of the medical word and the first character of the word to be compared. The first character of the medical word is k, and the first character of the word to be compared is s. The smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left of the cell T11 is 0, and the sum value with the value 1 is 1. Therefore, the value 1 is used as the matrix element of the cell T11.
[0074] Taking the matrix element in the third row and third column of the above array as an example, Tij is T22, that is: the second character of the medical word and the second character of the word to be compared. The second character of the medical word is i, and the second character of the word to be compared is also i. The matrix cell in the upper left adjacent to the cell T22 is T11. Therefore, the matrix element in the cell T22 is 1.
[0075] By the above method, the matrix elements in the cells of the above matrix can be obtained. At this time, the matrix element in the lower right matrix cell of the above matrix can be obtained, and A is 7. Then, through Formula 1, the similarity when the medical term is "kitten" and the term to be compared is "sitting" can be obtained.
[0076] The first similarity can be used to determine the relevance of a certain alternative term to each medical term (i.e., the standard medical term). The second similarity can be used to determine the relevance of the keyword input by the user to each medical term. When the relevance of a certain alternative term to the keyword is relatively high, the difference in similarity between the two to the same medical term is generally relatively small. Therefore, in order to determine which alternative term has a relatively high relevance to the keyword input by the user, after obtaining the first similarity between each alternative term and each medical term and the second similarity between the keyword and each medical term, the medical term with the highest similarity to each alternative term (i.e., the target medical term) can be determined. At this time, each alternative term will determine a target medical term, and this target medical term is the term with the highest relevance to the corresponding medical report. In addition, the second similarity between the keyword and each medical term can be determined, and the maximum second similarity can be determined from these second similarities. After obtaining the target medical terms corresponding to each alternative term, the first similarity between each alternative term and the target medical term (subsequently referred to as the similarity to be compared) can be obtained. Then, the difference between the determined maximum second similarity and each similarity to be compared is calculated, and the target medical term corresponding to the similarity to be compared with the smallest difference is determined. This target medical term can be understood as the medical term with the highest relevance to the keyword. Since each alternative term corresponds to a target medical term, after obtaining the target medical term, the alternative term corresponding to this target medical term can be determined, and then it can be determined which medical report this alternative term comes from, so as to use this medical report as the target medical report. After obtaining the target medical report, a CRF form is constructed based on the terms with values matched in the headings at all levels and the text content corresponding to the headings at all levels in this target medical report.
[0077] Since the target medical report has guiding content for the disease corresponding to the keyword input by the user, the guiding items can be obtained through the headings at all levels and the text content corresponding to the headings at all levels in this target medical report. Thus, the items in the constructed CRF form have a relatively high degree of relevance to the disease corresponding to the keyword and are relatively complete for the user to use. Through the above method, the user only needs to input a keyword to obtain the corresponding CRF form, which helps to reduce the manual workload.
[0078] In a feasible implementation Figure 2It is a schematic flowchart of another method for generating a CRF form for a rare disease provided by an embodiment of the present application. As shown in Figure 2 In step 106, it can be achieved through the following steps:
[0079] Step 201: After obtaining the titles at all levels in the target medical report, perform word segmentation on the titles at all levels to obtain candidate word segments included in the titles at all levels.
[0080] Step 202: For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment.
[0081] Step 203: Use the following formula two to calculate the similarity between the candidate word segment and the preset medical field words:
[0082] S= ; (Formula two)
[0083] Where K is the first word vector corresponding to the candidate word segment, L is the second word vector corresponding to the preset medical field word, is the modulus length of the first word vector, is the modulus length of the second word vector, and d is the weight.
[0084] Step 204: For the text content corresponding to each title at all levels, perform text processing on the text content to obtain the dependency relationship between the text word segments and the text word segments.
[0085] Step 205: According to the part of speech of each text word segment, determine candidate text word segments that have a dependency relationship with the text word segments whose part of speech is a numeral from each text word segment.
[0086] Step 206: Select a specified type of part of speech from the candidate text word segments as the target text word segment.
[0087] Step 207: Use the candidate word segments with a similarity higher than the preset threshold as the parent items. For each parent item, use the target text word segments included in the text content under the title corresponding to the parent item as the child items to construct the CRF form.
[0088] Specifically, after obtaining the target medical report, parent items can be constructed based on the headings at all levels, and then sub-items under each parent item can be constructed according to the content under the headings at all levels, thereby constructing a CRF form. For the headings at all levels, word segmentation is first performed on the headings at all levels to obtain the candidate words included in the headings at all levels. In the medical field, the usage frequencies of different parts of speech are different. The higher the usage frequency, the higher the importance of the part of speech in the medical field. Therefore, weights can be assigned to the corresponding candidate analyses according to the distribution ratio of various parts of speech in the medical field. For example, if the distribution ratio of nouns in various parts of speech in the medical field is 30%, the assigned weight for the noun part of speech is 0.3. After obtaining the weights corresponding to each candidate word segment, the similarity between each candidate word segment and the preset medical field words is calculated using Formula 2. The higher the similarity, the higher the degree of relevance of the candidate word to the preset medical field.
[0089] After obtaining the headings at all levels, text segmentation is performed on the text content under the headings at all levels, and the dependency relationship between each text segment is determined. The dependency relationship can represent the syntactic interdependence between the word segments. Since the CRF form generally fills in content related to numerical values, after obtaining the text segments, the numerical word segments related to numerals in the text segments are determined, and then the text segments that have a dependency relationship with the numerical word segments are used as candidate text segments (the candidate text segments can be used as sub-items).
[0090] Since the sub-items are generally words of a certain part of speech or several parts of speech, such as nouns, after obtaining the candidate text segments, the parts of speech of the specified type are selected from the candidate text segments as the target text segments, and then the candidate word segments with a similarity higher than the preset threshold are used as the parent items. For each parent item, the target text segments included in the text content under the heading corresponding to the parent item are used as the sub-items to construct the CRF form, so that the node relationship expressed by the constructed CRF is the same as the table of contents relationship where, for each parent item, the target text segments included in the text content under the heading corresponding to the parent item are used as the sub-items to construct the CRF form.
[0091] It should be noted that different search words and the preset medical field words corresponding to each search word can be set in advance. After obtaining the candidate word segments, the cosine similarity is used to determine the similarity between each candidate word segment and each search word, and the search word with the highest similarity is used as the target search word for the candidate word segment. Then, the preset medical field word corresponding to the target search word is used as the preset medical field word for the candidate word segment. Or multiple preset medical field words can be set in advance according to experience. Regarding how to obtain the preset medical field words, it can be set according to actual needs and will not be specifically limited here.
[0092] In a feasible implementation scheme, after obtaining the CRF form, the CRF form is displayed on the user terminal for the user to view.
[0093] In a feasible implementation, the user terminal includes a mobile terminal and a fixed terminal.
[0094] In a feasible implementation, after obtaining the CRF form, it is also possible to respond to the user's addition operation and add the item corresponding to the addition operation to the position specified by the addition operation to meet the specific needs of the user.
[0095] Figure 3 It is a schematic structural diagram of a device for generating a CRF form for a rare disease provided by an embodiment of the present application. The generating device is applied in a form generation system. The form generation system includes a user terminal and a server. The server is provided with a medical report database and a medical dictionary database. The generating device is configured on the server. As Figure 3 shown, the generating device includes:
[0096] A statistics unit 31, configured to respond to keywords input by the user at the user terminal, perform word frequency statistics on each medical report in the medical report database, and determine alternative words with the highest word frequency in each medical report;
[0097] A construction unit 32, configured to construct an (m + 1) * (n + 1) matrix for each medical word in the medical dictionary and each word to be compared, where m is the length of the first string of the medical word, n is the length of the second string of the word to be compared, and the words to be compared include the keyword and the alternative word;
[0098] A first determination unit 33, configured to initialize the first row matrix elements of the matrix to integers from 0 to n in sequence, initialize the first column matrix elements of the matrix to integers from 0 to m in sequence, and for other matrix elements Tij in the matrix, determine other matrix elements Tij according to a preset rule; where the preset rule includes: if the i-th character of the medical word is the same as the j-th character of the word to be compared, the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical word is different from the j-th character of the word to be compared, the matrix element Tij is equal to the sum of the smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left sides of the matrix cell where the matrix element Tij is located and the value 1;
[0099] A calculation unit 34, configured to calculate the similarity between the medical word and the word to be compared by using the following formula after determining the matrix elements in each matrix cell of the matrix:
[0100] S = 1 - B / A;
[0101] Wherein, B is the matrix element in the matrix cell located at the lower right corner of the matrix, and A is the maximum value of m and n;
[0102] A second determination unit 35, configured to determine, after obtaining the first similarity between each alternative word and each medical word and the second similarity between the keyword and each medical word, the medical word corresponding to the first similarity with the smallest difference from the target similarity among the first similarities between each alternative word and the target medical word, wherein the target medical word is the medical word corresponding to the maximum first similarity of each alternative word, and the target similarity is the maximum second similarity;
[0103] A generation unit 36, configured to construct a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical word.
[0104] In a feasible implementation, when the generation unit is configured to construct a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical word, it includes:
[0105] After obtaining the titles at all levels in the target medical report, perform word segmentation on the titles at all levels to obtain candidate word segments included in the titles at all levels;
[0106] For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment;
[0107] Use the following formula to calculate the similarity between the candidate word segment and the preset medical field word:
[0108] S = ;
[0109] Wherein, K is the first word vector corresponding to the candidate word segment, L is the second word vector corresponding to the preset medical field word, is the modulus of the first word vector, is the modulus of the second word vector, and d is the weight;
[0110] For the text content corresponding to each level of title, perform text processing on the text content to obtain the dependency relationship between the text word segments and the text word segments;
[0111] According to the part of speech of each text word segment, determine candidate text word segments that have a dependency relationship with the text word segments with the part of speech of numeral from each text word segment;
[0112] Select the specified type of part of speech from the candidate text word segmentation as the target text word segmentation;
[0113] Use the candidate word segmentations with similarity higher than the preset threshold as the parent items. For each parent item, use the target text word segmentations included in the text content under the corresponding title of the parent item as the child items to construct the CRF form.
[0114] In a feasible implementation, the generating device further includes:
[0115] A display unit for displaying the CRF form on the user terminal.
[0116] In a feasible implementation, the user terminal includes a mobile terminal and a fixed terminal.
[0117] In a feasible implementation, the generating device further includes:
[0118] An adding unit for adding the item corresponding to the adding operation to the position specified by the adding operation in response to the user's adding operation.
[0119] Regarding Figure 3 the related principle description can refer to Figure 1 - Figure 2 the description of the relevant content, which will not be elaborated in detail here.
[0120] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other forms.
[0121] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0122] In addition, the functional units in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0123] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0124] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and cannot be understood as indicating or implying relative importance.
[0125] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solution of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for generating a CRF form for a rare disease, characterized in that, The generation method is applied to a form generation system, which includes a user terminal and a server. The server is provided with a medical report database and a medical dictionary library. The generation method runs on the server and includes: Responding to keywords input by a user on the user terminal, performing word frequency statistics on each medical report in the medical report database, and determining alternative words with the highest word frequencies in each medical report; For each medical term in the medical dictionary and each term to be compared, a matrix of (m + 1) * (n + 1) is constructed according to the first string length of the medical term and the second string length of the term to be compared, where m is the first string length, n is the second string length, and the terms to be compared include the keywords and alternative words; Initializing the matrix elements in the first row of the matrix to integers from 0 to n in sequence, initializing the matrix elements in the first column of the matrix to integers from 0 to m in sequence, and for other matrix elements Tij in the matrix, determining other matrix elements Tij according to a preset rule; where the preset rule includes: if the i-th character of the medical term is the same as the j-th character of the term to be compared, the matrix element Tij is equal to the matrix element in the upper left matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical term is different from the j-th character of the term to be compared, the matrix element Tij is equal to the sum of the smallest matrix element in the three matrix cells adjacent to the upper left, upper, and left sides of the matrix cell where the matrix element Tij is located and the value 1; After determining the matrix elements in each matrix cell of the matrix, calculate the similarity between the medical term and the term to be compared using the following formula: S = 1 - B / A; where B is the matrix element in the matrix cell located in the lower right corner of the matrix, and A is the maximum value of m and n; After obtaining the first similarities between each alternative word and each medical term and the second similarities between the keyword and each medical term, determine the medical term corresponding to the first similarity with the smallest difference from the target similarity among the first similarities between each alternative word and the target medical term, where the target medical term is the medical term corresponding to the maximum first similarity of each alternative word, and the target similarity is the maximum second similarity; Construct a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical term.
2. The generation method according to claim 1, wherein, The constructing a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical term includes: After obtaining the titles at all levels in the target medical report, perform word segmentation on the titles at all levels to obtain candidate word segments included in the titles at all levels; For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment; Calculate the similarity between the candidate word segmentation and the preset medical domain words using the following formula: S= ; Wherein, K is the first word vector corresponding to the candidate word segmentation, and L is the second word vector corresponding to the preset medical domain word. is the norm of the first word vector. is the norm of the second word vector, and d is the weight. For the text content corresponding to each level of heading, perform text processing on the text content to obtain the dependency relationship between text segmentations; According to the part-of-speech of each text segmentation, determine the candidate text segmentations that have a dependency relationship with the text segmentation whose part-of-speech is a numeral from each text segmentation; Select the part-of-speech of the specified type from the candidate text segmentations as the target text segmentation; Use the candidate word segmentations with a similarity higher than the preset threshold as the parent items. For each parent item, use the target text segmentations included in the text content under the heading corresponding to the parent item as the sub-items to construct the CRF form.
3. The generation method according to claim 1, characterized in that, The generation method further includes: Display the CRF form on the user terminal.
4. The generation method according to claim 1, characterized in that, The user terminal includes a mobile terminal and a fixed terminal.
5. The generation method according to claim 1, wherein The generation method further includes: In response to the user's addition operation, add the item corresponding to the addition operation to the position specified by the addition operation.
6. A generating device for a CRF form of a rare disease, characterized in that, The generation device is applied to a form generation system. The form generation system includes a user terminal and a server. The server is provided with a medical report database and a medical dictionary database. The generation device is configured on the server. The generation device includes: A statistics unit for responding to keywords input by the user on the user terminal, performing word frequency statistics on each medical report in the medical report database, and determining the alternative words with the highest word frequencies in each medical report; A construction unit for constructing an (m + 1) * (n + 1) matrix for each medical word and each word to be compared in the medical dictionary according to the first string length of the medical word and the second string length of the word to be compared, where m is the first string length and n is the second string length, and the words to be compared include the keywords and alternative words; A first determination unit for initializing the matrix elements of the first row of the matrix to integers from 0 to n in sequence, initializing the matrix elements of the first column of the matrix to integers from 0 to m in sequence, and for other matrix elements Tij in the matrix, determining other matrix elements Tij according to a preset rule; where the preset rule includes: if the i-th character of the medical word is the same as the j-th character of the word to be compared, the matrix element Tij is equal to the matrix element in the upper left corner matrix cell adjacent to the matrix cell where the matrix element Tij is located; if the i-th character of the medical word is different from the j-th character of the word to be compared, the matrix element Tij is equal to the sum of the smallest matrix element among the three matrix cells in the upper left corner, upper side, and left side adjacent to the matrix cell where the matrix element Tij is located and the value 1; A calculation unit for, after determining the matrix elements in each matrix cell of the matrix, calculating the similarity between the medical word and the word to be compared using the following formula: S = 1 - B / A; Where B is the matrix element in the matrix cell located at the lower right corner of the matrix, and A is the maximum value of m and n; A second determination unit, configured to, after obtaining the first similarity between each alternative word and each medical word and the second similarity between the keyword and each medical word, determine the medical word corresponding to the first similarity with the smallest difference from the target similarity among the first similarities between each alternative word and the target medical word, where the target medical word is the medical word corresponding to the maximum first similarity of each alternative word, and the target similarity is the maximum second similarity; A generation unit, configured to construct a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical word.
7. The generating device according to claim 6, characterized in that, When the generation unit is configured to construct a CRF form according to the words with numerical values matched in the titles at all levels and the text content corresponding to the titles at all levels in the target medical report corresponding to the target medical word, it includes: After obtaining the titles at all levels in the target medical report, perform word segmentation on the titles at all levels to obtain candidate word segments included in the titles at all levels; For each candidate word segment, after determining the part of speech of the candidate word segment, determine the distribution ratio of the part of speech in various parts of speech in the medical field as the weight of the candidate word segment; Use the following formula to calculate the similarity between the candidate word segment and a preset medical field word: S= ; Wherein, K is the first word vector corresponding to the candidate word segmentation, and L is the second word vector corresponding to the preset medical field word. is the norm of the first word vector, is the norm of the second word vector, and d is the weight. For the text content corresponding to each level of title, perform text processing on the text content to obtain the dependency relationship between the text word segments and the text word segments; According to the parts of speech of the text word segments, determine candidate text word segments that have a dependency relationship with the text word segments with the part of speech being a numeral from the text word segments; Select a specified type of part of speech from the candidate text word segments as the target text word segment; Use the candidate word segments with a similarity higher than a preset threshold as the parent items, and for each parent item, use the target text word segments included in the text content corresponding to the title of the parent item as the child items to construct the CRF form.
8. The generating device according to claim 6, wherein The generation device further includes: A display unit, configured to display the CRF form on the user terminal.
9. The generating device according to claim 6, wherein The user terminal includes a mobile terminal and a fixed terminal.
10. The generating device according to claim 6, characterized in that The generation device further includes: An addition unit, configured to, in response to an addition operation of the user, add the item corresponding to the addition operation to the position specified by the addition operation.
Citation Information
Patent Citations
Method and device for automatically obtaining medical record template and storage medium
CN111696640A
Medical record text similarity retrieval method and system and computer equipment
CN111949759A