Text abbreviation data processing method and device

By constructing a database of abbreviation full-term terms and using machine learning models to identify and complete abbreviations in text, the problem of insufficient semantic information in abbreviations is solved, the efficiency of text parsing and knowledge extraction is improved, the semantic association of text is enhanced, and in-depth knowledge mining is promoted.

CN115936010BActive Publication Date: 2026-03-31NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, abbreviations contain relatively little semantic information, which affects text parsing and knowledge extraction, resulting in low efficiency in recognizing and understanding abbreviation data in text.

Method used

By constructing a database of abbreviation and full term pairs, and using a pre-trained machine learning model to identify and complete abbreviations in text, a correspondence between abbreviations and full terms is established, thereby improving the efficiency of identifying and understanding abbreviation data in text.

Benefits of technology

It enhances the semantic connections between texts, eliminates the difficulties in parsing the "rich" semantics and relationships of texts caused by inconsistent terminology, improves the efficiency of identifying and understanding abbreviated data in texts, and provides possibilities for deep knowledge mining of the whole text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115936010B_ABST
    Figure CN115936010B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a text abbreviation data processing method and device. The method comprises: obtaining a reference text set belonging to a target knowledge field, the reference text set comprising at least one reference text; identifying abbreviation and full name term pairs distributed in each reference text through a pre-trained abbreviation and full name term pair identification model, the abbreviation and full name term pair comprising an abbreviation term and a full name term corresponding to the abbreviation term; based on the identified abbreviation and full name term pairs, constructing an abbreviation and full name term pair library, the abbreviation and full name term pair library recording the correspondence between the abbreviation term and at least one full name term; obtaining a text to be processed belonging to the target knowledge field, and based on the abbreviation and full name term pair library, completing the full name term for the abbreviation term independently distributed in the text to be processed. The technical solution of the embodiments of the present application can improve the efficiency of identifying and understanding the abbreviation data in the text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and artificial intelligence technology, and more specifically, to a method and apparatus for processing text abbreviation data. Background Technology

[0002] In academic texts, abbreviations are often used to replace frequently occurring long terms, thus avoiding the reading comprehension difficulties caused by long and complex terms. However, this also introduces problems: abbreviations carry less semantic information, hindering text semantic representation, affecting text parsing and knowledge extraction, and reducing the efficiency of recognizing and understanding abbreviations in text. Therefore, improving the efficiency of recognizing and understanding abbreviations in text is an urgent technical problem to be solved. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, computer program product or computer program, computer-readable medium and electronic device for processing abbreviated text data, which can at least to some extent improve the efficiency of recognizing and understanding abbreviated data in text.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, a method for processing text abbreviation data is provided. The method includes: acquiring a set of reference texts belonging to a target knowledge domain, the set of reference texts including at least one reference text; identifying pairs of abbreviations and full names distributed in each reference text using a pre-trained abbreviation full name term pair recognition model, the pairs of abbreviations and full names including abbreviations and corresponding full names; constructing a library of abbreviation and full name terms based on the identified pairs of abbreviations and full names, the library recording the correspondence between abbreviations and at least one full name; acquiring text to be processed belonging to the target knowledge domain, and, based on the library of abbreviation and full name terms, completing the full names for abbreviations independently distributed in the text to be processed.

[0006] According to one aspect of the embodiments of this application, a text abbreviation data processing apparatus is provided, the apparatus comprising: a first acquisition unit, configured to acquire a set of reference texts belonging to a target knowledge domain, the set of reference texts including at least one reference text; an identification unit, configured to identify pairs of abbreviations and full names distributed in each reference text using a pre-trained abbreviation-full-name term pair identification model, the pairs of abbreviations and full names including abbreviations and corresponding full names; a construction unit, configured to construct an abbreviation-full-name term pair library based on the identified pairs of abbreviations and full names, the library recording the correspondence between abbreviations and at least one full name; and a second acquisition unit, configured to acquire text to be processed belonging to the target knowledge domain, and, based on the abbreviation-full-name term pair library, complete the full names of abbreviations independently distributed in the text to be processed.

[0007] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text abbreviation data processing method described in the above embodiments.

[0008] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the text abbreviation data processing method as described in the above embodiments.

[0009] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the text abbreviation data processing method as described in the above embodiments.

[0010] In some embodiments of this application, the technical solutions provided involve constructing a database of abbreviation-full-name term pairs that records the correspondence between abbreviations and at least one full-name term by identifying abbreviation-full-name term pairs. Based on this database, full-name terms are added to abbreviations that are independently distributed in the text to be processed. This avoids situations where abbreviations themselves carry little semantic information, which could affect text parsing and knowledge extraction. It helps to enhance the semantic connections between texts, eliminates problems such as the difficulty in parsing the "rich" semantics and connections of texts due to incomplete terminology, improves the efficiency of identifying and understanding abbreviation data in texts, and provides the possibility for deep knowledge mining of the entire text.

[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0013] Figure 1 A flowchart of a text abbreviation data processing method according to an embodiment of this application is shown;

[0014] Figure 2 A block diagram of a text abbreviation data processing apparatus according to an embodiment of the present application is shown;

[0015] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0016] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0017] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0018] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0019] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0020] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0021] It should also be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such uses of these terms can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described.

[0022] Before elaborating on the text abbreviation data processing scheme in this application, the relevant concepts involved in this application will be briefly introduced below.

[0023] In academic literature, long, recurring terms are often replaced with abbreviations. Using abbreviations concisely expresses the intended meaning and helps to accurately grasp the article's structure. In this application, long terms are defined as full terms, and abbreviations as abbreviated terms. For example, "HIV" stands for "Human Immunodeficiency Virus," "Pro" stands for "Protein," and "TLR" stands for "Toll-like receptor." The word formation characteristics of abbreviations are often irregular, generally falling into three categories: abbreviations, acronyms, and alphanumeric combinations. In academic literature, abbreviations typically appear in pairs with their corresponding full terms upon their first appearance, and thereafter the abbreviation replaces the full term.

[0024] It is evident that abbreviations are more common than full terms, which poses a significant challenge to the identification of full abbreviations. Finding the correct extension boundaries of abbreviations is the key to identifying abbreviations and their corresponding full terms.

[0025] The implementation details of the technical solutions in the embodiments of this application are described in detail below:

[0026] Figure 1 A flowchart of a text abbreviation data processing method according to an embodiment of this application is shown. This text abbreviation data processing method can be executed by a device with computational processing capabilities. (Refer to...) Figure 1 As shown, this text abbreviation data processing method includes at least steps 110 to 170, which are detailed below:

[0027] In step 110, a set of reference texts belonging to the target knowledge domain is obtained, the set of reference texts including at least one reference text.

[0028] In this application, the target knowledge domain can refer to a specific professional field, such as the medical field, the military field, the artificial intelligence field, the political field, etc., and this application does not impose further limitations on this. Furthermore, the text within the target knowledge domain can refer to multiple professional texts within a specific professional field, such as English papers and journal articles in the medical field.

[0029] For example, in the field of virology within the medical specialty, data association technologies such as semantic web can be used to associate and organize potential knowledge in virology research. Specifically, the reference text corpus in the virology knowledge domain can be obtained from the PubMed database. By referring to the list of virology journals in the Journal Citation Reports section of Web of Science, open access datasets in XML format can be downloaded in batches through the PMC FTP service server to obtain several full-text articles in XML format. The XML files can then be parsed using the xmltodict package in Python to obtain a set of reference texts.

[0030] Continue to refer to Figure 1 In step 130, the abbreviation and full name term pair recognition model is used to identify the abbreviation and full name term pairs distributed in each reference text. The abbreviation and full name term pair includes the abbreviation term and the full name term corresponding to the abbreviation term.

[0031] In one embodiment of this application, the abbreviation term recognition model and the abbreviation full term pair recognition model can be trained according to the following steps 121 to 123:

[0032] Step 121: Obtain a training text set and a verification text set belonging to the target knowledge domain. The training text set includes at least one training text, and the verification text set includes at least one verification text. Both the training text and the verification text include abbreviation full name term pairs and abbreviation terms, as well as annotation tags for abbreviation full name term pairs and abbreviation terms.

[0033] Step 122: Based on the training texts in the training text set, train the pre-built machine learning model to obtain at least one candidate abbreviation term recognition model and at least one candidate abbreviation full name term pair recognition model.

[0034] Step 123: Based on the verification text in the verification text set, select the abbreviation term recognition model and the abbreviation full term pair recognition model from at least one candidate abbreviation term recognition model and at least one candidate abbreviation full term pair recognition model, respectively.

[0035] In one specific implementation of this embodiment, the training text set and the validation text set can also be obtained from the PubMed database. After obtaining the training text and validation text, they are further segmented using "." and "?" as sentence breaks to form a first corpus to be labeled, used for abbreviation term recognition. Sentences containing bracket pairs are further filtered using regular expressions to form a second corpus to be labeled, used for abbreviation full term word pair recognition.

[0036] Furthermore, for the first corpus to be labeled, labels are added for abbreviations; for the second corpus to be labeled, labels are added for full abbreviations. To avoid incorrect labeling, single letters (m, gn, p, G, M, F, S, etc.) and letter combinations where the first letter is capitalized and the following letters are all lowercase (Cre, Flu, Can, Mab, This, etc.) are removed from the corpus.

[0037] In this application, the base model used for the machine learning model can be the BB-BLC model (i.e., the model used to train abbreviation recognition). Instead of the BERT general-domain pre-trained model, the BioBert model can be used, selecting the deep learning framework from bioBert-base-cased-v1.2-pytorch.

[0038] In the BB-BLC model, the output of the BERT layer transforms each word of the sentence into three embeddeddings, which are then summed: token embeddings, segment embeddings, and position embeddings. The summed sequence vector is then input into a true bidirectional Transformer attention mechanism for feature extraction, and a fine-tuning mode is used to obtain a sequence vector rich in contextual semantic information. The BiLSTM layer is responsible for obtaining feature vectors, used to bidirectionally encode the BERT input vector to represent context-related semantic information. Using BiLSTM can better capture long-distance order dependencies. The CRF layer effectively expresses label transition relationships, considering the label information of preceding and following characters during decoding. The CRF layer's role is to obtain the final term labels; its input is a given set of observation sequences X, and it calculates the probability of the corresponding state sequence Y. For each possible state sequence, its score s is calculated, and the sequence with the highest score is taken as the recognition result.

[0039] Based on this, a hybrid model combining rules and deep learning, BBF-BLC-R (i.e., a model for training abbreviation full term pair recognition), is proposed to improve the recognition and correction of abbreviation full term pairs. This model mainly uses the BBF-BLC model to discover candidate abbreviation full term pairs by fusing external features, and then uses the R model to correct the extracted abbreviation full term pairs through user-defined rules.

[0040] In this application, external features mainly include part-of-speech features, boundary word features, word formation features, stem features, symbol features, and prefix word features.

[0041] In this application, the rules for modifying abbreviation-full term pairs are mainly divided into two steps: First, expanding the scope of the definition based on the initial letter structure of the abbreviation-full term pair. Second, constraining the scope of expansion through the subset relationship and order relationship between the letters of the abbreviation and the full name.

[0042] In this application, a machine learning model is trained by training text to enable the model to recognize abbreviations and full abbreviation pairs in the text, resulting in at least one candidate abbreviation recognition model and at least one candidate full abbreviation pair recognition model. In order to obtain the model with the best recognition ability, the abbreviation recognition model and the full abbreviation pair recognition model can be selected from at least one candidate abbreviation recognition model and at least one candidate full abbreviation pair recognition model respectively by verifying the text.

[0043] Continue to refer to Figure 1In step 150, based on the identified abbreviation full term pairs, an abbreviation full term pair library is constructed, which records the correspondence between abbreviations and at least one full term.

[0044] In such Figure 1 In one embodiment of step 150 shown, based on the identified abbreviation full name term pair, a library of abbreviation full name term pairs is constructed, which can be performed according to the following steps 151 to 152:

[0045] Step 151: For each target abbreviation full name term pair, query whether the target abbreviation term in the target abbreviation full name term pair exists in the abbreviation full name term pair library. The target abbreviation full name term pair is any one of the identified abbreviation full name term pairs.

[0046] Step 152: If the target abbreviation term in the target abbreviation term pair does not exist in the abbreviation full name term pair library, then the target abbreviation full name term pair is added to the abbreviation full name term pair library.

[0047] In this embodiment, steps 153 to 156 may also be performed:

[0048] Step 153: If the target abbreviation term in the target abbreviation term pair exists in the abbreviation full name term pair library, then the full name term in the target abbreviation full name term pair is taken as the first full name term, and the full name term in the abbreviation full name term pair library corresponding to the target abbreviation term is taken as the second full name term.

[0049] Step 154: Calculate the comprehensive similarity between the first and second universal terms.

[0050] Step 155: If the overall similarity does not exceed the similarity threshold, then the target abbreviation full name term pair is added to the abbreviation full name term pair library.

[0051] Step 156: If the overall similarity exceeds the similarity threshold, then the target abbreviation full name term pair will not be included in the abbreviation full name term pair library.

[0052] In one embodiment of step 154 ​​above, calculating the comprehensive similarity between the first full term and the second full term can be performed according to steps 1541 to 1543 as follows:

[0053] Step 1541: Calculate the semantic similarity between the first universal term and the second universal term.

[0054] Step 1542: Calculate the structural similarity between the first and second universal terms.

[0055] Step 1543: Calculate the comprehensive similarity based on the semantic similarity and the structural similarity using a linear weighted method.

[0056] In this application, since the same full term may have multiple variations, for example, the full term corresponding to the abbreviation "MC" can be "Microphone Controller" or "Move the Crowd". Therefore, when querying whether the target abbreviation exists in the abbreviation full term pair library, a full term alignment strategy based on semantic similarity and a full term alignment strategy based on structural similarity can be used to determine whether the meanings of different forms of the full term corresponding to the abbreviation are the same. If they are different, the abbreviation full term pair needs to be included in the abbreviation full term pair library; if they are the same, it is not necessary to include the abbreviation full term pair in the abbreviation full term pair library.

[0057] In this application, corresponding solutions are designed for different alignment strategies. A linearly weighted hybrid method can be used to integrate the two similarities to obtain the final comprehensive similarity of full-name terms. The alignment of abbreviation terms is obtained by summarizing and organizing the results of the full-name term alignment. Specifically, a set of abbreviation full-name term pairs {{L} can be defined. i ,S i}, {L j ,S j}, ..., {L n S n}}, where {L i ,S i} represents the i-th universal term L i The abbreviation is S i , {L j ,S j} represents the j-th universal term L j The abbreviation is S j .

[0058] Specifically, custom rules can be constructed to determine the abbreviation of the term S. i and S j Are the structures similar? The custom rules are as follows:

[0059] a. Remove punctuation marks from abbreviations, including: "(", ")", "-", "''", ",", ".", " / ", "%" and "".

[0060] b. If the last character of the abbreviation is a lowercase "s", then remove the "s".

[0061] Under the constraints of the above rules, if the abbreviation S i and S j Similar in structure, calculate the full term L i and L j The similarity.

[0062] Furthermore, the basic idea of ​​the semantic similarity-based full term alignment strategy is to use BioBert vectors to calculate the full term L. i and L j semantic similarity Sim sem (L i, L j ).

[0063] The basic idea of ​​the full-term alignment strategy based on structural similarity is as follows: First, to reduce the variation caused by meaningless stop words, stop words are uniformly removed from all full-term terms; second, to mitigate the variation caused by changes in morphology such as singular / plural and tense, morphological restoration is uniformly performed on all full-term terms; finally, the Jaccard fuzzy matching algorithm is used to calculate the full-term term alignment. i and l j Structural similarity Sim str (L i ,L).

[0064] In this application, a linearly weighted hybrid strategy can be used to calculate the full term l. i and l j The overall similarity Sim(L) i L j As shown in formula (1).

[0065] Sim(L i ,L j )=αSim sem (L i ,L j )+βSim str (L i ,L j (1)

[0066] Where α and β are adjustable parameters.

[0067] Set a threshold γ, and Sim(L) i L j The full term L of ) > γ j Considered as a full term L i Variations of . At this point, the alignment of the full term is complete. If the full term L i and L j Alignment, then the abbreviation S i and S jIt also aligns automatically. Based on this, aligned full terms are treated as full terms with the same meaning, and aligned abbreviations are treated as abbreviations with the same meaning.

[0068] In this application, based on the results of abbreviation term recognition, abbreviation-full term pair recognition, and abbreviation-full term alignment, a standardized abbreviation-full term pair database construction rule is further designed according to the correspondence between the number of abbreviation terms and full terms. Among them, one-to-one means that the same abbreviation term corresponds to the same full term, one-to-two means that the same abbreviation term corresponds to two full terms, one-to-three means that the same abbreviation term corresponds to three full terms, and one-to-many means that the same abbreviation term corresponds to multiple full terms.

[0069] Based on this, the standardized abbreviation full name term pair library can be divided into: a general abbreviation full name term pair library and a general abbreviation full name term pair library. Among them, abbreviation full name term pairs that conform to a one-to-one relationship are included in the general abbreviation full name term pair library, while others are included in the general abbreviation full name term pair library.

[0070] Continue to refer to Figure 1 In step 170, the text to be processed belonging to the target knowledge domain is obtained, and the full names of the abbreviations and full names are completed for the abbreviations and full names that are independently distributed in the text to be processed, based on the abbreviation and full name term pair library.

[0071] In such Figure 1 In one embodiment of step 170 shown, based on the abbreviation full term pair library, to complete the full terms for abbreviations independently distributed in the text to be processed, the following steps 171 to 173 can be performed:

[0072] Step 171: Identify abbreviations that are independently distributed in the text to be processed using a pre-trained abbreviation recognition model, and use them as abbreviations to be completed.

[0073] Step 172: Search the abbreviation full term database for the full term corresponding to the abbreviation to be completed, and use it as a candidate full term.

[0074] Step 173: Based on the candidate full name terms, determine the target full name term, and use the target full name term as the full name term corresponding to the abbreviation term to be completed to complete the full name term.

[0075] In this application, the text to be processed refers to the text in which the full names of abbreviations distributed independently need to be completed. In some practical application scenarios, if some abbreviations in the text to be processed do not have their abbreviation definitions extended through the text features of the full abbreviation pairs, it will be impossible to correctly identify the full names corresponding to the abbreviations based on the text features of the full abbreviation pairs. Therefore, by completing the full names of the abbreviations in the text to be processed, the deeper information in the text can be fully explored. At the same time, it can also improve the efficiency of identifying and understanding the abbreviations in the text to be processed. For example, in translation scenarios, based on the full names of the abbreviations in the text to be processed, the efficiency and accuracy of text translation can be greatly improved.

[0076] Before step 171 above, that is, before identifying the abbreviations independently distributed in the text to be processed by the pre-trained abbreviation recognition model, steps 161 to 162 can also be performed as follows:

[0077] Step 161: Identify the pairs of abbreviations and full terms distributed in the text to be processed using the abbreviation and full term pair recognition model.

[0078] Step 162: Add the abbreviation full name term pairs distributed in the text to be processed to the abbreviation full name term pair library.

[0079] In this application, since the pre-built abbreviation full name term pair library cannot exhaust all abbreviation full name term pairs in the target knowledge domain, the abbreviation full name term pairs distributed in the text to be processed are included in the abbreviation full name term pair library. This is beneficial to enrich the abbreviation full name term pairs in the abbreviation full name term pair library and provide stronger support for the subsequent full name term completion of the abbreviation terms in the text to be processed.

[0080] In step 173 above, the target full term is determined based on the candidate full term, which can be performed according to steps 1731 to 1732 as follows:

[0081] Step 1731: If the number of candidate full names is one, then the candidate full name is determined as the target full name.

[0082] Step 1732: If there are multiple candidate full-name terms, then select the full-name term that matches the semantic features of the text to be processed from the multiple candidate full-name terms as the target full-name term.

[0083] In one embodiment of this application, step 1721 may also be performed:

[0084] Step 1721: If no full term corresponding to the abbreviation to be completed is found in the abbreviation full term pair database, the full term is completed for the abbreviation independently distributed in the text to be processed by manual term completion. Based on the abbreviation to be completed and the full term completed by manual term completion, an abbreviation full term pair is constructed and the constructed abbreviation full term pair is included in the abbreviation full term pair database.

[0085] If multiple candidate full names of the abbreviation are found in the abbreviation full name term pair database (the meanings of the multiple candidate full names of the same abbreviation in the term pair database may differ), the specific meaning of the abbreviation in the current context can be inferred by considering the contextual semantic information of the abbreviation in the text to be processed. Then, the full name term that matches the semantic features of the text to be processed can be selected from the multiple candidate full names. In other words, the same abbreviation expresses different semantic information in different contexts. The abbreviation term word vectors can be generated using BioBert.

[0086] Specifically, taking a single text to be processed as a unit, a set of abbreviations to be completed is defined as {S}. 11 ,S 21 ,...,S ij}, where S ij S represents the j-th abbreviation in the i-th text to be processed. ij The full term L corresponding to it was not identified. ij The dataset to be completed {{SeS} 11}, {SeS 21},...,{SeS ij}}, where SeS ij S represents the j-th abbreviation in the i-th text to be processed. ij The sentence set and SeS ij The complete set of abbreviations and full terms has been added. 11 L′ 11},{S′ 21 L′ 21},...,{S′ mn L′ mn}}, where S′ mn L′ mn S′ represents the nth abbreviation in the m-th text to be processed. mn The corresponding full term is L′ mn The dataset {{Se′S′} has been completed. 11},{SeS′ 21},...,{Se′S′ mn}}, where Se′S′ mnS′ represents the nth abbreviation in the m-th text to be processed. mn The sentence set and Se′S′ mn .

[0087] First, based on the constructed abbreviation full term pair library, determine the abbreviation term S to be completed. ij If a term belongs to the general abbreviation full name database, and a match is found, then these abbreviation terms are directly mapped to their full abbreviations.

[0088] Secondly, based on the constructed abbreviation full term pair library, determine the abbreviation term S to be completed. ij Does it belong to the abbreviation full name general word pair library? If the match is successful, find the dataset to be completed, SeS. ij and the completed dataset Se′S′ mn For all sentences containing the same abbreviation, obtain the different semantic vectors of the same abbreviation in different sentences according to different contexts, and calculate S. ij and S′ mn Similarity score sim .

[0089] Then, set the threshold δ to set the score. sim The abbreviation S′ is included in the completed dataset of >δ. mn The corresponding full term L′ mn The abbreviation term S ij Candidate full names of terms.

[0090] Finally, the statistical candidate full term set L′ mn The frequency of occurrence of identical full-name terms in Chinese will be used to calculate the average similarity score. sim The highest universal term L′ mn As the abbreviation term in this document, S ij The most appropriate full term L ij .

[0091] If the abbreviation to be completed does not belong to either the general abbreviation full name database or the general abbreviation full name database, then manual assistance is required to complete the corresponding full name.

[0092] In this application, by identifying abbreviation-full term pairs, a library of abbreviation-full term pairs is constructed, recording the correspondence between abbreviations and at least one full term. Based on this library, full terms are added to abbreviations that are independently distributed in the text to be processed. This avoids the situation where abbreviations themselves carry little semantic information, which affects text parsing and knowledge extraction. It helps to enhance the semantic connections between texts, eliminates the difficulties in parsing the "rich" semantics and connections of texts caused by incomplete terminology, improves the efficiency of identifying and understanding abbreviation data in texts, and provides the possibility for deep knowledge mining of the whole text.

[0093] The following describes an apparatus embodiment of this application, which can be used to execute the text abbreviation data processing method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the text abbreviation data processing method described above.

[0094] Figure 2 A block diagram of a text abbreviation data processing apparatus according to an embodiment of this application is shown.

[0095] Reference Figure 2 As shown, a text abbreviation data processing apparatus 200 according to an embodiment of this application includes: a first acquisition unit 201, an identification unit 202, a construction unit 203, and a second acquisition unit 204.

[0096] The system comprises the following components: a first acquisition unit 201, which acquires a set of reference texts belonging to the target knowledge domain, the set of reference texts including at least one reference text; an identification unit 202, which identifies pairs of abbreviations and full names of terms distributed in each reference text using a pre-trained abbreviation and full name term pair recognition model, the pairs of abbreviations and full names of terms including abbreviations and corresponding full names of terms; a construction unit 203, which constructs a library of abbreviation and full name term pairs based on the identified pairs of abbreviations and full names of terms, the library recording the correspondence between abbreviations and at least one full name of term; and a second acquisition unit 204, which acquires text to be processed belonging to the target knowledge domain and, based on the library of abbreviation and full name term pairs, completes the full names of abbreviations independently distributed in the text to be processed.

[0097] In some embodiments of this application, based on the foregoing scheme, the construction unit 203 is configured to: for each target abbreviation full name term pair, query whether the target abbreviation term in the target abbreviation full name term pair exists in the abbreviation full name term pair library, wherein the target abbreviation full name term pair is any one of the identified abbreviation full name term pairs; if the target abbreviation term in the target abbreviation full name term pair does not exist in the abbreviation full name term pair library, then the target abbreviation full name term pair is included in the abbreviation full name term pair library.

[0098] In some embodiments of this application, based on the foregoing scheme, the construction unit 203 is further configured to: if the target abbreviation term in the target abbreviation term pair exists in the abbreviation full name term pair library, then the full name term in the target abbreviation full name term pair is taken as the first full name term, and the full name term in the abbreviation full name term pair library corresponding to the target abbreviation term is taken as the second full name term; calculate the comprehensive similarity between the first full name term and the second full name term; if the comprehensive similarity does not exceed the similarity threshold, then the target abbreviation full name term pair is included in the abbreviation full name term pair library; if the comprehensive similarity exceeds the similarity threshold, then the target abbreviation full name term pair is not included in the abbreviation full name term pair library.

[0099] In some embodiments of this application, based on the foregoing scheme, the construction unit 203 is further configured to: calculate the semantic similarity between the first full-name term and the second full-name term; calculate the structural similarity between the first full-name term and the second full-name term; and calculate the comprehensive similarity based on the semantic similarity and the structural similarity using a linear weighted method.

[0100] In some embodiments of this application, based on the foregoing scheme, the second acquisition unit 204 is configured to: identify abbreviations independently distributed in the text to be processed through a pre-trained abbreviation recognition model, as abbreviations to be completed; query the full-name term corresponding to the abbreviation to be completed in the abbreviation full-name term pair library, as candidate full-name terms; determine the target full-name term based on the candidate full-name term, and complete the full-name term by using the target full-name term as the full-name term corresponding to the abbreviation to be completed.

[0101] In some embodiments of this application, based on the foregoing scheme, the construction unit 203 is further configured to: identify abbreviations independently distributed in the text to be processed by the pre-trained abbreviation recognition model, identify abbreviation full name term pairs distributed in the text to be processed by the abbreviation full name term pair recognition model; and include the abbreviation full name term pairs distributed in the text to be processed in the abbreviation full name term pair library.

[0102] In some embodiments of this application, based on the foregoing scheme, the second acquisition unit 204 is configured to: if the number of candidate full-name terms is one, then determine the candidate full-name term as the target full-name term; if the number of candidate full-name terms is multiple, then select the full-name term that matches the semantic features of the text to be processed from the multiple candidate full-name terms as the target full-name term.

[0103] In some embodiments of this application, based on the foregoing scheme, the construction unit 203 is further configured to: if no full term corresponding to the abbreviation to be completed is found in the abbreviation full term pair library, complete the full term for the abbreviation independently distributed in the text to be processed by manual term completion, and construct abbreviation full term pairs based on the abbreviation to be completed and the full term completed by manual term completion, so as to include the constructed abbreviation full term pairs in the abbreviation full term pair library.

[0104] In some embodiments of this application, based on the foregoing scheme, the apparatus further includes: a training unit, configured to acquire a training text set and a verification text set belonging to the target knowledge domain, wherein the training text set includes at least one training text, the verification text set includes at least one verification text, and both the training text and the verification text include abbreviation full name term pairs and abbreviation terms, as well as annotation labels for the abbreviation full name term pairs and abbreviation terms; based on the training text in the training text set, a pre-built machine learning model is trained to obtain at least one candidate abbreviation term recognition model and at least one candidate abbreviation full name term pair recognition model; based on the verification text in the verification text set, the abbreviation term recognition model and the abbreviation full name term pair recognition model are selected respectively from the at least one candidate abbreviation term recognition model and the at least one candidate abbreviation full name term pair recognition model.

[0105] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0106] It should be noted that, Figure 3 The computer system 300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0107] like Figure 3As shown, the computer system 300 includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0108] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0109] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs various functions defined in the system of this application.

[0110] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0112] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0113] In another aspect, this application also provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the text abbreviation data processing method described in the above embodiments.

[0114] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the text abbreviation data processing method described in the above embodiments.

[0115] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0116] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0117] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0118] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method of processing text abbreviation data, characterized by, The method comprises: acquiring a reference text set belonging to a target knowledge field, the reference text set comprising at least one reference text; identifying abbreviation-full term pairs distributed in each reference text through a pre-trained abbreviation-full term pair identification model, the abbreviation-full term pair comprising an abbreviation term and a full term corresponding to the abbreviation term; the abbreviation-full term pair identification model is trained based on a BBF-BLC-R model, wherein external features and an R model are fused in the BBF-BLC-R model; the external features comprise part-of-speech features, boundary word features, word formation features, stem features, symbol features and prefix features; the R model corrects the extracted abbreviation-full term pair through user-defined rules; the user-defined rules comprise: expanding the definition range according to the initial letter structure features of the abbreviation-full term pair; and constraining the expanded range through the subset relationship and order relationship of the letters of the abbreviation and the full term; for each target abbreviation-full term pair, querying whether the target abbreviation term in the target abbreviation-full term pair exists in an abbreviation-full term pair library, the abbreviation-full term pair library recording the correspondence between an abbreviation term and at least one full term, the target abbreviation-full term pair being any one of the identified abbreviation-full term pairs; if the target abbreviation term in the target abbreviation-full term pair does not exist in the abbreviation-full term pair library, listing the target abbreviation-full term pair in the abbreviation-full term pair library; if the target abbreviation term in the target abbreviation-full term pair exists in the abbreviation-full term pair library, taking the full term in the target abbreviation-full term pair as a first full term, taking the full term corresponding to the target abbreviation term in the abbreviation-full term pair library as a second full term, calculating the semantic similarity of the first full term and the second full term, calculating the structural similarity of the first full term and the second full term, calculating the comprehensive similarity in a linear weighting manner based on the semantic similarity and the structural similarity, listing the target abbreviation-full term pair in the abbreviation-full term pair library if the comprehensive similarity does not exceed a similarity threshold, and not listing the target abbreviation-full term pair in the abbreviation-full term pair library if the comprehensive similarity exceeds the similarity threshold; acquiring a to-be-processed text belonging to a target knowledge field, and identifying an abbreviation term independently distributed in the to-be-processed text through a pre-trained abbreviation term identification model as a to-be-completed abbreviation term; querying the full term corresponding to the to-be-completed abbreviation term in the abbreviation-full term pair library as a candidate full term; determine a target full-term based on the candidate full-terms, and perform full-term completion on the target full-term as the full-term corresponding to the abbreviation to be completed; if the number of the candidate full-terms is one, the candidate full-term is determined as the target full-term; if the number of the candidate full-terms is more than one, a full-term matching the semantic feature of the text to be processed is selected from the candidate full-terms as the target full-term.

2. The method of claim 1, wherein, Before identifying the abbreviation term independently distributed in the text to be processed by the pre-trained abbreviation term recognition model, the method further comprises: identifying the abbreviation-full-term pair distributed in the text to be processed by the abbreviation-full-term pair recognition model; listing the abbreviation-full-term pair distributed in the text to be processed in the abbreviation-full-term pair library.

3. The method of claim 1, wherein, The method further comprises: if the full-term corresponding to the abbreviation to be completed is not found in the abbreviation-full-term pair library, performing full-term completion on the abbreviation term independently distributed in the text to be processed by artificial term completion, and constructing the abbreviation-full-term pair based on the abbreviation to be completed and the full-term completed by artificial term completion, so as to list the constructed abbreviation-full-term pair in the abbreviation-full-term pair library.

4. The method according to any one of claims 1 to 3, characterized in that, The abbreviation term recognition model and the abbreviation-full-term pair recognition model are trained according to the following steps: obtain a training text set and a verification text set belonging to a target knowledge field, the training text set includes at least one training text, the verification text set includes at least one verification text, and the training text and the verification text both include an abbreviation-full-term pair and an abbreviation term, as well as a labeled label of the abbreviation-full-term pair and the abbreviation term; train a pre-constructed machine learning model based on the training text in the training text set, to obtain at least one candidate abbreviation term recognition model and at least one candidate abbreviation-full-term pair recognition model; select the abbreviation term recognition model and the abbreviation-full-term pair recognition model from the at least one candidate abbreviation term recognition model and the at least one candidate abbreviation-full-term pair recognition model respectively based on the verification text in the verification text set.

5. A text abbreviation data processing apparatus characterized by comprising: The device comprises: a first obtaining unit configured to obtain a reference text set belonging to a target knowledge field, the reference text set including at least one reference text; The identifying unit is configured to identify an abbreviation-full form term pair, which comprises an abbreviation term and a full form term corresponding to the abbreviation term, in each reference text by using a pre-trained abbreviation-full form term pair identification model. The abbreviation-full form term pair identification model is trained based on a BBF-BLC-R model, in which external features and an R model are fused. The external features include part-of-speech features, boundary word features, word formation features, stem features, symbol features and prefix word features. The R model corrects the extracted abbreviation-full form term pair by using user-defined rules. The user-defined rules include: expanding the definition range according to the initial letter structure features of the abbreviation-full form term pair; and constraining the expanded range by using the subset relationship and the order relationship between the letters of the abbreviation and the full form. The constructing unit is configured to, for each target abbreviation-full form term pair, query whether a target abbreviation term in the target abbreviation-full form term pair exists in an abbreviation-full form term pair library, the abbreviation-full form term pair library recording the correspondence between an abbreviation term and at least one full form term, the target abbreviation-full form term pair being any one of the identified abbreviation-full form term pairs. If the target abbreviation term in the target abbreviation-full form term pair does not exist in the abbreviation-full form term pair library, the target abbreviation-full form term pair is listed in the abbreviation-full form term pair library. If the target abbreviation term in the target abbreviation-full form term pair exists in the abbreviation-full form term pair library, a full form term in the target abbreviation-full form term pair is taken as a first full form term, and a full form term corresponding to the target abbreviation term in the abbreviation-full form term pair library is taken as a second full form term. The semantic similarity of the first full form term and the second full form term is calculated. The structural similarity of the first full form term and the second full form term is calculated. The comprehensive similarity is calculated by using linear weighting based on the semantic similarity and the structural similarity. If the comprehensive similarity does not exceed a similarity threshold, the target abbreviation-full form term pair is listed in the abbreviation-full form term pair library. If the comprehensive similarity exceeds the similarity threshold, the target abbreviation-full form term pair is not listed in the abbreviation-full form term pair library. The second acquisition unit is configured to acquire a to-be-processed text belonging to a target knowledge field, and identify an abbreviation term independently distributed in the to-be-processed text as a to-be-completed abbreviation term by using a pre-trained abbreviation term recognition model; query a full name term corresponding to the to-be-completed abbreviation term as a candidate full name term in the abbreviation full name term word pair library; determine a target full name term based on the candidate full name term, and complete the full name term of the to-be-completed abbreviation term by using the target full name term as the full name term corresponding to the to-be-completed abbreviation term; if the number of the candidate full name terms is one, the candidate full name term is determined as the target full name term; if the number of the candidate full name terms is more than one, a full name term matching a semantic feature of the to-be-processed text is selected from the plurality of candidate full name terms as the target full name term.

Citation Information

Patent Citations

  • Medical question-answering method based on ontology semantic similarity

    CN110706807A

  • Initial abbreviation disambiguation method and system, electronic equipment and storage medium

    CN113449516A