Chinese Word Segmentation Method, Apparatus and Device
By performing word embedding processing of DR reports and graph attention neural network model word segmentation, the precise word segmentation problem of medical image reports is solved, and better word segmentation effect and semantic expression are achieved.
Patent Information
- Application Number
- CN202211728809.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-12-30
AI Technical Summary
The existing Chinese word segmentation method is difficult to accurately segment medical imaging reports (DR reports) and cannot effectively express its original semantics, mainly due to the large differences between medical vocabulary and daily terms and insufficient labeling data.
By obtaining the text of the DR report, word embedding processing is performed, the correlation relationship between word vectors is determined, and word segmentation is used to use the graph attention neural network model to perform word segmentation, multiple word segmentation labels and their probability are obtained, and the word segmentation result is finally determined.
Accurate word segmentation for DR reports is achieved, ensuring that word segmentation results can accurately express the original semantics and improve word segmentation effect.
Smart Images

Figure CN116205223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical imaging technology, and in particular to a Chinese word segmentation method, device and equipment. Background Art
[0002] Chinese word segmentation divides Chinese text into sequences of words that represent precise semantic meanings based on established criteria. In Chinese natural language processing tasks, the effectiveness of Chinese word segmentation methods will affect the results of subsequent tasks such as entity recognition, semantic analysis, and information analysis.
[0003] Traditional Chinese word segmentation methods are primarily based on machine learning and deep learning algorithms. Both methods require converting word segmentation tasks into character annotation tasks. Insufficient training data can affect the final Chinese word segmentation results. Furthermore, medical vocabulary differs from everyday spoken or written language. Furthermore, medical imaging reports are specialized written documents in the field of medical imaging, containing a wealth of specialized information. Therefore, using general word segmentation methods to segment medical vocabulary is difficult to achieve ideal results. Summary of the Invention
[0004] The embodiments of the present invention provide a Chinese word segmentation method, apparatus, and device, which can achieve accurate Chinese word segmentation of DR reports, so that the word segmentation results can accurately express the original semantics.
[0005] In a first aspect, an embodiment of the present invention provides a Chinese word segmentation method, the method comprising:
[0006] Get the DR report of the word to be segmented;
[0007] Performing word embedding processing on the text in the DR report to obtain a word vector corresponding to each word in the text;
[0008] Determining the association between each word vector in the text;
[0009] Based on the association relationship, the text is segmented to obtain a plurality of segmentation labels and probabilities corresponding to the plurality of segmentation labels;
[0010] The word segmentation result of the DR report to be segmented is determined according to the multiple word segmentation labels and the probabilities corresponding to the multiple word segmentation labels.
[0011] In a second aspect, an embodiment of the present invention provides a Chinese word segmentation device, the device comprising:
[0012] The acquisition module is used to obtain the DR report to be segmented;
[0013] The first processing module is used to perform word embedding processing on the text in the DR report to obtain the word vectors corresponding to each word in the text;
[0014] The first determination module is used to determine the correlation relationship between the word vectors in the text;
[0015] The second processing module is used to perform word segmentation processing on the text based on the correlation relationship to obtain a plurality of word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively;
[0016] The second determination module is used to determine the word segmentation result of the DR report to be word-segmented according to the plurality of word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively.
[0017] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory and a processor; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the Chinese word segmentation method as described in the first aspect.
[0018] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the Chinese word segmentation method as described in the first aspect.
[0019] The Chinese word segmentation technical solution provided by the embodiment of the present invention first obtains the DR report to be word-segmented, performs word embedding processing on the text in the DR report to obtain the word vectors corresponding to each word in the text. Then, it determines the correlation relationship between the word vectors in the text, and based on the correlation relationship, performs word segmentation processing on the text to obtain a plurality of word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively. Finally, according to the plurality of word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively, it determines the word segmentation result of the DR report to be word-segmented. In the above solution, through the correlation relationship between the word vectors in the text of the DR report, multiple word segmentations are performed on the text to obtain a plurality of word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively. Based on the individual word segmentation tags and the probabilities corresponding to the plurality of word segmentation tags respectively, the final word segmentation result of the DR report to be word-segmented is determined, which can not only achieve accurate word segmentation of the DR report, but also make the word segmentation result accurately express the original semantics, thus having a better word segmentation effect. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of a Chinese word segmentation method provided by an embodiment of the present invention;
[0022] Figure 2 It is a schematic application diagram of a Chinese word segmentation method provided by an embodiment of the present invention;
[0023] Figure 3 It is a schematic application diagram of establishing a word graph provided by an embodiment of the present invention;
[0024] Figure 4 It is a schematic flowchart of another Chinese word segmentation method provided by an embodiment of the present invention;
[0025] Figure 5 It is a schematic structural diagram of a Chinese word segmentation device provided by an embodiment of the present invention;
[0026] Figure 6 For Figure 5 It is a schematic structural diagram of an electronic device corresponding to the Chinese word segmentation device shown in the embodiment. Detailed implementation manners
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0028] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms of "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. "Multiple" generally includes at least two.
[0029] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0030] It can also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a commodity or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the commodity or system comprising the element.
[0031] In addition, the step timings in the following method embodiments are only examples, not strictly limited.
[0032] In practical applications, a computer cannot directly recognize the semantics of Chinese text, so the text must be decomposed into meaningful words for the computer to perform subsequent retrieval, reasoning and other tasks. Chinese word segmentation refers to splitting a Chinese text composed of a sequence of Chinese characters into individual words. For example, the sentence "I love my motherland" is manually segmented into three Chinese words: "I", "love", and "motherland". Through word segmentation technology, a sentence or even a paragraph can be split into various meaningful words, and then the decomposed words can be recognized and processed by the computer.
[0033] Chinese word segmentation is the basis of Chinese text mining, and text mining is the mining and discovery of knowledge based on text information. As in the above example, when the Chinese text "I love my motherland" is input into the computer's retrieval system and segmented into three Chinese words "I", "love", and "motherland" by a word segmenter, the computer can accurately recognize the semantics contained in this sentence and can perform a system retrieval with the retrieval strategy of "I" AND "love" AND "motherland".
[0034] Digital radiography (DR for short) is developed on the basis of digital fluorography. It uses an image intensifier as the information carrier to receive the X-ray information passing through the human body. After being collected by a video camera, it is converted into a digital signal and then digitized to obtain a digital human body graph. In medicine, DR is often used for imaging examinations to check the physical conditions of patients, such as whether there are bone fractures, whether the heart is enlarged, whether there is fluid in the lungs, whether there is stenosis or obstruction in the gastrointestinal tract, etc. The examination volume is huge and it has great clinical application value. Usually, after a patient undergoes a DR examination, a doctor will generate a DR report based on the DR examination results. That is, a DR report refers to a collection of various text information including descriptions of human tissues, descriptions of lesions, and professional diagnostic opinions given by a licensed radiologist who is legally registered under the current national laws and regulations system after viewing the patient's DR graph and combining their professional knowledge. Its carrier can be digital or in the form of paper, etc. There is a large amount of free text in the DR report, that is, text that has not been segmented. Therefore, the DR report cannot be directly recognized by a computer for further processing, so a word segmenter is needed to segment the DR report.
[0035] Due to the word segmentation method and the characteristics of DR corpus, existing popular and general Chinese word segmenters (such as NLPIR of the Institute of Computing Technology, Chinese Academy of Sciences, LTP of Harbin Institute of Technology, THULAC of Tsinghua University, Stanford Word Segmenter, Hanlp Word Segmenter, jieba, IKAnalyzer, etc.) cannot accurately segment the Chinese words in the DR report. Since the basic principle of Chinese word segmentation is usually based on the dictionary method or the statistical frequency method or a combination of the two, but the corpus in the DR report is quite different from the daily colloquial expressions or literary works. For example, "The physiological curvature of the sacrococcyx exists, the alignment sequence is normal, and no obvious scoliosis or spondylolisthesis signs are seen; no obvious osteophyte or bone destruction signs are seen in each vertebral body." The dictionary and word frequency of this kind of corpus are naturally different from those of general texts. Especially, understanding the semantics of this kind of text requires a strong medical background, and even must be interpreted by a radiologist. Therefore, it is not appropriate to use the current popular and general Chinese word segmenters for word segmentation of DR reports.
[0036] In addition, due to the extremely strong professionalism and barriness in the medical field, it is difficult to obtain a large amount of accurately labeled medical imaging report texts. Therefore, using traditional Chinese word segmentation methods will result in the segmented text not being able to accurately express the original semantics because the amount of labeled DR report text data is insufficient and not trained and modeled for the field of medical imaging, and the expected word segmentation effect cannot be achieved.
[0037] To solve the above technical problems, this embodiment provides a method for Chinese word segmentation applicable to DR reports. In this technical solution, the correlation relationship between each word vector in the text of the DR report is determined, and based on the correlation relationship between each word vector, multiple word segmentations are performed on the text to obtain multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels. Based on the word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, the final word segmentation result of the DR report to be segmented is determined, which can not only achieve accurate word segmentation of the DR report, but also enable the word segmentation result to accurately express the original semantics, thereby achieving a better word segmentation effect.
[0038] The Chinese word segmentation method provided by the present invention can be executed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, etc., or a server. The server can be a physical server including an independent host, or can also be a virtual server, or can also be a cloud server or a server cluster. The following will exemplarily illustrate the Chinese word segmentation process of the DR report with reference to the accompanying drawings.
[0039] Figure 1 It is a flowchart of a Chinese word segmentation method provided by an embodiment of the present invention; Figure 2 It is an application schematic diagram of a Chinese word segmentation method provided by an embodiment of the present invention; Refer to the attached Figure 1-2 As shown, this embodiment provides a Chinese word segmentation method. The execution subject of this method can be a Chinese word segmentation device. It can be understood that this device can be implemented as software, or a combination of software and hardware. Specifically, the Chinese word segmentation method can include:
[0040] Step 101: Obtain the DR report to be segmented.
[0041] Step 102: Perform word embedding processing on the text in the DR report to obtain the word vector corresponding to each word in the text.
[0042] Step 103: Determine the correlation relationship between each word vector in the text.
[0043] Step 104: Based on the correlation relationship, perform word segmentation processing on the text to obtain multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels.
[0044] Step 105: According to the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, determine the word segmentation result of the DR report to be segmented.
[0045] In practical applications, after a DR imaging examination of a patient, a practicing radiologist will issue a professional DR report based on the DR imaging examination results and medical expertise. As an information carrier transformed by human expertise, the DR report contains a large amount of imaging detail information about human structures or lesions. In-depth exploration of this information is helpful for clinical research and clinical decision-making, and further promotes the development of precision diagnosis and treatment. On the other hand, by deeply exploring the text content of the DR report, the quality of the DR report can be evaluated, which has a positive significance for improving the level of imaging diagnosis and treatment work. However, since each physician has different writing habits, the obtained DR reports contain a large amount of free text, and computers cannot directly recognize and process the DR reports. Therefore, when conducting further in-depth research on DR reports, it is first necessary to accurately segment a large number of unannotated DR reports to obtain relatively accurate word segmentation results, so that computers can directly recognize and carry out the next step of research work. Therefore, how to accurately segment DR reports is very important.
[0046] When performing Chinese word segmentation on a DR report, first obtain the DR report to be segmented. Then, perform word embedding processing on the text in the DR report to obtain the word vectors corresponding to each word in the text. Among them, the text in the DR report contains a large amount of text, and computers cannot directly recognize the text. However, for text, how to effectively input it into the computer and be recognized by the computer for further operation is very important. Then, when performing word segmentation on the text in the DR report, the text in the DR report can be converted into a language that the computer can recognize and then input into the computer to facilitate computer processing.
[0047] Among them, word embedding processing refers to converting each word in the text into a real number vector. In the embodiments of the present invention, each Chinese character in the text is subjected to word embedding operation through the word embedding technology to obtain a plurality of word vectors, so that the computer can directly recognize the text in the DR report. After obtaining the word vectors corresponding to each word in the text, then determine the correlation relationship between the various word vectors in the text. Among them, the correlation relationship can refer to the connection relationship between word vectors, the position relationship of word vectors in the text, the context relationship corresponding to word vectors, the front-back relationship corresponding to word vectors, etc.
[0048] In an alternative embodiment, the specific implementation of determining the association relationship between each word vector in the text may be as follows: Based on the DR report corpus, determine the neighborhood features of each word vector in the text, obtain the connection relationship between each word vector, and based on the connection relationship between each word vector, determine the weight corresponding to the domain features. Based on the weights, determine the association relationship between each word vector in the text. A word vector may correspond to multiple domain features, and each domain feature plays a certain role in determining the association relationship of the word vector. Among the numerous domain features, the domain features that are more critical to the current word vector can be focused on, the attention to other domain features can be reduced, and even irrelevant information can be filtered out. Therefore, the weights corresponding to each domain feature can be determined, and based on the weights corresponding to each domain feature, the association relationship between each word vector in the text can be determined.
[0049] In addition, the dictionary corresponding to the DR report corpus and / or the word frequency appearing in the DR report corpus can also be used, and combined with semantics to determine the association relationship between each word vector in the text. Specifically, when using the dictionary-based word segmentation method, a dictionary corresponding to the DR report corpus is required to match the text to be segmented with the entries in the dictionary. In order to make the word segmentation result more accurate, when using the dictionary to determine the association relationship between each word vector, the shortest path can be combined to jointly determine the association relationship between each word vector in the text. Specifically, before using the shortest path method, in addition to the dictionary, a word graph is also required. Combining Figure 3 An exemplary illustration of this is as Figure 3 shown. A directed acyclic word graph with N + 1 nodes is established. The word graph is a fully connected graph, each vertex represents a word in the dictionary, and the edge between vertices represents the distance between two combined words. For example, the distance between "weather" and "good" is less than the distance between "weather" and "big" because the combination "weather good" is more likely. Among them, the shortest path of the word graph can be solved using methods such as Dijkstra. In addition, the method of statistical word frequency can be used to determine the association relationship between each word vector in the text. Specifically, based on the statistical word segmentation method, a generation model of the learning sample is first established, and then the model is used to indirectly infer the prediction result, so as to determine the association relationship between each word vector.
[0050] Next, based on the correlation relationships between the individual word vectors, word segmentation is performed on the text to obtain multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels. First, based on the correlation relationships between the individual word vectors, multiple segmentation methods are determined, and the text is segmented based on the multiple segmentation methods respectively, and the probability corresponding to each segmentation method is predicted. For example, when performing word segmentation on the text "I love him", there are four corresponding segmentation methods as follows: the first one is "I / love / him"; the second one is "I love / him"; the third one is "I / love him"; the fourth one is "I love him", and the probabilities corresponding to each word segmentation method are P1, P2, P3, and P4 respectively.
[0051] Finally, based on the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, the word segmentation result of the DR report to be segmented is determined. The DR report text is segmented in multiple ways to obtain multiple word segmentation labels and the probabilities corresponding to each word segmentation label. The word segmentation label with the highest probability among the word segmentation labels can be selected as the final word segmentation result of the DR report. Optionally, the specific implementation method for determining the word segmentation result of the DR report to be segmented based on the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels may include: sorting the probabilities corresponding to the multiple word segmentation labels from largest to smallest; determining the word segmentation label with the largest probability as the target word segmentation label, and determining the target word segmentation label as the word segmentation result of the DR report to be segmented.
[0052] In another alternative embodiment, the probabilities corresponding to the multiple word segmentation labels can be weighted to obtain the final label probability, and the final word segmentation result of the DR report is determined based on the final label probability. In this way, the advantages of each segmentation method can be fully considered to improve the word segmentation efficiency. Among them, the weight corresponding to the weighting is greater than 0 and less than 1.
[0053] In the embodiment of the present invention, through word embedding processing on the text in the DR report, the word vector corresponding to each word in the text is obtained. Next, the correlation relationships between the individual word vectors in the text are determined, multiple word segmentations are performed on the text to obtain multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels. Based on the individual word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, the final word segmentation result of the DR report to be segmented is determined. In this way, not only can accurate word segmentation of the DR report be achieved, but also the word segmentation result can accurately express the original semantics, thus having a better word segmentation effect.
[0054] Figure 4 It is a flowchart showing another Chinese word segmentation method provided by the embodiment of the present invention; on the basis of the above embodiment, continue to refer to the attached Figure 4 As shown, this embodiment provides an implementation method that can perform Chinese word segmentation based on a graph attention neural network model, which specifically may include:
[0055] Step 401: Obtain the DR report to be segmented.
[0056] Step 402: Perform word embedding processing on the text in the DR report to obtain the word vectors corresponding to each word in the text.
[0057] Step 403: Input the word vectors corresponding to each word in the text into a pre-trained Chinese word segmentation model to obtain multiple word segmentation labels corresponding to the text and the probabilities corresponding to each of the multiple word segmentation labels, where the Chinese word segmentation model is trained to segment the text and predict the probabilities corresponding to the word segmentation labels.
[0058] Step 404: Determine the word segmentation result of the DR report to be segmented according to the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels.
[0059] In traditional NLP tasks, text sequences are considered as bags of tokens. With the emergence of word embedding technology, sentences are represented as sequences of tokens. A large number of NLP problems can be best expressed in graph structures, and graph-structured data can encode complex pairwise relationships between entity tokens to learn more informative representations for further selection. Therefore, in the embodiments of the present invention, a Chinese word segmentation model for DR reports is obtained based on a graph attention neural network, and the Chinese word segmentation model is used to perform Chinese word segmentation processing on a large number of unannotated DR reports. Specifically, the Graph Attention Network (GAT) is adapted for applicability to obtain a Chinese word segmentation model for DR examination report text data, which can improve the word segmentation effect of the word segmentation model to achieve the expected word segmentation effect.
[0060] The graph attention network, a new neural network architecture based on graph-structured data, uses hidden self-attention layers to address the deficiencies of previous methods based on graph convolution or its approximations. By stacking layers, nodes can participate in the features of their neighbors, and different weights can be (implicitly) assigned to different nodes in the neighborhood without any costly matrix operations (such as inversion) and without prior knowledge of the graph structure, making it suitable for inductive and deductive problems. Due to the advantages of the graph attention model network and the characteristics of the DR report corpus, it is determined that a graph attention neural network can be used for learning and training to generate a Chinese word segmentation model for DR reports.
[0061] As mentioned above, DR reports have a distinct corpus. Moreover, there are inherent professional knowledge associations between sentences or words. For example, in the sentence "The physiological curvature of the sacrococcyx exists, the alignment sequence is normal, no obvious scoliosis or spondylolisthesis signs are seen; no obvious osteophyte or bone destruction signs are seen in each vertebral body", through the professional knowledge of medical imaging, it can be judged that the word "sacrococcyx" has a strong association with the words "physiological curvature", "alignment sequence", "scoliosis", and "spondylolisthesis". That is, when writing a DR report, if a radiologist sees the sacrococcyx presented in the patient's DR image, then through professional knowledge and training, it is known that the sacrococcyx should be evaluated in four conventional aspects: "physiological curvature", "alignment sequence", "scoliosis", and "spondylolisthesis" (note that if there are other lesions, they should be described separately; this is just an example). Therefore, the word "sacrococcyx" has a strong connection with the words "physiological curvature", "alignment sequence", "scoliosis", and "spondylolisthesis", which is an important corpus feature of DR reports. From this, it can be seen that the relationship between words becomes particularly important, directly affecting the dictionary and word frequency. So, a model that can express complex relationships is needed and trained specifically for DR reports. Therefore, in the embodiments of the present invention, a graph attention neural network that introduces an attention mechanism to the relationships between each node of the model is selected for this task. That is, using the pre-annotated Chinese DR report corpus as the training set, it is input into the word segmentation model based on the input graph attention model network for training to obtain the Chinese word segmentation model of the Chinese DR report.
[0062] Specifically, when performing Chinese word segmentation on a DR report, first obtain the DR report to be segmented, and then perform character embedding processing on the text in the DR report to obtain the character vector corresponding to each character in the text. That is, each Chinese character in the text is converted into a vector form that can be directly recognized and processed by a computer. Then, input the character vectors corresponding to each character in the text into the pre-trained Chinese word segmentation model to use the Chinese word segmentation model to obtain multiple word segmentation labels corresponding to the text and the probabilities corresponding to each of the multiple word segmentation labels. Among them, the Chinese word segmentation model is trained to segment the text and predict the probability corresponding to the word segmentation label.
[0063] In an embodiment of the present invention, each word vector in the DR report text is processed by a pre-trained Chinese word segmentation model to determine the association relationship between the word vectors, and the text is segmented to obtain a plurality of word segmentation tags corresponding to the DR report text, and the probability corresponding to each of the plurality of word segmentation tags is predicted. Among them, the Chinese word segmentation model can be pre-trained to segment text and predict the probability corresponding to the word segmentation tag. In addition, a Chinese word segmentation model for the DR report can be generated by learning and training a graph attention neural network, that is, the graph attention neural network is learned and trained using a preset pre-annotated DR report corpus, so that a Chinese word segmentation model can be obtained. Specifically, the implementation process of obtaining the pre-trained Chinese word segmentation model can be: obtaining a pre-annotated DR report corpus, performing word embedding processing on the pre-annotated DR report corpus to obtain a plurality of word vectors corresponding to the DR report corpus, and inputting the plurality of word vectors into the graph attention neural network for training to obtain the pre-trained Chinese word segmentation model.
[0064] Among them, annotation means tagging. Common text annotation methods include 2-word position annotation method, 4-word position annotation method, 6-word position annotation method, etc., and the corresponding annotation method can be selected according to actual needs. Here, the 4-word position annotation method is used as an example. Among them, B represents the first position value of a word, M represents the middle position of a word, E represents the end position of a word, and S represents a single word. For example, for the text "I am Chinese", it is annotated according to the rules of the 4-word position annotation method as "I / S am / S Chinese / B people / M".
[0065] In addition, in the embodiment of the present invention, the specific implementation manner of obtaining the pre-annotated DR report corpus is not limited. Preferably, the specific implementation manner of obtaining the pre-annotated DR report corpus can include: obtaining a plurality of DR report texts, segmenting the plurality of DR report texts based on a preset word segmentation package to obtain a preliminary word segmentation result. Then, the preliminary word segmentation result is corrected to obtain a corrected word segmentation result. Finally, the 4-word annotation method is used to perform annotation processing on the corrected word segmentation result to obtain the pre-annotated DR report corpus.
[0066] In specific applications, the jieba word segmentation package on the Python platform is used to perform preliminary word segmentation on multiple obtained DR reports to obtain preliminary word segmentation results. Then, the preliminary word segmentation results are manually checked and corrected to ensure the accuracy of the word segmentation results, and the corrected word segmentation results are obtained. Finally, the corrected word segmentation results after manual correction are labeled using the 4-word position annotation method to obtain the training set of the model. For example, for the text "The physiological curvature of the sacrococcyx exists", it is first imported into the jieba word segmentation package on the Python platform for preliminary word segmentation to obtain the preliminary word segmentation results "sacrum / tail / vertebra / physiology / curvature / exists". Then, the manual word segmentation opinions of radiologists are adopted for calibration to obtain the corrected word segmentation results as "sacrococcyx / physiology / curvature / exists". Finally, it is labeled using the 4-word position annotation method to obtain the labeled results "sacB / tailM / vertebraE / lifeB / physiologyE / curveB / degreeE / existB / existE".
[0067] After the above annotation, using the word embedding technology, multiple DR reports are converted into multiple word vectors based on the DR report corpus. The DR report corpus is input into the graph attention neural network for training to obtain a pre-trained Chinese word segmentation model. The specific training process may include: inputting the multiple word vectors corresponding to the DR report corpus into the graph attention neural network model, and based on the multiple word vectors, determining the directed graph corresponding to the multiple word vectors. Then, based on the directed graph, determining the weights corresponding to each parameter in the graph attention neural network. Finally, based on the weights corresponding to each parameter, generating a pre-trained Chinese word segmentation model.
[0068] In an optional embodiment, the directed graph corresponding to the multiple word vectors can be determined by determining the edge weights. Specifically, determining the first edge weights between each word vector, determining the second edge weights between the word vector and the text vector, and based on the first edge weights and the second edge weights, determining the directed graph corresponding to the multiple word vectors. Among them, both the text and the word vector are used as nodes of the directed graph. Then the edge weights include two categories. One category is the edge weights between word nodes, and the other category is the edge weights between word nodes and text nodes. Based on these two categories of edge weights, a complete directed graph can be established. Among them, based on the directed graph, not only the association relationship between word vectors and the association relationship between word vectors and text can be determined, but also the semantics of word vectors can be understood based on the directed graph.
[0069] Then, using the multi-head attention mechanism algorithm, based on the directed graph, determine the weights corresponding to each parameter in the graph attention neural network, and based on each parameter and the weights corresponding to each parameter, determine a new word segmentation function. This new word segmentation function is the Chinese word segmentation model, and the probability corresponding to each word segmentation label can be output at the fully connected layer based on the above-normalized attention neural network, denoted as Pt. That is, based on the Chinese word segmentation model, not only can multiple word segmentation labels corresponding to the text be determined based on various segmentation methods, but also the probabilities corresponding to each segmentation method can be determined.
[0070] After establishing the Chinese word segmentation model, the Chinese word segmentation model can be used to analyze and process each word vector corresponding to the DR report to be segmented, so that multiple word segmentation labels corresponding to the DR report to be segmented and the probabilities corresponding to each word segmentation label can be obtained. In the embodiment of the present invention, by using the pre-trained Chinese word segmentation model to analyze and process multiple word vectors corresponding to the DR report to be segmented, obtaining multiple word segmentation labels corresponding to the DR report to be segmented and the probabilities corresponding to each of the multiple word segmentation labels can not only effectively ensure accurate word segmentation of the DR report, ensure the accuracy of the word segmentation result, enable the word segmentation result to accurately express the original semantics, but also further improve the usage stability of this Chinese word segmentation method.
[0071] Finally, according to multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, determine the word segmentation result of the DR report to be segmented. Specifically, the probabilities corresponding to each of the multiple word segmentation labels can be sorted from largest to smallest, and the word segmentation label with the largest probability is determined as the target word segmentation label, and the target word segmentation label is determined as the word segmentation result of the DR report to be segmented. The word segmentation result determined in this way is more accurate and can better restore the original semantics.
[0072] In addition, it should be noted that: the process of determining the word segmentation result of the DR report to be segmented according to multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels can be executed in the Chinese word segmentation model or directly by the Chinese word segmentation device, and no specific limitation is made on this, and it can be set according to actual usage needs. When the process of determining the word segmentation result of the DR report to be segmented according to multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels is executed in the Chinese word segmentation model, the Chinese word segmentation model determines multiple word segmentation labels corresponding to the text based on multiple segmentation methods, predicts the probabilities corresponding to each of the word segmentation labels, then sorts the probabilities corresponding to each of the multiple word segmentation labels from largest to smallest, and outputs the word segmentation label with the largest probability.
[0073] For the detailed execution process and technical effects related thereto, refer to the description in the foregoing embodiments, and will not be elaborated herein.
[0074] In summary, in the embodiments of the present invention, by obtaining a DR report to be segmented, word embedding processing is performed on the text in the DR report to obtain word vectors corresponding to each word in the text. Then, the word vectors corresponding to each word in the text are input into a pre-trained Chinese word segmentation model to obtain multiple word segmentation labels corresponding to the text and probabilities corresponding to each of the multiple word segmentation labels, where the Chinese word segmentation model is trained to segment the text and predict the probabilities corresponding to the word segmentation labels. Finally, according to the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels, the word segmentation result of the DR report to be segmented is determined. The word segmentation result obtained in this way is more accurate, enabling the obtained word segmentation result to accurately express the original semantics, so that the Chinese word segmentation method can have a better word segmentation effect.
[0075] In addition, since the amount of labeled image DR report text data is insufficient, it will affect the word segmentation effect of the obtained Chinese word segmentation model. To enable the Chinese word segmentation model to have a better word segmentation effect, the Chinese word segmentation method provided in the embodiments of the present invention may further include: performing inductive learning on the Chinese word segmentation model obtained after learning and training based on a graph attention neural network to obtain an inductive learning result, then performing an attention coefficient normalization operation on the inductive result to obtain a visualized inductive relationship, and determining the final word segmentation result based on the transductive result and the inductive learning result. Among them, inductive learning is an unsupervised learning. Based on a part of the DR report corpus with labels and the DR report text with annotation labels obtained, learning and training are performed to obtain a more accurate Chinese word segmentation model for DR reports. This way of inductive learning requires inputting the DR report data without annotation labels into the Chinese word segmentation model for continued learning and training, which can increase the amount of DR report corpus data and thus improve the performance of the Chinese word segmentation model. In this way, when the DR report to be segmented is segmented based on the Chinese word segmentation model, a better word segmentation effect can be obtained.
[0076] The Chinese word segmentation device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.
[0077] Figure 5 The structural schematic diagram of a Chinese word segmentation device provided for the embodiments of the present invention is as Figure 5 shown. The device includes: an acquisition module 51, a first processing module 52, a first determination module 53, a second processing module 54, and a second determination module 55.
[0078] The acquisition module 51 is used to acquire a DR report to be segmented;
[0079] The first processing module 52 is configured to perform word embedding processing on the text in the DR report to obtain word vectors corresponding to each word in the text;
[0080] The first determination module 53 is configured to determine the association relationship between the word vectors in the text;
[0081] The second processing module 54 is configured to perform word segmentation processing on the text based on the association relationship to obtain a plurality of word segmentation tags and probabilities corresponding to the plurality of word segmentation tags respectively;
[0082] The second determination module 55 is configured to determine the word segmentation result of the DR report to be word-segmented according to the plurality of word segmentation tags and probabilities corresponding to the plurality of word segmentation tags respectively.
[0083] Preferably, the first determination module 53 may specifically be configured to: determine the neighborhood features of the word vectors in the text based on the DR report corpus; determine the weights corresponding to the domain features based on the connection relationship between the word vectors; and determine the association relationship between the word vectors in the text based on the weights.
[0084] Preferably, the first determination module 53 and the second processing module 54 may specifically be configured to: input the word vectors corresponding to each word in the text into a pre-trained Chinese word segmentation model to obtain a plurality of word segmentation tags corresponding to the text and probabilities corresponding to the plurality of word segmentation tags respectively by using the Chinese word segmentation model, where the Chinese word segmentation model is trained to obtain the probabilities corresponding to the word segmentation tags.
[0085] Preferably, the device may further include a training module, which may specifically be configured to: obtain a pre-annotated DR report corpus; perform word embedding processing on the pre-annotated DR report corpus to obtain a plurality of word vectors corresponding to the DR report corpus; input the plurality of word vectors into a graph attention neural network model for training to obtain a pre-trained Chinese word segmentation model.
[0086] Preferably, the training module may specifically further be configured to: obtain a plurality of DR report texts; perform word segmentation on the plurality of DR report texts based on a preset word segmentation package to obtain a preliminary word segmentation result; perform correction processing on the preliminary word segmentation result to obtain a corrected word segmentation result; and use the 4-word annotation method to perform annotation processing on the corrected word segmentation result to obtain a pre-annotated DR report corpus.
[0087] Preferably, the training module may specifically be further configured to: input the multiple word vectors corresponding to the DR report corpus into a graph attention neural network model; determine a directed graph corresponding to the multiple word vectors based on the multiple word vectors; determine weights corresponding to each parameter in the graph attention neural network model based on the directed graph; and generate a pre-trained Chinese word segmentation model based on the weights corresponding to each parameter.
[0088] Preferably, the training module may specifically be further configured to: determine a first edge weight value between each of the word vectors; determine a second edge weight value between the word vector and the text vector; and determine a directed graph corresponding to the multiple word vectors based on the first edge weight value and the second edge weight value.
[0089] Preferably, the second determination module 55 may specifically be configured to: sort the probabilities corresponding to the multiple word segmentation labels from largest to smallest; determine the word segmentation label with the largest probability as the target word segmentation label; and determine the target word segmentation label as the word segmentation result of the DR report to be segmented.
[0090] Figure 5 The illustrated device may execute the Chinese word segmentation method provided in the foregoing Figures 1 to 4 For the detailed execution process and technical effects, refer to the description in the foregoing embodiments, which will not be elaborated herein.
[0091] In a possible design, the structure of the foregoing Figure 5 illustrated Chinese word segmentation device may be implemented as an electronic device. The electronic device provided in this embodiment may obtain a DR report to be segmented;
[0092] perform word embedding processing on the text in the DR report to obtain word vectors corresponding to each word in the text;
[0093] determine the association relationship between each word vector in the text;
[0094] perform word segmentation processing on the text based on the association relationship to obtain multiple word segmentation labels and probabilities corresponding to each of the multiple word segmentation labels;
[0095] determine the word segmentation result of the DR report to be segmented according to the multiple word segmentation labels and the probabilities corresponding to each of the multiple word segmentation labels.
[0096] As Figure 6 illustrated, the electronic device may include: a processor 61 and a memory 62. Among them, an executable code is stored on the memory 62. When the executable code is executed by the processor 61, the processor 61 can at least implement the Chinese word segmentation method provided in the foregoing Figures 1 to 4 illustrated embodiments.
[0097] In an embodiment of the present invention, the memory 62 is used to store computer programs and can be configured to store various other data to support operations on the device where it is located. Among them, the processor 61 can execute the computer programs stored in the memory to implement corresponding control logic. The memory 62 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0098] In an embodiment of the present invention, the processor can be any hardware processing device capable of executing the above method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a field-programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), a complex programmable logic device (CPLD), etc.; or an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system-on-chip (SoC), etc., but not limited thereto.
[0099] Optionally, the electronic device may further include a communication interface 63 for communicating with other devices. The communication interface is configured to facilitate communication between the device where it is located and other devices in a wired or wireless manner. The device where the communication interface is located can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, or other technologies.
[0100] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the Chinese word segmentation method provided in the foregoing Figures 1 to 4 embodiment as shown.
[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A Chinese word segmentation method, characterized in that, including: Obtain a DR report to be segmented; Perform word embedding processing on the text in the DR report to obtain word vectors corresponding to each word in the text; Based on the DR report corpus, determine the neighborhood features of each word vector in the text; Obtain the connection relationships between the word vectors; Based on the connection relationships between the word vectors, determine the weights corresponding to the domain features; Based on the weights, determine the association relationships between the word vectors in the text; Based on the association relationships, determine multiple segmentation methods for segmenting the text, and segment the text based on the multiple segmentation methods respectively to obtain multiple segmentation labels and probabilities corresponding to the multiple segmentation labels; According to the multiple segmentation labels and the probabilities corresponding to the multiple segmentation labels, determine the segmentation result of the DR report to be segmented.
2. The method according to claim 1, wherein The determining the association relationships between the word vectors in the text; based on the association relationships, segmenting the text to obtain multiple segmentation labels and probabilities corresponding to the multiple segmentation labels includes: Input the word vectors corresponding to each word in the text into a pre-trained Chinese word segmentation model to obtain multiple segmentation labels corresponding to the text and probabilities corresponding to the multiple segmentation labels by using the Chinese word segmentation model, where the Chinese word segmentation model is trained to segment text and predict the probabilities corresponding to the segmentation labels.
3. The method according to claim 2, wherein The method further includes: Obtain a pre-annotated DR report corpus; Perform word embedding processing on the pre-annotated DR report corpus to obtain multiple word vectors corresponding to the DR report corpus; Input the multiple word vectors into a graph attention neural network for training to obtain a pre-trained Chinese word segmentation model.
4. The method according to claim 3, characterized in that, The obtaining the pre-annotated DR report corpus includes: Obtain multiple DR report texts; Based on a preset segmentation package, segment the multiple DR report texts to obtain preliminary segmentation results; Perform correction processing on the preliminary segmentation results to obtain corrected segmentation results; Use the 4-word annotation method to perform annotation processing on the corrected segmentation results to obtain a pre-annotated DR report corpus.
5. The method according to claim 4, wherein The inputting the multiple word vectors into a graph attention neural network model for training to obtain a pre-trained Chinese word segmentation model includes: Input the multiple word vectors corresponding to the DR report corpus into the graph attention neural network model; Based on the multiple word vectors, determine a directed graph corresponding to the multiple word vectors; Based on the directed graph, determine the weights corresponding to each parameter in the graph attention neural network model; Based on the weights corresponding to each parameter, generate a pre-trained Chinese word segmentation model.
6. The method according to claim 5, wherein The based on the multiple word vectors, determining a directed graph corresponding to the multiple word vectors includes: Determine the first edge weights between the word vectors; Determine the second edge weights between the word vectors and the text vectors; Based on the first edge weights and the second edge weights, determine the directed graph corresponding to the multiple word vectors.
7. The method according to claim 1, characterized in that Determining the word segmentation result of the DR report to be segmented according to the multiple word segmentation tags and the probabilities respectively corresponding to the multiple word segmentation tags includes: Sorting the probabilities respectively corresponding to the multiple word segmentation tags from largest to smallest; Determining the word segmentation tag with the largest probability as the target word segmentation tag; Determining the target word segmentation tag as the word segmentation result of the DR report to be segmented.
8. A Chinese word segmentation device, characterized in that, Including: An acquisition module, configured to acquire a DR report to be segmented; A first processing module, configured to perform word embedding processing on the text in the DR report to obtain word vectors corresponding to each word in the text; A first determination module, configured to determine the neighborhood features of each word vector in the text based on a DR report corpus, obtain the connection relationships between the word vectors, determine the weights corresponding to the neighborhood features based on the connection relationships between the word vectors, and determine the association relationships between the word vectors in the text based on the weights; A second processing module, configured to determine multiple segmentation methods for segmenting the text based on the association relationships, and segment the text respectively based on the multiple segmentation methods to obtain multiple word segmentation tags and the probabilities respectively corresponding to the multiple word segmentation tags; A second determination module, configured to determine the word segmentation result of the DR report to be segmented according to the multiple word segmentation tags and the probabilities respectively corresponding to the multiple word segmentation tags.
9. An electronic device, characterized in that, Including: A memory and a processor; wherein, the memory is used to store one or more computer instructions, and when the one or more computer instructions are executed by the processor, the Chinese word segmentation method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Deep learning-based merger and acquisition announcement abstract method
CN110956041A
Chinese word segmentation method and device and storage medium
CN111209751A