Text merging judgment model training method and text merging judgment method
By constructing a text merging judgment model trained with positive and negative sample groups, and utilizing encoders such as BERT and LSTM and fully connected layers to process features, the problem of insufficient accuracy in merging judgment caused by text segmentation errors is solved, achieving more efficient and accurate text merging judgment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT WEALTH (SHANGHAI) FINANCIAL INFORMATION SERVICES CO LTD
- Filing Date
- 2022-11-22
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, text merging judgments due to text segmentation errors are not accurate enough, affecting the effectiveness of applications such as text deduplication and intelligent question answering.
By constructing positive and negative sample groups, a text merging judgment model is trained. Self-supervised learning and multiple rounds of training are used to improve the robustness and accuracy of the model, including feature processing using BERT models, LSTM encoders, and fully connected layers.
This improves the training efficiency and accuracy of the text merging judgment model, ensuring that the merged text has complete semantics and is easy for users to understand.
Smart Images

Figure CN115905865B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of natural language processing technology, and in particular to a training method and a text merging judgment method for a text merging judgment model. Background Technology
[0002] Typically, a long text can be segmented into multiple sentences using delimiters such as periods, exclamation marks, questions, or commas. However, due to the complexities of text generation environments, input text may contain incorrect delimiters. For example, users inputting text via a mobile device's touchscreen might use incorrect delimiters, excessive spaces, or incorrect line breaks. Similarly, users inputting text via voice might experience segmentation errors due to poor voice input conditions or abnormal pauses. Therefore, determining whether two sentences (i.e., two short texts) can be merged has always been a fundamental task in the field of artificial intelligence natural language processing, serving as a foundational technology for higher-level applications such as text deduplication and intelligent question answering. Summary of the Invention
[0003] This specification provides a text merging judgment method, apparatus, storage medium, and electronic device, which can train a text merging judgment model, improve the robustness of the text merging model, and increase the accuracy of the text merging judgment model in determining whether two documents can be merged. The technical solution is as follows:
[0004] Firstly, embodiments of this specification provide a training method for a text merging judgment model, the method comprising:
[0005] Obtain at least one positive sample group and at least one negative sample group, wherein the positive sample group comprises two texts that cannot be merged and the negative sample group comprises two texts that can be merged;
[0006] The text merging judgment model is trained using at least one positive sample group and at least one negative sample group until the text merging judgment model converges.
[0007] Secondly, embodiments of this specification provide a method for determining text merging, the method comprising:
[0008] Obtain two texts to be detected;
[0009] The two texts to be detected are input into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model described in the first aspect.
[0010] Thirdly, embodiments of this specification provide a training apparatus for a text merging judgment model, the method comprising:
[0011] A sample acquisition module is used to acquire at least one positive sample group and at least one negative sample group, wherein the positive sample group includes two texts that cannot be merged, and the negative sample group includes two texts that can be merged.
[0012] The model training module is used to train the text merging judgment model using the at least one positive sample group and the at least one negative sample group until the text merging judgment model converges.
[0013] Fourthly, embodiments of this specification provide a device for determining text merging, the device comprising:
[0014] The text acquisition module is used to acquire two texts to be detected.
[0015] The result acquisition module is used to input the two texts to be detected into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model described in the first aspect.
[0016] Fifthly, embodiments of this specification provide a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.
[0017] Sixthly, embodiments of this specification provide a computer program product that stores multiple instructions adapted for loading by a processor and executing the above-described method steps.
[0018] In a seventh aspect, embodiments of this specification provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0019] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0020] This specification's embodiments reasonably construct at least one positive sample group and one negative sample group. The positive sample group includes texts that cannot be merged, and the negative sample group includes texts that can be merged. By using at least one positive and negative sample group, the text merging judgment model can learn in a self-supervised manner whether there is a mergeable relationship between two texts until the text merging judgment model converges, thereby improving the training efficiency of the text merging judgment model. Furthermore, by using at least one positive and negative sample group, the text merging judgment model can be trained in multiple rounds, so that the trained text merging judgment model has better anti-interference and robustness, and higher accuracy in performing the task of judging whether two texts are merged, thus obtaining merged texts with complete semantics, which is convenient for users to read and understand. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating a text merging judgment model provided in the embodiments of this specification to determine whether texts can be merged;
[0023] Figure 2 This is a training method for a text merging judgment model provided in the embodiments of this specification;
[0024] Figure 3 This is a schematic diagram of a process for obtaining a negative sample group provided in the embodiments of this specification;
[0025] Figure 4 This is a schematic diagram of a process for obtaining a negative sample group provided in the embodiments of this specification;
[0026] Figure 5 This is a schematic diagram of the structure of a text merging judgment model provided in the embodiments of this specification;
[0027] Figure 6 This is a flowchart illustrating a text merging judgment model provided in the embodiments of this specification to determine whether texts can be merged;
[0028] Figure 7 This is a schematic diagram illustrating a scenario of a text merging judgment method provided in the embodiments of this specification;
[0029] Figure 8 This is a flowchart illustrating a text merging judgment method provided in the embodiments of this specification;
[0030] Figure 9This is a schematic diagram of the structure of a training device for a text merging judgment model provided in the embodiments of this specification;
[0031] Figure 10 This is a schematic diagram of the structure of a text merging judgment device provided in the embodiments of this specification;
[0032] Figure 11 This is a schematic diagram of the structure of an electronic device provided in the embodiments of this specification. Detailed Implementation
[0033] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0034] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0035] The present specification will now be described in detail with reference to specific embodiments.
[0036] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0037] With the continuous development of network technology, artificial intelligence technology has been applied to various fields, such as the technology to determine whether two texts can be merged. Normally, a long text can be split into multiple sentences using delimiters such as ".", "!", "?", or even ",". However, due to the highly complex environment in which text is generated, the input text may contain incorrect delimiters. For example, a user inputting text through a mobile device's touchscreen might use incorrect delimiters, excessive spaces, or incorrect line breaks. Similarly, a user inputting text via voice might experience segmentation errors due to poor voice input conditions or abnormal pauses.
[0038] For example, if the user inputs the text "Regarding this issue, I have a different opinion, and I hope everyone will listen," a period is incorrectly used as a separator between text 1 "Regarding this issue" and text 2 "I have a different opinion, and I hope everyone will listen." In fact, text 1 and text 2 should be combined, and the correct combined text is "Regarding this issue, I have a different opinion, and I hope everyone will listen."
[0039] Therefore, a text merging judgment model for determining whether two texts to be detected can be merged has emerged. For example... Figure 1 As shown, Figure 1 This is a flowchart illustrating a text merging judgment model provided in this embodiment of the specification, which determines whether texts can be merged. Text 1011 and text 1012 are input into the text merging judgment model 102, so that the text merging judgment model 102 can determine whether text 1011 and text 1012 can be merged, and outputs a judgment result 103. The judgment result 103 includes at least two results: one is "can be merged" and the other is "cannot be merged".
[0040] For example, if the input text 1011 is "Regarding this issue" and the input text 1012 is "I have a different opinion, and I hope everyone will listen", the text merging judgment model 102 determines whether text 1011 and text 1012 can be merged, outputs the judgment result 103 as "can be merged", and performs subsequent text processing tasks accordingly.
[0041] Text merging judgment models in related technologies are mainly divided into two categories: models built based on machine learning methods in artificial intelligence, and models built based on deep learning methods in artificial intelligence. Specifically, the judgment process of models built based on machine learning methods is as follows: the text merging judgment problem is broken down into feature engineering and a classifier. Feature engineering includes two parts: text preprocessing, feature extraction, and text representation. First, the two texts are cleaned separately, and word segmentation tools are used to segment each text. Then, methods such as bag-of-words and TF-IDF are used to represent each text as a vector, which is then input into a classifier such as SVM or decision tree to obtain the final result. The judgment process of models built based on deep learning methods is as follows: Neural networks are used to obtain the effective features corresponding to the two texts, such as convolutional neural networks and recurrent neural networks. First, the two texts are cleaned and segmented separately. Then, methods based on neural network ideas, such as word2vec, are used to convert the two texts into dense distributed word vectors. Finally, neural networks such as CNN or LSTM are used to train the data corresponding to the word vectors to obtain the final result.
[0042] In one embodiment, such as Figure 2 As shown, Figure 2 This specification presents a training method for a text merging judgment model. This method can be implemented using a computer program and can run on a text merging judgment training device based on the von Neumann architecture. The computer program can be integrated into applications or run as a standalone utility application.
[0043] Specifically, the training methods for the text merging judgment model include:
[0044] S102. Obtain at least one positive sample group and at least one negative sample group.
[0045] Each positive sample group consists of two texts that cannot be merged, each text possessing independent and complete semantics. For example, if the two texts in a positive sample group come from text paragraphs published in textbooks, newspapers, or news websites, and are connected by ".", "!", or "?", then the two texts in this positive sample group are correctly segmented and cannot be merged. When a positive sample group is input into the text merging judgment model, the judgment result of the converged text merging judgment model should be "cannot be merged".
[0046] Each negative sample group consists of two texts that can be merged, meaning the two texts are semantically related and only have complete semantics when merged. For example, a long text may contain the symbols ",", "、", ":", and "——". Segmenting the text at any of these symbols yields two texts, each of which cannot express a complete meaning on its own. When a negative sample group is input into a text merging judgment model, the judgment result of the converged text merging judgment model should be "can be merged". In other words, in one embodiment, the method for obtaining two texts that can be merged to form a negative sample group is as follows: obtain at least one sample text to be segmented, and segment the sample text to be segmented according to a preset symbol in the at least one sample text to be segmented, thereby obtaining at least one negative sample group. The preset symbol can be any one of ",", "、", ":", and "——", or other symbols set as needed by those skilled in the art.
[0047] In one embodiment, such as Figure 3 The diagram shown is a schematic representation of a process for obtaining a negative sample group according to an embodiment of this specification, including the following steps:
[0048] S1022. Obtain the sample text to be segmented.
[0049] The sample text consists of multiple characters. The method for obtaining the sample text to be segmented can be any known and feasible method, and the specific content of the sample text can be arbitrary. Taking a medical scenario as an example: the sample text contains patient medical records, and the sample text is at least one question-and-answer pair for the case, containing questions raised by the patient regarding the case and answers given by the doctor to those questions. Alternatively, the sample text may be the doctor's diagnostic results and treatment plan for the case, such as "The patient exhibits severe anemia symptoms and should pay attention to their diet and meal times."
[0050] S1024. Determine the character located in the middle position of each sample text to be segmented as the target character.
[0051] Each sample text to be segmented consists of multiple characters, which are divided into symbolic characters and non-symbolic characters. Based on the number of characters in each sample text to be segmented, and using the reading order as the judgment order, the character located in the middle position is selected as the target character.
[0052] For example, the sample text to be segmented is "The patient exhibits severe anemia symptoms and should pay attention to diet and meal times". This text to be segmented includes 26 characters, including 2 characters of type symbol. Therefore, the target character "should" located at the 14th character is determined to be the target character.
[0053] S1026. Detect whether there is a preset symbol among the N characters to the left of the target character, and detect whether there is a preset symbol among the N characters to the right of the target character.
[0054] N is a positive integer greater than 1, set as needed by those skilled in the art. For example, N can be 3, 4, or 5. Using N as the window, the system detects whether a preset symbol exists in the left and right windows of the target character. The preset symbol can be any one of ",", "、", ":", and "—", or other symbols set as needed by those skilled in the art.
[0055] The order of detecting whether a preset symbol exists in the N characters to the left of the target character and whether a preset symbol exists in the N characters to the right of the target character can be according to the reading order. For example, if the reading order is from left to right, first detect whether a preset symbol exists in the N characters to the left of the target character. If yes, execute S1028. If no, that is, there is no preset symbol in the N characters to the left of the target character, then continue to detect whether a preset symbol exists in the N characters to the right of the target character. If yes, execute S1028. If no, that is, there is no preset symbol in the N characters to the left of the target character and no preset symbol in the N characters to the right of the target character, then execute S1022 and obtain the sample text to be segmented.
[0056] S1028. Segment each sample text to be segmented using a preset symbol as the boundary to obtain at least one negative sample group.
[0057] If a preset symbol is detected in the N characters to the left of the target character, or if a preset symbol is detected in the N characters to the right of the target character, the sample text to be segmented is divided into two texts using the preset symbol as the boundary. These two texts are then combined into a negative sample group.
[0058] In another embodiment, when a preset symbol is detected in N characters to the left of the target character and a preset symbol is detected in N characters to the right of the target character, the sample text to be segmented is segmented with the preset symbol as the boundary to obtain two negative sample groups.
[0059] For example, such as Figure 4 The diagram shown is a flowchart illustrating a method for obtaining negative sample groups according to an embodiment of this specification. The sample text 200 to be segmented is obtained, and the sample text 200 is divided into sample text 201 and sample text 202, using the target character located in the middle of the sample text 200 as the boundary. Further, it is determined whether a preset symbol exists in the left window 2011 and right window 2021 of the target character. For example, in… Figure 4In the target character, there is a preset symbol in the left window 2011. According to the preset symbol, the sample text 200 to be segmented is divided into sample text 203 and sample text 204. Sample text 203 and sample text 204 are input into the text merging judgment model as a negative sample group.
[0060] This embodiment provides a more reasonable and zero-cost sample construction method, which can not only reduce the manual annotation cost of constructing positive and negative sample groups, but also avoid the problems of under-segmentation, over-segmentation, and mis-segmentation of sample text when simply performing text segmentation according to preset symbols if there is misuse of symbols in the sample text.
[0061] In another embodiment, based on the position of each preset character in each sample text to be segmented, the sample text is segmented using each preset symbol as a boundary, resulting in a negative sample group corresponding to each preset symbol. For example, the sample text to be segmented is "The patient exhibits severe anemia symptoms and should pay attention to diet and meal times." This text includes 26 characters, including two symbols. The sample text to be segmented is divided into two negative sample groups: the first negative sample group includes "The patient exhibits severe anemia symptoms" and "should pay attention to diet," and the other negative sample group includes "should pay attention to diet" and "and meal times." The text segmentation method provided in this embodiment has simple logic and high efficiency in creating negative sample groups.
[0062] S104. Train the text merging judgment model using at least one positive sample group and at least one negative sample group until the text merging judgment model converges.
[0063] Obtain at least one positive sample group and at least one negative sample group. Input each positive and negative sample group into the text merging judgment model during training. Adjust the text merging judgment model based on the ideal results until the model converges. In this specification, the condition for the text merging judgment model to converge can be a pre-set number of training rounds or a stopping condition determined during training. The stopping condition can be that the loss function of the text merging judgment model converges to the expected value, or that the loss function reaches a stable value and then shows discrepancies.
[0064] The training process can include transfer learning, multi-task learning, and adversarial training, with data augmentation applied to at least one positive and one negative sample group. Transfer learning uses a model trained on a similar task as the initial model and retrains it on the original task. By sharing the knowledge learned by the model, transfer learning can accelerate the model's learning efficiency and improve its generalization ability. Multi-task learning also uses a model trained on a similar task as the initial model and retrains it on the original task. By sharing the knowledge learned by the model, transfer learning can accelerate the model's learning efficiency and improve its generalization ability. Data augmentation includes a series of techniques for generating new training samples. These techniques achieve this by randomly jittering and perturbing the original data while keeping the class labels unchanged. The goal of applying data augmentation is to increase the model's generalization ability. Adversarial training is an important method for enhancing model robustness. During adversarial training, at least one positive and at least one negative sample group are given small perturbations, causing the text merging judgment model to make mistakes. This allows the text merging judgment model to adapt to the perturbations during training, thereby enhancing its robustness.
[0065] In one embodiment, such as Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of a text merging judgment model provided in the embodiments of this specification. The text merging judgment model 40 includes: multiple encoders, at least one fully connected layer 402 and a judge 403. The multiple encoders include encoder 4011, encoder 4012, encoder 4013, ..., encoder 401M, where M is a positive integer greater than or equal to 2.
[0066] Multiple encoders 401 are used to encode the input text to be detected, so as to obtain multiple feature vectors corresponding to each text. The multiple encoders are one or more of the following: a bidirectional encoder representation of the BERT model, a recurrent neural network encoder, or a convolutional neural network encoder. Specifically, the Bidirectional Encoder Representation from Transformers (BERT) is a pre-trained language model based on Transformers, trained on a large-scale corpus using a masked language model (MLM) and next sentence prediction (NSP) multi-tasks. The recurrent neural network (RNN) is a type of recurrent neural network that takes sequence data as input, recurses in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain-like manner. It is understood that the embodiments in this specification also include other types of encoders, which are not limited thereto.
[0067] Fully connected layer 402 is used to perform fully connected processing on multiple feature vectors corresponding to two texts respectively, to obtain at least one connection result. In one embodiment, the number of fully connected layers 403 is one or more, and at least one fully connected layer 402 includes one or more of the following fully connected layers: a fully connected layer that connects all feature vectors in sequence, a fully connected layer that connects the feature vectors corresponding to the first character of each text, and a fully connected layer that connects the feature vector corresponding to the first character of one text with the feature vector corresponding to the last character of another text.
[0068] The judge 403 is used to determine whether at least two texts can be merged based on at least one connection result. Specifically, the judge 403 performs constraint processing on at least one connection result to obtain the probability that at least two texts can be merged; based on the probability that at least two texts can be merged, it determines whether at least two texts can be merged. For example, two texts to be detected are input into the text merging judgment model 40, multiple encoders 401 obtain multiple feature vectors corresponding to each text, at least one fully connected layer 402 connects the multiple feature vectors corresponding to each text to obtain at least one connection result, and finally the judge 403 performs constraint processing on at least one connection result to obtain a judgment result, determining whether the two texts to be detected can be merged.
[0069] Specifically, such as Figure 6 As shown, Figure 6This is a flowchart illustrating a text merging judgment model provided in the embodiments of this specification, which determines whether texts can be merged.
[0070] First, two texts to be tested are obtained: text 501 and text 502. Then, according to the word segmentation rules, texts 501 and 502 are segmented at the lowest granularity to obtain multiple word tokens corresponding to text 501 and text 502. A [CLS] category is set at the beginning of the multiple word tokens corresponding to text 501, and the multiple word tokens corresponding to text 501 and text 502 are connected by [SEP]. Finally, [SEP] is set as the end of the multiple word tokens corresponding to text 502.
[0071] Furthermore, the multiple encoders of the encoding layer 401 of the text classification model encode multiple word segmentation tokens corresponding to the text to be detected 501 and the text to be detected 502, respectively, to obtain the vector embedding corresponding to each word segmentation token. For example, the encoding layer 401 first outputs a 1 for each word segmentation token. × The 1024-bit vector is used as the first feature vector of this token segmentation. Then, multiple transformer layers encode multiple first feature vectors to produce a second feature vector, such as... Figure 6 As shown, the results include T1 to T N And T / 1 to T / M The transformer layers consist of 12 layers. The method for obtaining the second feature vectors from the first feature vector can be as follows: Identify the part-of-speech tags of keywords in the texts 501 and 502 to be detected. Keywords tend to contain more valid information, and part-of-speech tags include nouns, verbs, adjectives, adverbs, numbers, or foreign words. Input the first feature vector into the encoding layer 401. Through the keyword highlighting operation introduced in the encoding layer 401, based on the feature vectors in the first feature vector that represent keywords, highlight the feature vectors in the first feature vector that represent text information to obtain multiple second feature vectors corresponding to the texts 501 and 502 to be detected. It is understood that... Figure 6 The number of transformer layers and fully connected layer 402 shown is for illustrative purposes only, and this embodiment does not impose any limitations on this.
[0072] Finally, the character [CLS] is set at the beginning of the multiple second feature vectors corresponding to the text to be detected 501, and the multiple second feature vectors corresponding to the text to be detected 502 are connected by the character [SEP]. The character [CLS] is set as the end. The set vector statement is used as the input of the fully connected layer 402, and then the class label judge 403 performs the text merging judgment task to obtain the final output, that is, the judgment result of whether the text to be detected 501 and the text to be detected 502 can be merged.
[0073] This specification's embodiments reasonably construct at least one positive sample group and one negative sample group. The positive sample group includes texts that cannot be merged, and the negative sample group includes texts that can be merged. By using at least one positive and negative sample group, the text merging judgment model can learn in a self-supervised manner whether there is a mergeable relationship between two texts until the text merging judgment model converges, thereby improving the training efficiency of the text merging judgment model. Furthermore, by using at least one positive and negative sample group, the text merging judgment model can be trained in multiple rounds, so that the trained text merging judgment model has better anti-interference and robustness, and higher accuracy in performing the task of judging whether two texts are merged, thus obtaining merged texts with complete semantics, which is convenient for users to read and understand.
[0074] After introducing the design concept of the text merging judgment model in this specification, the application scenarios set up in this application will be briefly described below.
[0075] like Figure 7 The diagram illustrates a scenario in which a text merging judgment model provided in this application is applied. This application scenario includes a terminal device 602 and a server 601. The terminal device 602 and the server 601 can communicate via a communication network. In one embodiment, the communication network is a wired network or a wireless network. The terminal device 602 and the server 601 can be directly or indirectly connected via wired or wireless communication methods; this embodiment does not impose any limitations on this.
[0076] In this embodiment, the terminal device 602 is an electronic device used by a user. This electronic device can be a personal computer, mobile phone, tablet computer, laptop, e-book reader, or other computer device with a certain computing capability that runs instant messaging software and websites or social networking software and websites. The terminal device 602 can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these.
[0077] Server 601 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0078] The text classification model can be deployed on server 601 for training. Server 601 can store a large number of training samples, including at least one positive sample group and one negative sample group, for training the text merging judgment model. Optionally, after training the text merging judgment model based on the training method in the embodiments of this specification, the trained text merging judgment model can be directly deployed on server 601 or terminal device 602. Generally, the text merging judgment model is directly deployed on server 602. In the embodiments of this application, the text merging judgment model is often used to analyze the user-input question and the corresponding two texts to be detected, in order to determine whether the two texts to be detected can be merged.
[0079] In one possible application scenario, to reduce communication latency, servers 601 can be deployed in various regions, or for load balancing, different servers 601 can serve the regions corresponding to each terminal device 602. Multiple servers 601 can share data through blockchain, forming a data-sharing system. For example, terminal device 602 located at location a communicates with server 601, while terminal device 602 located at location b communicates with other servers 601.
[0080] In one embodiment, such as Figure 8 As shown, Figure 8 This specification provides a method for determining text merging in an embodiment. This method can be implemented using a computer program and can run on a text merging device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.
[0081] Specifically, the methods for determining text merging include:
[0082] S202. Obtain two texts to be detected.
[0083] The method for obtaining two texts to be detected can be to obtain texts input by the user on the mobile terminal 602 via voice, touch input, or other means, or to receive texts to be detected sent from the mobile terminal 602.
[0084] S204. Input the two texts to be detected into the text merging judgment model to obtain the judgment result of whether the two texts to be detected can be merged.
[0085] This embodiment trains a text merging judgment model using at least one positive sample group and one negative sample group until the model converges. This allows the model to be used to determine whether two texts should be merged, improving the accuracy of the judgment. Furthermore, the text merging judgment model provided in this embodiment combines currently popular natural language processing models by custom-adding one or more fully connected layers after multiple encoding layers. This performs feature compression processing on the multiple feature vectors obtained from the multiple encoding layers, further enhancing the algorithmic performance of the text merging judgment model.
[0086] It should be noted that the text merging judgment model provided in this application embodiment can be applied to various application scenarios that include text merging judgment. For example, it can be used for basic tasks such as text merging judgment in various natural language processing tasks in the medical, financial, or educational fields. However, such basic tasks are often crucial to subsequent tasks.
[0087] The following are embodiments of the apparatus described in this specification, which can be used to execute the embodiments of the methods described in this specification. For details not disclosed in the apparatus embodiments of this specification, please refer to the embodiments of the methods described in this specification.
[0088] Please see Figure 9 This diagram illustrates the structure of a training apparatus for a text merging judgment model provided in an exemplary embodiment of this specification. The text merging judgment apparatus can be implemented as all or part of the apparatus through software, hardware, or a combination of both. The apparatus includes a sample acquisition module 901 and a model training module 902.
[0089] The sample acquisition module 901 is used to acquire at least one positive sample group and at least one negative sample group, wherein the positive sample group includes two texts that cannot be merged and the negative sample group includes two texts that can be merged.
[0090] The model training module 902 is used to train the text merging judgment model using the at least one positive sample group and the at least one negative sample group until the text merging judgment model converges.
[0091] In one embodiment, the sample acquisition module 901 includes:
[0092] The sample acquisition unit is used to acquire at least one sample text to be segmented;
[0093] The sample segmentation unit is used to segment the sample text to be segmented according to preset symbols in the at least one sample text to be segmented, so as to obtain at least one negative sample group.
[0094] In one embodiment, the sample segmentation unit includes:
[0095] The target determination subunit is used to determine the character located in the middle position of each of the sample texts to be segmented as the target character;
[0096] A symbol detection subunit is used to detect whether the preset symbol exists in the N characters to the left of the target character, and to detect whether the preset symbol exists in the N characters to the right of the target character, where N is an integer greater than 1;
[0097] The target segmentation subunit is used to segment each sample text to be segmented with the preset symbol as the boundary if a preset symbol exists in N characters to the left of the target character or in N characters to the right of the target character, thereby obtaining at least one negative sample group.
[0098] In one embodiment, the sample segmentation unit includes:
[0099] The symbol segmentation subunit is used to segment the sample text to be segmented based on the position of each preset character in each sample text to be segmented, with each preset symbol as the boundary, to obtain the negative sample group corresponding to each preset symbol.
[0100] In one embodiment, the text merging judgment model includes: multiple encoders, at least one fully connected layer, and a judge;
[0101] The plurality of encoders are used to encode the text to obtain a plurality of feature vectors corresponding to the text;
[0102] The at least one fully connected layer is used to perform fully connected processing on multiple feature vectors corresponding to the two texts respectively, to obtain at least one connection result;
[0103] The judge is used to determine whether the at least two texts can be merged based on the at least one connection result.
[0104] In one embodiment, the at least one fully connected layer includes one or more of the following fully connected layers: a fully connected layer that sequentially connects all feature vectors, a fully connected layer that connects the feature vectors corresponding to the first character of each text, and a fully connected layer that connects the feature vector corresponding to the first character of one text with the feature vector corresponding to the last character of another text.
[0105] In one embodiment, the determiner is specifically used for:
[0106] Constrain the at least one connection result to obtain the probability that the at least two texts can be merged;
[0107] Based on the probability that the at least two texts can be merged, determine whether the at least two texts can be merged.
[0108] In one embodiment, the plurality of encoders is one or more of the following: a bidirectional encoder representing an encoder for a BERT model, an encoder for a recurrent neural network, or an encoder for a convolutional neural network.
[0109] This specification describes an embodiment that constructs at least one positive sample group and one negative sample group. The positive sample group includes text that cannot be merged, and the negative sample group includes text that can be merged. By using at least one positive and negative sample group, the text merging judgment model can self-supervisedly learn whether a mergeable relationship exists between two texts until the model converges, thereby improving the training efficiency of the text merging judgment model. Furthermore, by using at least one positive and negative sample group, the text merging judgment model undergoes multiple rounds of training, resulting in a trained model with better anti-interference and robustness, and higher accuracy in judging whether two texts are merged. This yields merged text with complete semantics, facilitating user reading and understanding. It should be noted that the training device for the text merging judgment model provided in the above embodiment is only illustrative of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the training device for the text merging judgment model and the training method for the text merging judgment model provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.
[0110] Please see Figure 10 This diagram illustrates the structure of a text merging and determining device provided in an exemplary embodiment of this specification. The text merging and determining device can be implemented as all or part of a device through software, hardware, or a combination of both. The device includes a text acquisition module 1001 and a result acquisition module 1002.
[0111] The text acquisition module 1001 is used to acquire two texts to be detected.
[0112] The result acquisition module 1002 is used to input the two texts to be detected into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model described in the above embodiment.
[0113] This embodiment trains a text merging judgment model using at least one positive sample group and one negative sample group until the model converges. This allows the model to be used to determine whether two texts should be merged, improving the accuracy of the judgment. Furthermore, the text merging judgment model provided in this embodiment combines currently popular natural language processing models by custom-adding one or more fully connected layers after multiple encoding layers. This performs feature compression processing on the multiple feature vectors obtained from the multiple encoding layers, further enhancing the algorithmic performance of the text merging judgment model.
[0114] It should be noted that the text merging judgment device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the text merging judgment method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the text merging judgment device and the text merging judgment method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.
[0115] The example numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the examples.
[0116] This specification also provides a computer storage medium that can store multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-8 The text merging judgment method described in the illustrated embodiment can be found in the following document for its specific execution process. Figures 1-8 The specific details of the illustrated embodiments will not be elaborated here.
[0117] This specification also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-8 The text merging judgment method described in the illustrated embodiment can be found in the following document for its specific execution process. Figures 1-8 The specific details of the illustrated embodiments will not be elaborated here.
[0118] Please see Figure 11 This document provides a schematic diagram of the structure of an electronic device as an embodiment of the present specification. Figure 11As shown, the electronic device 1110 may include: at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105, and at least one communication bus 1102.
[0119] The communication bus 1102 is used to realize the connection and communication between these components.
[0120] The user interface 1103 may include a display screen and a camera. Optionally, the user interface 1103 may also include a standard wired interface and a wireless interface.
[0121] The network interface 1104 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0122] The processor 1101 may include one or more processing cores. The processor 1101 connects to various parts of the server 1100 via various interfaces and lines, and performs various functions and processes data of the server 1100 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1105, and by calling data stored in the memory 1105. Optionally, the processor 1101 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 1101 may integrate one or a combination of several of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 1101 and may be implemented as a separate chip.
[0123] The memory 1105 may include random access memory (RAM) or read-only memory. Optionally, the memory 1105 may include a non-transitory computer-readable storage medium. The memory 1105 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1105 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1105 may also be at least one storage device located remotely from the aforementioned processor 1101. Figure 11 As shown, the memory 1105, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program. The application program is an application program for training a text merging judgment model and / or an application program for a text merging judgment method.
[0124] exist Figure 11 In the illustrated electronic device 1100, the user interface 1103 is mainly used to provide an input interface for the user and to obtain the user's input data; while the processor 1101 can be used to call the training application of the text merging judgment model stored in the memory 1105, and specifically perform the following operations:
[0125] Obtain at least one positive sample group and at least one negative sample group, wherein the positive sample group comprises two texts that cannot be merged and the negative sample group comprises two texts that can be merged;
[0126] The text merging judgment model is trained using at least one positive sample group and at least one negative sample group until the text merging judgment model converges.
[0127] In one embodiment, the processor 1101 performs the acquisition of at least one negative sample group, specifically by:
[0128] Obtain at least one sample text to be segmented;
[0129] Based on preset symbols in the at least one sample text to be segmented, the sample text to be segmented is segmented to obtain at least one negative sample group.
[0130] In one embodiment, the processor 1101 executes the step of segmenting the sample text to be segmented according to preset characters in the at least one sample text to be segmented, to obtain at least one negative sample group, specifically:
[0131] The character located in the middle position of each of the sample texts to be segmented is determined as the target character;
[0132] Detect whether the preset symbol exists in the N characters to the left of the target character, and detect whether the preset symbol exists in the N characters to the right of the target character, where N is an integer greater than 1;
[0133] If a preset symbol exists in the N characters to the left of the target character, or in the N characters to the right of the target character, then each of the sample texts to be segmented is segmented using the preset symbol as the boundary, to obtain at least one negative sample group.
[0134] In one embodiment, the processor 1101 executes the step of segmenting the sample text to be segmented according to preset characters in the at least one sample text to be segmented, to obtain at least one negative sample group, specifically:
[0135] Based on the position of each preset character in each sample text to be segmented, the sample text to be segmented is divided with each preset symbol as the boundary, to obtain the negative sample group corresponding to each preset symbol.
[0136] In one embodiment, the text merging judgment model includes: multiple encoders, at least one fully connected layer, and a judge;
[0137] The plurality of encoders are used to encode the text to obtain a plurality of feature vectors corresponding to the text;
[0138] The at least one fully connected layer is used to perform fully connected processing on multiple feature vectors corresponding to the two texts respectively, to obtain at least one connection result;
[0139] The judge is used to determine whether the at least two texts can be merged based on the at least one connection result.
[0140] In one embodiment, at least one fully connected layer includes one or more of the following fully connected layers: a fully connected layer that sequentially connects all feature vectors, a fully connected layer that connects the feature vectors corresponding to the first character of each text, and a fully connected layer that connects the feature vector corresponding to the first character of one text with the feature vector corresponding to the last character of another text.
[0141] In one embodiment, the determiner is specifically used for:
[0142] Constrain the at least one connection result to obtain the probability that the at least two texts can be merged;
[0143] Based on the probability that the at least two texts can be merged, determine whether the at least two texts can be merged.
[0144] In one embodiment, the plurality of encoders is one or more of the following: a bidirectional encoder representing an encoder for a BERT model, an encoder for a recurrent neural network, or an encoder for a convolutional neural network.
[0145] In one embodiment, processor 1101 can be used to invoke a text merging and judgment application stored in memory 1105, and specifically perform the following operations:
[0146] Obtain two texts to be detected;
[0147] The two texts to be detected are input into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model described in the above embodiment.
[0148] This specification describes embodiments that reasonably construct at least one positive sample group and one negative sample group. The positive sample group includes text that cannot be merged, and the negative sample group includes text that can be merged. Through at least one positive and negative sample group, the text merging judgment model can self-supervisedly learn whether a mergeable relationship exists between two texts until the text merging judgment model converges, thereby improving the training efficiency of the text merging judgment model. Furthermore, by using at least one positive and negative sample group, the text merging judgment model can be trained in multiple rounds, resulting in a trained text merging judgment model with better anti-interference and robustness, and higher accuracy in performing the task of judging whether two texts are merged, thus obtaining merged text with complete semantics, which is convenient for users to read and understand. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0149] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A training method for a text merging judgment model, the method comprising: Obtain at least one positive sample group and at least one negative sample group, wherein the positive sample group comprises two texts that cannot be merged and the negative sample group comprises two texts that can be merged; The text merging judgment model is trained using the at least one positive sample group and the at least one negative sample group until the text merging judgment model converges. The step of obtaining at least one negative sample group includes: Obtain at least one sample text to be segmented; Based on the preset symbols in the at least one sample text to be segmented, the sample text to be segmented is segmented to obtain at least one negative sample group; The step of segmenting the sample text to be segmented according to preset characters in the at least one sample text to be segmented, to obtain at least one negative sample group, includes: The character located in the middle position of each of the sample texts to be segmented is determined as the target character; Detect whether the preset symbol exists in the N characters to the left of the target character, and detect whether the preset symbol exists in the N characters to the right of the target character, where N is an integer greater than 1; If a preset symbol exists in the N characters to the left of the target character, or in the N characters to the right of the target character, then each of the sample texts to be segmented is segmented using the preset symbol as the boundary, to obtain at least one negative sample group.
2. The method according to claim 1, wherein the step of segmenting the sample text to be segmented according to preset characters in the at least one sample text to be segmented, to obtain at least one negative sample group, includes: Based on the position of each preset character in each sample text to be segmented, the sample text to be segmented is divided with each preset symbol as the boundary, to obtain the negative sample group corresponding to each preset symbol.
3. The method according to claim 1, wherein the text merging judgment model comprises: Multiple encoders, at least one fully connected layer, and a decision layer; The plurality of encoders are used to encode the text to obtain a plurality of feature vectors corresponding to the text; The at least one fully connected layer is used to perform fully connected processing on multiple feature vectors corresponding to the two texts respectively, to obtain at least one connection result; The judge is used to determine whether the at least two texts can be merged based on the at least one connection result.
4. The method according to claim 3, wherein the at least one fully connected layer comprises one or more of the following fully connected layers: a fully connected layer that sequentially connects all feature vectors, a fully connected layer that connects the feature vectors corresponding to the first character of each text, and a fully connected layer that connects the feature vector corresponding to the first character of one text with the feature vector corresponding to the last character of another text.
5. The method according to claim 3, wherein the determiner is specifically used for: Constrain the at least one connection result to obtain the probability that the at least two texts can be merged; Based on the probability that the at least two texts can be merged, determine whether the at least two texts can be merged.
6. The method according to claim 3, wherein the plurality of encoders is one or more of the following: a bidirectional encoder representing an encoder of a BERT model, an encoder of a recurrent neural network, and an encoder of a convolutional neural network.
7. A method for determining text merging, the method comprising: Obtain two texts to be detected; The two texts to be detected are input into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model according to any one of claims 1 to 6.
8. A training device for a text merging judgment model, the device comprising: A sample acquisition module is used to acquire at least one positive sample group and at least one negative sample group, wherein the positive sample group includes two texts that cannot be merged, and the negative sample group includes two texts that can be merged. The model training module is used to train the text merging judgment model using the at least one positive sample group and the at least one negative sample group until the text merging judgment model converges. The sample acquisition module includes: The sample acquisition unit is used to acquire at least one sample text to be segmented; A sample segmentation unit is used to segment the sample text to be segmented according to preset symbols in the at least one sample text to be segmented, so as to obtain at least one negative sample group; The sample segmentation unit includes: The target determination subunit is used to determine the character located in the middle position of each of the sample texts to be segmented as the target character; A symbol detection subunit is used to detect whether the preset symbol exists in the N characters to the left of the target character, and to detect whether the preset symbol exists in the N characters to the right of the target character, where N is an integer greater than 1; The target segmentation subunit is used to segment each sample text to be segmented with the preset symbol as the boundary if a preset symbol exists in N characters to the left of the target character or in N characters to the right of the target character, thereby obtaining at least one negative sample group.
9. An apparatus for determining text merging, the apparatus comprising: The text acquisition module is used to acquire two texts to be detected. The result acquisition module is used to input the two texts to be detected into the text merging judgment model to obtain a judgment result on whether the two texts to be detected can be merged; wherein, the text merging judgment model is a model trained using the training method of the text merging judgment model according to any one of claims 1 to 6.
10. A computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 7.
11. A computer program product storing a plurality of instructions adapted for loading by a processor and executing the method steps of any one of claims 1 to 7.
12. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Text processing method and device, equipment and storage medium
CN113361260A