Method for training a language representation model, method and apparatus for searching for a statement

By using training statements in the target field and adjusting the model structure in the language representation model, the problem of large amount of model calculation and insufficient accuracy is solved, and faster and more accurate information retrieval is achieved.

CN114648030BActive Publication Date: 2025-07-08阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210302920.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-07-08
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

The existing language representation model has a large amount of computation and cannot accurately represent relevant statements in the target field, resulting in insufficient application accuracy of the model in the target field.

Method used

By obtaining training statements in the target field, the pre-trained language representation model is trained and the model structure is adjusted, so that the semantic feature extraction layer directly receives the output of the phrase feature extraction layer, adjusts the weights of word features and text features based on the weight value, and optimizes the model structure to adapt to the target field.

Benefits of technology

It improves the running speed and accuracy of the language representation model in the target field, and improves the speed and accuracy of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114648030B_ABST
    Figure CN114648030B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method for training a language representation model, a method for searching for a statement, and an apparatus. The method includes: obtaining a target training statement, where the target training statement is obtained by collecting statements in a target field to which the language representation model is applied; training a pre-trained language representation model according to the target training statement to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer, and the input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, and i and j are integers greater than or equal to 1. Through some embodiments of the present application, the running speed of the language representation model can be improved, and the parameters in the target language representation model can be made more suitable for application in the target field, thereby improving the accuracy of the language representation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of natural language processing, and particularly to a method for training a language representation model, a method for searching for a statement, and an apparatus therefor. Background Art

[0002] In related technologies, with the wide popularization of artificial intelligence in various industries, language representation models are well known. However, existing language representation models have at least the following defects: they have a huge amount of computation, and in the process of using a language representation model, they are usually limited to models trained with training statements in the general domain, resulting in the inability of the language representation model to accurately represent relevant statements in the target domain.

[0003] Therefore, how to obtain a better-performing language representation model has become a problem to be solved. Summary of the Invention

[0004] Embodiments of the present application provide a method for training a language representation model, a method for searching for a statement, and an apparatus therefor. Through some embodiments of the present application, at least the running speed of the language representation model can be improved, and the parameters in the target language representation model can be made more suitable for application in the target domain, thereby improving the accuracy of the language representation model.

[0005] In a first aspect, embodiments of the present application provide a method for training a language representation model. The method includes: obtaining target training statements, where the target training statements are obtained by collecting statements in the target domain to which the language representation model is applied; training a pre-trained language representation model according to the target training statements to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer, and the input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, where i and j are integers greater than or equal to 1.

[0006] Therefore, on the one hand, by inputting statements in the target domain into the model for training in embodiments of the present application, the pre-trained language representation model can accurately express statement features related to the target domain, thereby improving the accuracy of the model. On the other hand, in some embodiments of the present application, by directly inputting the features output by the j-th layer of the phrase feature extraction layer into the corresponding nodes in the i-th layer, that is, the input of each node in some nodes included in the i-th layer is the output of the corresponding node in the phrase feature extraction layer, the pre-trained language representation model can better learn the phrase features of the target training statements and improve the prediction accuracy of the model.

[0007] In combination with the first aspect, in an implementation manner of the present application, each layer in the pre-trained language representation model includes two types of nodes. Among them, the first type of nodes is used to extract text features, and the second type of nodes is used to extract character features; wherein, the input of the second type of nodes included in the i-th layer is the output of the second type of nodes included in the j-th layer; the input of the first type of nodes included in the i-th layer is the output of the first type of nodes included in the i-1-th layer; wherein, the text features are used to represent the overall semantic features of the target training statement, and the character features are used to represent the semantic features of a character in the target training statement.

[0008] Therefore, by using the phrase features extracted in the phrase feature extraction layer as the input of the semantic feature extraction layer in the embodiments of the present application, the semantic feature extraction layer can better learn the phrase features, thereby improving the expression ability of the target language representation model for the phrase features.

[0009] In combination with the first aspect, in an implementation manner of the present application, the semantic feature extraction layer includes L layers, and the phrase feature extraction layer includes K layers, where L and K are integers greater than 1. Among them, the i-th layer in the semantic feature extraction layer is the first layer in the K layers; the j-th layer in the phrase feature extraction layer is the last layer in the L layers.

[0010] Therefore, by using the output of the last layer in the phrase feature extraction layer as the input of the first layer in the semantic feature extraction layer in the embodiments of the present application, the semantic feature extraction layer can obtain the phrase features with better learning degree in the phrase feature extraction layer, thereby expressing the input statement more accurately.

[0011] In a second aspect, embodiments of the present application provide a method for finding a statement. The method includes: obtaining a statement to be matched; inputting the statement to be matched into the target language representation model obtained in the implementation manner of the first aspect, and obtaining a target statement that matches the statement to be matched through the target language representation model.

[0012] Therefore, by using the target language representation model to search for the target statement in the embodiments of the present application, the speed and accuracy of the search can be improved, thereby realizing efficient retrieval of information.

[0013] In combination with the second aspect, in an implementation manner of the present application, obtaining the target statement that matches the statement to be matched through the target language representation model includes: extracting a to-be-matched representation vector of the statement to be matched; matching the to-be-matched representation vector with at least one group of candidate representation vectors to obtain the target statement, where one group of candidate representation vectors is used to represent a candidate statement, and one group of candidate representation vectors corresponds to one candidate statement.

[0014] Therefore, in the embodiments of the present application, by matching the to-be-matched characterization vector with at least one group of candidate characterization vectors, one or more target sentences that match the to-be-matched sentence can be found from at least one candidate sentence.

[0015] Combined with the second aspect, in an implementation manner of the present application, the matching the to-be-matched characterization vector with at least one group of candidate characterization vectors to obtain the target sentence includes: calculating a target similarity value between the to-be-matched characterization vector and each group of candidate characterization vectors in the at least one group of candidate characterization vectors based on a weight value, where the target similarity value is used to characterize the similarity degree between the to-be-matched characterization vector and each group of candidate characterization vectors, and the weight value is used to adjust the weight between the extracted character features and the extracted text features; obtaining the target sentence from the at least one candidate sentence through the target similarity value.

[0016] Therefore, different from the related art, in the embodiments of the present application, the weight between the extracted character features and the text features is adjusted through the weight value to obtain the target sentence, and the weight between the character features and the text features can be adjusted arbitrarily according to the actual situation, so that more accurate target sentences can be recommended.

[0017] Combined with the second aspect, in an implementation manner of the present application, the to-be-matched characterization vector includes a to-be-matched text semantic characterization sub-vector and a to-be-matched character semantic characterization sub-vector, the candidate characterization vector corresponding to the K-th candidate sentence includes a K-th candidate text semantic characterization sub-vector and a K-th candidate character semantic characterization sub-vector, the weight value corresponding to the K-th candidate sentence includes a K-th text weight value and a K-th character weight value, and the sum of the K-th text weight value and the K-th character weight value is 1; the calculating the target similarity value between the to-be-matched characterization vector and each group of candidate characterization vectors in the at least one group of candidate characterization vectors includes: calculating a K-th text similarity value between the to-be-matched text semantic characterization sub-vector and the K-th candidate text semantic characterization sub-vector, where K is an integer greater than or equal to 1; calculating a K-th character similarity value according to the to-be-matched character semantic characterization sub-vector and the K-th candidate character semantic characterization sub-vector; calculating the product of the K-th text similarity value and the K-th text weight value to obtain a first product; calculating the product of the K-th character similarity value and the K-th character weight value to obtain a second product; calculating the sum of the first product and the second product to obtain the target similarity value corresponding to the K-th candidate sentence.

[0018] Therefore, in the embodiments of the present application, the target similarity value corresponding to the K-th candidate sentence is calculated through the K-th text similarity value and the K-th character similarity value, and the weight between the overall semantic feature and each character feature can be allocated, so that the calculated target similarity value can better conform to the actual situation of the target field, and the found target sentence is more accurate.

[0019] In a third aspect, an embodiment of the present application provides an apparatus for training a language representation model. The apparatus includes: a training statement acquisition module configured to acquire a target training statement, where the target training statement is obtained by collecting statements in a target field to which the language representation model is applied; a model training module configured to train a pre-trained language representation model according to the target training statement to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer, and the input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, and i and j are integers greater than or equal to 1.

[0020] In combination with the third aspect, in an implementation manner of the present application, each layer in the pre-trained language representation model includes two types of nodes, where the first type of nodes is used to extract text features, and the second type of nodes is used to extract character features; among them, the input of the second type of nodes included in the i-th layer is the output of the second type of nodes included in the j-th layer; the input of the first type of nodes included in the i-th layer is the output of the first type of nodes included in the i-1-th layer; among them, the text features are used to represent the overall semantic features of the target training statement, and the character features are used to represent the semantic features of a character in the target training statement.

[0021] In combination with the third aspect, in an implementation manner of the present application, the semantic feature extraction layer includes L layers, and the phrase feature extraction layer includes K layers, where L and K are integers greater than 1, where the i-th layer in the semantic feature extraction layer is the first layer in the K layers; the j-th layer in the phrase feature extraction layer is the last layer in the L layers.

[0022] In a fourth aspect, an embodiment of the present application provides an apparatus for searching for a statement. The apparatus includes: a statement acquisition module configured to acquire a statement to be matched; a statement matching module configured to input the statement to be matched into the target language representation model obtained by using the implementation manner of the first aspect, and obtain a target statement that matches the statement to be matched through the target language representation model.

[0023] In combination with the fourth aspect, in an implementation manner of the present application, the statement matching module is further configured to: extract a to-be-matched representation vector of the statement to be matched; match the to-be-matched representation vector with at least one group of candidate representation vectors to obtain the target statement, where one group of candidate representation vectors is used to represent a candidate statement, and one group of candidate representation vectors corresponds to one candidate statement.

[0024] In combination with the fourth aspect, in an implementation of this application, the statement matching module is further configured to: calculate a target similarity value between the to-be-matched characterization vector and each group of candidate characterization vectors in the at least one group of candidate characterization vectors based on a weight value, where the target similarity value is used to characterize the similarity degree between the to-be-matched characterization vector and each group of candidate characterization vectors, and where the weight value is used to adjust the weight between the extracted word features and the extracted text features; obtain the target statement from the at least one candidate statement through the target similarity value.

[0025] In combination with the fourth aspect, in an implementation of this application, the to-be-matched characterization vector includes a to-be-matched text semantic characterization sub-vector and a to-be-matched word semantic characterization sub-vector, the candidate characterization vector corresponding to the Kth candidate statement includes a Kth candidate text semantic characterization sub-vector and a Kth candidate word semantic characterization sub-vector, the weight value corresponding to the Kth candidate statement includes a Kth text weight value and a Kth word weight value, and the sum of the Kth text weight value and the Kth word weight value is 1; the statement matching module is further configured to: calculate a Kth text similarity value between the to-be-matched text semantic characterization sub-vector and the Kth candidate text semantic characterization sub-vector, where K is an integer greater than or equal to 1; calculate a Kth word similarity value according to the to-be-matched word semantic characterization sub-vector and the Kth candidate word semantic characterization sub-vector; calculate the product of the Kth text similarity value and the Kth text weight value to obtain a first product; calculate the product of the Kth word similarity value and the Kth word weight value to obtain a second product; calculate the sum of the first product and the second product to obtain a target similarity value corresponding to the Kth candidate statement.

[0026] Fifth aspect, an embodiment of this application provides an electronic device, including: a processor, a memory, and a bus; the processor is connected to the memory through the bus, and the memory stores computer-readable instructions, which, when executed by the processor, are used to implement the methods described in the implementations of the first aspect and the second aspect.

[0027] Sixth aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, it is used to implement the methods described in the implementations of the first aspect and the second aspect. Description of the Drawings

[0028] Figure 1 It is a schematic diagram of a scenario for finding a statement shown in an embodiment of this application;

[0029] Figure 2 It is a flowchart of a method for training a language characterization model shown in an embodiment of this application;

[0030] Figure 3 One of the schematic diagrams showing the structural composition of the language representation model shown in the embodiments of the present application;

[0031] Figure 4 Another of the schematic diagrams showing the structural composition of the language representation model shown in the embodiments of the present application;

[0032] Figure 5 The flowchart of the method for finding statements shown in the embodiments of the present application;

[0033] Figure 6 One of the schematic diagrams of specific embodiments of the method for finding statements shown in the embodiments of the present application;

[0034] Figure 7 Another of the schematic diagrams of specific embodiments of the method for finding statements shown in the embodiments of the present application;

[0035] Figure 8 The block diagram showing the composition of the device for training the language representation model shown in the embodiments of the present application;

[0036] Figure 9 The block diagram showing the composition of the device for finding statements shown in the embodiments of the present application;

[0037] Figure 10 The schematic diagram showing the composition of the electronic device shown in the embodiments of the present application. Detailed implementation manners

[0038] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings below is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.

[0039] Embodiments of the present application can be applied to the field of matching sentence search. For example, some embodiments of the present application search for a target sentence based on a sentence to be matched input by a user, and feed back the target sentence to the user's scenario or feed back the response sentence corresponding to the target sentence (the correspondence between the response sentence and the target sentence needs to be stored in advance) to the user. In order to improve the problems in the background technology, in some embodiments of the present application, first, it is necessary to collect sentences in the target field to which the neural network model is specifically applied, and train the model with these sentences as training sentences (i.e., target training sentences) to obtain a target language representation model. Finally, the target sentence or response sentence that matches the input sentence to be matched is obtained according to the trained target language representation model.

[0040] For example, in some embodiments of the present application, the process of obtaining a matching sentence for a sentence to be matched exemplarily includes: first, retraining the pre-trained language representation model using the training sentence of the target domain to obtain a target language representation model suitable for the target domain. Then, the sentence to be matched is input into the target language representation model, and a representation vector to be matched that can represent the sentence to be matched is obtained through model calculation. Next, at least one group of candidate representation vectors is obtained, and the representation vector to be matched is matched with each group of at least one group of candidate representation vectors respectively to obtain a corresponding target similarity value. Finally, the target sentence is determined based on the target similarity value.

[0041] It should be noted that the target domain is the specific application domain of the language representation model. For example, if the language representation model is to be applied in the insurance industry, the target domain refers to the insurance field, and the corresponding target training sentences are the training sentences related to the insurance field.

[0042] Figure 1 A schematic diagram of a search statement (specifically referring to a search target statement) in some embodiments of the present application is provided, and the scenario includes a user 110, a client 120, and a server 130. Specifically, when the user 110 needs to obtain a reply statement or needs to search for information, the user 110 inputs a question text (i.e., a sentence to be matched) to the client 120. After receiving the sentence to be matched, the client 120 sends the sentence to be matched to the server 130. Thereafter, the server 130 selects a target sentence from at least one candidate sentence based on the sentence to be matched, and feeds the target sentence back to the client 120 for display.

[0043] It should be noted that the target sentence can be a best matching result, or can be a plurality of matching results selected according to the sorting after being sorted according to the target similarity value.

[0044] For example, if user 110 wants to obtain the address of an insurance company, the user enters the question text "What is the address of the insurance company?" in the client 120. After receiving the question text, the server 130 selects the three target statements with the highest similarity to the above question text from at least one candidate statement, and feeds them back to the client 120 for display as reply statements.

[0045] Different from the embodiments of the present application, in the related art, the model used to represent statements usually has a large computational amount, and is usually limited to the model trained using training statements in the general domain, resulting in the inability of the language representation model to accurately represent relevant statements in the target domain. However, the embodiments of the present application train the model with training statements in the target domain and change the input of some nodes in the language representation model, thus solving the problems in the related art.

[0046] The following exemplarily elaborates on the solutions for training the language representation model provided by some embodiments of the present application. It can be understood that the technical solutions of the method for training the language representation model in the embodiments of the present application can be applied to any server, and the obtained target language representation model after training will be applied to the method for searching statements in the embodiments of the present application.

[0047] At least to solve the problems existing in the background technology, such as Figure 2 As shown, some embodiments of the present application provide a method for training a language representation model, and the method includes:

[0048] S210, obtain target training statements.

[0049] It should be noted that the target training statements are obtained by collecting statements in the target domain to which the language representation model is applied. As a specific embodiment of the present application, the target training statements may be a set of question-and-answer statements. As another specific embodiment of the present application, the target training statements may also be a document.

[0050] For example, if the target language representation model after training is used in the financial field, then the training statements related to the financial field (i.e., target training statements) are input into the pre-trained language representation model for training.

[0051] S220, train the pre-trained language representation model according to the target training statements to obtain the target language representation model.

[0052] It should be noted that the pre-trained language representation model is obtained by improving the existing BERT (Bidirectional Encoder Representation from Transformers) model in terms of structure (changing the fully connected structure of the original network model). In the embodiments of the present application, the training sentences in the target domain are input into the pre-trained language representation model for further training to obtain a target language representation model applicable to the target domain.

[0053] It can be understood that the BERT model also includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer. Different from the present application, each layer in the BERT model is connected and transmits features layer by layer in a fully connected form.

[0054] The structure of the pre-trained language representation model in the embodiments of the present application will be described below.

[0055] In some embodiments of the present application, the schematic structural diagram of the pre-trained language representation model is as Figure 3 shown. The pre-trained language representation model sequentially includes an input layer 310, a phrase feature extraction layer 320, a syntactic feature extraction layer 330, a semantic feature extraction layer 340, and an output layer 350. Among them, the input of each node in part of the nodes in the i-th layer of the semantic feature extraction layer 340 is the output of the corresponding node in the j-th layer in the phrase feature extraction layer 320.

[0056] That is to say, different from the structure of layer-by-layer transmission in the related BERT model, in the embodiments of the present application, the input of a part of the nodes in the i-th layer of the semantic feature extraction layer is not the output of the previous layer in the related art, but the output of the j-th layer in the phrase feature extraction layer. Adopting the connection method of the present application can enable some nodes in the semantic feature extraction layer to directly obtain the phrase features extracted by the phrase feature extraction layer, so as to achieve the purpose of optimizing the model structure and improving the model operation speed.

[0057] It should be noted that for those skilled in the art, it belongs to common general knowledge to determine which layers of a specific language feature model belong to the phrase feature extraction layer, the syntactic feature extraction layer, and the semantic feature extraction layer. The i-th layer in the above semantic feature extraction layer can be obtained by changing the structure of the original layer in the BERT model (that is, directly changing the connection relationship between the original layer and the previous layer), or by adding a layer in the semantic feature extraction layer of the BERT model.

[0058] It can be understood that each layer in the pre-trained language representation model includes two types of nodes. Among them, the first type of nodes is used to extract text features, and the second type of nodes is used to extract character features. That is to say, during the training of the pre-trained language representation model, the nodes (i.e., the first type of nodes) in a certain layer of the semantic feature extraction layer that are used to extract the overall text features receive the features from the previous layer of this layer, while the nodes (i.e., the second type of nodes) in a certain layer of the semantic feature extraction layer that are used to extract the character features of each word receive the phrase features from the phrase feature extraction layer.

[0059] As Figure 3 shown, the first type of nodes included in the j-th layer of the phrase feature extraction layer is node 301, and the second type of nodes is node 302; the first type of nodes included in the i-th layer of the semantic feature extraction layer is node 303, and the second type of nodes is node 304. Among them, preferably, there is one first type of node in each layer and multiple second type of nodes. As Figure 3 shown, in some embodiments of the present application, the input of the second type of nodes 304 included in the i-th layer is the output of the second type of nodes 302 included in the j-th layer, and the input of the first type of nodes 303 included in the i-th layer is the output of the first type of nodes 305 included in the (i - 1)-th layer. Different from the embodiments of the present application, if the i-th layer is a layer in the existing network, the input of the second type of nodes 304 included in the i-th layer of the prior art is the output of the second type of nodes included in the (i - 1)-th layer.

[0060] It should be noted that the text features calculated by the first type of nodes in each layer of the pre-trained language representation model are used to represent the overall semantic features of the target training sentence. The character features calculated by the second type of nodes in each layer are used to represent the semantic features of each character in the target training sentence. Among them, in the case where the target training sentence is segmented into multiple words, the character features calculated by the second type of nodes can also be used to represent the semantic features of each word in the target training sentence.

[0061] For example: the target training sentence is "What types of insurance are included", then in order to accurately express the overall meaning of the above target training sentence, a first type of node is placed in each layer to represent the overall meaning of the above target training sentence. And in the model, the above target training sentence will also be split into multiple characters. For example, the multiple characters include "insurance", "of", etc., and the second type of nodes corresponding to the above multiple characters in the model are used to calculate the semantic features of each character.

[0062] Therefore, in the embodiments of the present application, by inputting the phrase features in the phrase feature extraction layer into the semantic feature extraction layer, the semantic feature extraction layer can better learn the phrase features, thereby improving the expression ability of the target language representation model for the phrase features. Moreover, since the above neural network is not fully connected, the structure of the target language representation model can be simplified, thereby improving the running speed of the target language representation model.

[0063] In some embodiments of the present application, for all layers other than the i-th layer in the semantic feature extraction layer in the pre-trained language representation model, the input of the first type of nodes included in all other layers is the output of the previous layer, and the input of the second type of nodes included in all other layers is also the output of the previous layer.

[0064] That is to say, as Figure 3 shown, both the first type of nodes 301 and the second type of nodes 302 included in the j-th layer in the phrase feature extraction layer 320 are transmitted to the syntactic feature extraction layer in a fully connected form. At the same time, the features of each layer included in the phrase feature extraction layer and the syntactic feature extraction layer are transmitted in the form of a fully connected layer.

[0065] Therefore, in the related art, the features of each layer in the language representation model are transmitted in a fully connected form. Different from the related art, in the embodiments of the present application, the i-th layer in the semantic feature extraction layer is not connected to the (i - 1)-th layer, but receives the phrase features of a certain layer in the phrase feature extraction layer.

[0066] In one embodiment of the present application, the semantic feature extraction layer includes L layers, and the phrase feature extraction layer includes K layers. Among them, the j-th layer included in the phrase feature extraction layer can be any one of the K layers included in the phrase feature extraction layer, and the i-th layer in the semantic feature extraction layer can be any one of the L layers.

[0067] Preferably, as Figure 3 shown, the i-th layer in the semantic feature extraction layer is the first layer of the K layers, and the j-th layer in the phrase feature extraction layer is the last layer of the L layers. That is to say, in some embodiments of the present application, the output of the last layer of the phrase feature extraction layer is used as the input of the first layer of the semantic feature extraction layer.

[0068] Therefore, by setting the i-th layer as the first layer of the K layers and the j-th layer as the last layer of the L layers in the embodiments of the present application, it is possible to obtain the phrase features with better learning degree in the phrase feature extraction layer and continue to learn in the semantic feature extraction layer, thereby more accurately expressing the input sentence.

[0069] The structure of the pre-trained language representation model in the embodiments of the present application is described in detail above. Next, the process of training the pre-trained language expression model in the embodiments of the present application will be described.

[0070] Step 1: Collect target training sentences corresponding to the target domain.

[0071] In an implementation manner of the present application, first, it is necessary to collect the original data related to the insurance field (i.e., the target domain), including important documents related to the insurance field, insurance product materials, insurance field Q&A, etc. Then, convert the obtained original data into text format and clean meaningless special characters, such as spaces, garbled codes, etc. Finally, divide 80% of the above data into the training set and the remaining 20% into the validation set.

[0072] Step 2: Input the target training sentences in the training set into the pre-trained language representation model for training to obtain the target language representation model.

[0073] First, the pre-trained language representation model divides the target training sentences into multiple characters at the input layer. One character corresponds to a second type of node, and a first type of node is placed before the multiple characters. The first type of node is used to represent the overall semantics of the target training sentence, denoted as [CLS]. A set of target training sentences may include two sentences. For example, sentence A is "Life insurance claim settlement practitioners with excellent qualities" and sentence B is "Insurance". The two sentences are separated by the separator [SEP]. Then, use the masking symbol [Mask] to mask M characters. Among them, during the training process, the masked characters will carry labels for supervised learning.

[0074] It should be noted that during the training process, a pair of sentences in the training set includes sentence A and sentence B. Sentence A and sentence B may carry labels, and the labels include that sentence A and sentence B are in an upper and lower sentence relationship, etc., or sentence A and sentence B are two unrelated sentences.

[0075] It should be noted that during the training process, due to the randomness of the masking symbol [Mask], when the number of target training sentences is small, a larger random masking number can be set to obtain more training samples.

[0076] As a specific embodiment of the present application, the main masking object of the masking symbol [Mask] during the masking process is phrases. Therefore, it is possible to predict the masked phrases during the training process, so that the training model can better capture the relationship between phrases. For example, the phrase "Insurance" is masked by the masking symbol [Mask].

[0077] Then, the pre-trained language representation model needs to complete two tasks during the training process, namely predicting the masked characters and predicting whether sentence B is the next sentence of sentence A. And by calculating the cross-entropy loss function for the prediction results, when it is confirmed that the prediction accuracy reaches the preset accuracy threshold, the training is ended to obtain the target language representation model.

[0078] As a specific embodiment of the present application, during the training of the pre-trained language representation model, the first type of nodes and the second type of nodes can be subjected to word embeddings (Token Embeddings), segment embeddings (Segment Embeddings), and position embeddings (Position Embeddings) to obtain the representation vectors of each word and the overall semantics. Among them, since the pre-trained language representation model also needs to have a classification function, segment embeddings (Segment Embeddings) are used to classify the two sentences in a pair of sentences (for example, sentence A and sentence B).

[0079] It should be noted that Token Embeddings is the representation of word vectors. Segment Embeddings is the vector representation for distinguishing the two sentences included in a pair of sentences. Position Embeddings is to learn the sequential attributes of the input.

[0080] It should be noted that during the process of representing the input target training sentence, it is also a process of encoding each word and the overall semantics, as shown in formula (1):

[0081] h 0 = Embedding(x) (1)

[0082] Among them, x represents the input target training sentence, and h 0 represents the feature vector obtained after encoding the target training sentence.

[0083] The following will be combined with Figure 4 to describe a specific embodiment of training the pre-trained language representation model in the embodiments of the present application.

[0084] As Figure 4 shown, the pre-trained language representation model includes an input layer 310, the first layer 321 in the phrase feature extraction layer, the last layer 322 in the phrase feature extraction layer, the first layer 331 in the syntactic feature extraction layer, the last layer 332 in the syntactic feature extraction layer, the first layer 341 in the semantic feature extraction layer, the last layer 342 in the semantic feature extraction layer, and an output layer 350.

[0085] Specifically, the initialization parameters of the pre-trained language representation model are from the parameters of the original Bert model. During the training process, the input of the first type of node [CLS] included in the first layer 341 of the semantic feature extraction layer comes from the first type of node [CLS] included in the last layer 332 of the syntactic feature extraction layer; the input of the second type of nodes other than the first type of node [CLS] included in the first layer 341 of the semantic feature extraction layer all comes from the last layer 322 of the phrase feature extraction layer. Moreover, there is no connection between the first layer 341 of the semantic feature extraction layer and the last layer 332 of the syntactic feature extraction layer, and the remaining layers are all connected in a fully connected manner.

[0086] For example, based on the original Bert model, a feature acquisition layer (i.e., the i-th layer) is added to the semantic feature extraction layer. Then, the pre-trained language representation model sequentially includes an input layer, the 1st to 6th hidden layers (i.e., the phrase feature extraction layer), the 7th to 10th hidden layers (i.e., the syntactic feature extraction layer), and the semantic feature extraction layer (i.e., the feature acquisition layer, the 11th and 12th hidden layers).

[0087] Step 1: The input target training sentence is x. A node [CLS] (i.e., the first type of node) representing the overall semantics is set before x, and then [CLS] and x are encoded together, as shown in formula (2):

[0088]

[0089] where cls represents the first type of node, x represents the second type of node, h 0 represents the feature vector obtained after encoding the target training sentence in the input layer, represents the representation vector of the overall semantics of the target training sentence in the input layer.

[0090] Step 2: Transmit h 0 and to the 1st hidden layer, and continue to encode h 0 and and then transmit it to the 2nd hidden layer for encoding, and so on, until it is transmitted to the 6th hidden layer to obtain the phrase features of each word and the overall semantics, as shown in formula (3):

[0091]

[0092] where h 0 represents the feature vector obtained after encoding the target training sentence in the input layer, represents the representation vector of the overall semantics of the target training sentence in the input layer, h 1-6 represents the phrase features of each word, Phrase features representing the overall semantics.

[0093] Step 3: Transmit h 1-6 and to the 7th hidden layer, and after encoding h 1-6 and and transmitting them to the 7th hidden layer for encoding, and so on, until transmitting to the 10th hidden layer to obtain the syntactic features of each word and the overall semantics, as shown in formula (4):

[0094]

[0095] where h 1-6 represents the phrase features of each word, represents the phrase features of the overall semantics, h 7-10 represents the syntactic features of each word, represents the syntactic features of the overall semantics.

[0096] Step 4: Transmit to the first type of nodes in the feature acquisition layer, and transmit h 1-6 to the second type of nodes in the feature acquisition layer, and continue to encode to obtain the primary semantic features of each word and the overall semantics, as shown in formula (5):

[0097]

[0098] where, represents the syntactic features of the overall semantics, h 1-6 represents the phrase features of each word, represents the primary semantic features of the overall semantics, h c represents the primary semantic features of each word.

[0099] Step 5: Input and h c into the 11th hidden layer for encoding, and then input to the 12th hidden layer to continue encoding, and finally output by the output layer to obtain the feature vectors corresponding to each word in the target training sentence and the feature vector of the overall semantics.

[0100] It should be noted that except for the feature acquisition layer, the states of other hidden layers in the pre-trained language representation model are determined by the states of the previous hidden layer, as shown in formula (6) and formula (7):

[0101] h l = Transformer l (h l-1 ) (6)

[0102]

[0103] Among them, h l represents the character features of the current layer, and h l-1 represents the character features of the previous layer. represents the overall semantic feature of the current layer, and represents the overall semantic feature of the previous layer.

[0104] The above describes the process of obtaining the target language representation model after training the pre-trained language representation model in the embodiments of the present application. The following will describe the method for searching for statements using the target language representation model in the embodiments of the present application.

[0105] It should be noted that the method for searching for statements in the embodiments of the present application can be executed by a server.

[0106] At least to solve the problems existing in the background technology, such as Figure 5 As shown, some embodiments of the present application provide a method for searching for statements, and the method includes:

[0107] S510, obtain the statement to be matched.

[0108] It should be noted that the statement to be matched needs to calculate the similarity with at least one candidate statement in the database one by one, and then select the candidate statement with a higher similarity to the statement to be matched. As a specific embodiment of the present application, the statement to be matched can be a question that the user wants to consult, for example, what types of insurance are there. As another specific embodiment of the present application, the statement to be matched can be a document that needs to find similar texts, for example, query a similar document of an insurance industry report, etc.

[0109] As an implementation manner of the present application, the user inputs a statement or a document (i.e., the statement to be matched) into the client, and the client sends the statement to be matched input by the user to the server. After the server obtains the statement to be matched, it executes S520, and then returns the obtained target statement to the client for display.

[0110] S520, input the statement to be matched into the target language representation model, and obtain a target statement that matches the statement to be matched through the target language representation model.

[0111] In an implementation manner of the present application, S520 includes:

[0112] Step 1: Extract the representation vector to be matched of the statement to be matched.

[0113] That is to say, after obtaining the statement to be matched, the statement to be matched is input into the above-mentioned trained target language representation model for calculation to obtain a to-be-matched representation vector corresponding to the statement to be matched. The to-be-matched representation vector includes a to-be-matched text semantic representation sub-vector and a to-be-matched character semantic representation sub-vector. The to-be-matched text semantic representation sub-vector is used to represent the overall semantics of the statement to be matched, and the to-be-matched character semantic representation sub-vector is used to represent the semantics of each character in the statement to be matched.

[0114] Step 2: Match the to-be-matched representation vector with at least one group of candidate representation vectors to obtain a target statement.

[0115] It should be noted that a group of candidate representation vectors is used to represent a candidate statement.

[0116] As an implementation manner in the above Step 2 of the present application, as Figure 6 shown, first, each candidate statement in at least one candidate statement needs to be input into the target language representation model for calculation to obtain at least one group of candidate representation vectors, that is, the Kth candidate statement is input into the target language representation model to obtain the Kth candidate representation vector. Among them, a group of candidate representation vectors includes a candidate text semantic representation sub-vector and a candidate character semantic representation sub-vector, and the above at least one group of candidate representation vectors is pre-stored in a database in advance, so as to reduce the calculation amount in the matching process and improve the matching speed.

[0117] Then, the statement to be matched is input into the target language representation model to obtain a to-be-matched representation vector.

[0118] Finally, calculate the target similarity value between the to-be-matched representation vector and the Kth candidate representation vector, so as to obtain the target similarity values between the to-be-matched representation vector and each candidate representation vector. Then, use the target similarity values to sort the candidate statements to obtain the target statement.

[0119] It should be noted that the to-be-matched representation vector is obtained by averaging the representation vectors corresponding to multiple characters. The method for calculating the target similarity value includes a cosine similarity calculation method or an Euclidean distance calculation method. The present application does not limit the method for calculating the similarity value.

[0120] As another implementation manner in the above Step 2 of the present application, first, calculate the target similarity value between the to-be-matched representation vector and each group of candidate representation vectors in at least one group of candidate representation vectors based on the weight value.

[0121] That is to say, different from the related art, in the embodiment of the present application, the weight between the character features and the text features extracted is adjusted by the weight value, where the text feature is the overall semantic feature that can represent the statement to be matched. One statement to be matched corresponds to one text feature, and one candidate statement corresponds to another text feature.

[0122] It should be noted that the target similarity value is used to characterize the similarity degree between the representation vector to be matched and each group of candidate representation vectors.

[0123] As a specific implementation manner of the present application, the specific process for calculating the target similarity value based on the above is as follows:

[0124] It should be noted that the candidate representation vector corresponding to the K-th candidate statement includes the K-th candidate text semantic representation sub-vector and the K-th candidate character semantic representation sub-vector. The weight value corresponding to the K-th candidate statement includes the K-th text weight value and the K-th character weight value, and the sum of the K-th text weight value and the K-th character weight value is 1. Wherein, the K-th candidate statement is any one of at least one candidate statement.

[0125] Step 1: Calculate the K-th text similarity value between the text semantic representation sub-vector of the text to be matched and the K-th candidate text semantic representation sub-vector, where K is an integer greater than or equal to 1.

[0126] That is to say, as Figure 7 shown, input the statement to be matched into the target language representation model for calculation to obtain the text semantic representation sub-vector of the statement to be matched (i.e., Vcls a), read the K-th candidate text semantic representation sub-vector (i.e., Vcls b) from the database, and then use the existing similarity calculation method to calculate the similarity between Vcls a and Vcls b to obtain the K-th text similarity value.

[0127] Step 2: Calculate the K-th character similarity value according to the character semantic representation sub-vector of the text to be matched and the K-th candidate character semantic representation sub-vector.

[0128] That is to say, after inputting the statement to be matched into the target language representation model for calculation, obtain the character semantic representation sub-vector of the statement to be matched. The character semantic representation sub-vector includes representation sub-vectors corresponding to multiple characters of the statement to be matched (i.e., VT1, VT2... VTN). At the same time, read the K-th candidate character semantic representation sub-vector corresponding to the K-th candidate statement from the database. The K-th candidate character semantic representation sub-vector includes representation sub-vectors corresponding to multiple characters of the K-th candidate statement (i.e., Vt1, Vt2... VtN).

[0129] After that, calculate the similarity between each character semantic representation sub-vector of the statement to be matched and each K-th candidate character semantic representation sub-vector respectively, and select the maximum value. As shown in formula (8):

[0130]

[0131] Among them, S tok (a, b) represents the K-th character similarity value, Ea Represents the word semantic representation sub-vector of the statement to be matched, E b The K-th candidate word semantic representation sub-vector, and [cls] represents the text semantic representation sub-vector.

[0132] As a specific embodiment of the present application, such as Figure 7 As shown, VT1 is respectively calculated for similarity with each of Vt1, Vt2... VtN, the maximum value is selected, and the maximum value is used as the maximum similarity value of T1. As another specific embodiment of the present application, VT2 is respectively calculated for similarity with each of Vt1, Vt2... VtN, the maximum value is selected, and the maximum value is used as the maximum similarity value of T2. By analogy according to the above method, until VtN is respectively calculated for similarity with each of Vt1, Vt2... VtN, the maximum value is selected, and the maximum value is used as the maximum similarity value of TN. Finally, the maximum similarity values of T1, T2... TN can be obtained. Finally, the maximum similarity values of T1, T2... TN are added to obtain the K-th word similarity value.

[0133] Step three: Calculate the product of the K-th text similarity value and the K-th text weight value to obtain the first product.

[0134] That is to say, after obtaining the K-th text similarity value using the method in step one, if the K-th text weight value is confirmed to be β, then β is multiplied by the K-th text similarity value to obtain the first product.

[0135] It should be noted that the value range of the weight value β is [0, 1].

[0136] Step four: Calculate the product of the K-th word similarity value and the K-th word weight value to obtain the second product.

[0137] That is to say, after obtaining the K-th word similarity value using the method in step two, the K-th word similarity value is multiplied by the K-th word weight value to obtain the second product. Since the sum of the K-th text weight value and the K-th word weight value is 1, the K-th text weight value is 1 - β.

[0138] It should be noted that in the actual application process, if the weight value of the text feature needs to be increased, then β is set to a value greater than 0.5. If the weight value of the word feature needs to be increased, then β is set to a value less than 0.5.

[0139] Step five: Calculate the sum of the first product and the second product to obtain the target similarity value corresponding to the K-th candidate statement.

[0140] That is, add the first product obtained in Step 3 to the second product obtained in Step 4 to obtain the target similarity value corresponding to the K-th candidate statement. Repeat the above steps to calculate the target similarity values between the statement to be matched and all candidate statements. As shown in Formula (9):

[0141]

[0142] where S full (a, b) represents the target similarity value, β represents the K-th text weight value, represents the semantic representation sub-vector of the text to be matched, represents the semantic representation sub-vector of the K-th candidate text, and S tok (a, b) represents the K-th character similarity value.

[0143] Therefore, in the embodiment of the present application, the target similarity value corresponding to the K-th candidate statement is calculated through the K-th text similarity value and the K-th character similarity value, which can assign the weights between the overall semantic features and each character feature, so that the calculated target similarity value can better conform to the actual situation of the target field, and thus the found statement is more accurate.

[0144] Then, obtain the target statement from the at least one candidate statement through the target similarity value.

[0145] That is, sort the target similarity values corresponding to the at least one candidate statement, and according to the sorting result, select one or more target statements from the at least one candidate statement, and the one or more target statements are fed back to the client for display.

[0146] Therefore, different from the related art, in the embodiment of the present application, the weights between the character features and text features extracted are adjusted through the weight value to obtain the target statement, which can arbitrarily adjust the weights between the character features and text features according to the actual situation, so as to recommend more accurate target statements.

[0147] Therefore, in the embodiment of the present application, by matching the statement representation vector to be matched with at least one group of candidate representation vectors, one or more target statements matching the statement to be matched can be found from the at least one candidate statement.

[0148] The above describes a method for finding a statement in the embodiment of the present application. The following will describe the application scenarios of finding a statement in the embodiment of the present application.

[0149] As one of the multiple scenarios of this application, the method for searching for statements is applied to a document retrieval system. Specifically, in this scenario, the type of the statement to be matched is a document. First, all candidate documents are input into the target language representation model for calculation to obtain candidate representation vectors corresponding to the candidate documents, which are stored in the database. Then, after the user inputs the statement to be matched online in the document retrieval system, the server loads the target language representation model and inputs the statement to be matched into the target language representation model for calculation to obtain the representation vector to be matched. Finally, using the above method for searching for statements, based on the candidate representation vector and the representation vector to be matched, the target similarity value is calculated, and the statements are sorted according to the target similarity value to obtain the final target statement.

[0150] As another one of the multiple scenarios of this application, the method for searching for statements is applied to re - sort the statements sorted by a search engine. Specifically, first, the parameter file of the target language representation model is stored in the server. After receiving the statement to be matched input by the user, pre - processing is performed on the statement to be matched. For example, meaningless special characters, spaces, garbled codes, etc. are deleted. Then, after the search engine (such as elasticseach) uses the BM25 algorithm to query the statement to be matched and obtains N statements with relatively high matching degrees, the server obtains the above N statements with relatively high matching degrees. Finally, the method for searching for statements in the embodiments of this application is used to re - sort the N statements with relatively high matching degrees to obtain the corresponding N target similarities, and the N statements with relatively high matching degrees are re - sorted according to the N target similarities, and the sorting result is returned to the client for display.

[0151] The above describes the application scenarios of a method for searching for statements in the embodiments of this application. The following will describe a training device for a language representation model in the embodiments of this application.

[0152] As Figure 8 shown, the device 800 for training a language representation model includes: a training statement acquisition module 810 and a model training module 820.

[0153] In one embodiment of the present application, an apparatus 800 for training a language representation model is provided in an embodiment of the present application. The training apparatus includes: a training statement acquisition module 810 configured to acquire a target training statement, where the target training statement is obtained by collecting statements in the target field to which the language representation model is applied; a model training module 820 configured to train a pre-trained language representation model according to the target training statement to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer, and the input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, and i and j are integers greater than or equal to 1.

[0154] In one embodiment of the present application, each layer in the pre-trained language representation model includes two types of nodes. Among them, the first type of nodes is used to extract text features, and the second type of nodes is used to extract character features; among them, the input of the second type of nodes included in the i-th layer is the output of the second type of nodes included in the j-th layer; the input of the first type of nodes included in the i-th layer is the output of the first type of nodes included in the i-1-th layer; among them, the text features are used to represent the overall semantic features of the target training statement, and the character features are used to represent the semantic features of a character in the target training statement.

[0155] In one embodiment of the present application, the semantic feature extraction layer includes L layers, and the phrase feature extraction layer includes K layers, where L and K are integers greater than 1. Among them, the i-th layer in the semantic feature extraction layer is the first layer in the K layers; the j-th layer in the phrase feature extraction layer is the last layer in the L layers.

[0156] In an embodiment of the present application Figure 8 The modules shown can implement Figure 2 、 Figure 3 and Figure 4 each process in the method embodiment. Figure 8 The operations and / or functions of each module in Figure 2 、 Figure 3 and Figure 4 correspond to the corresponding processes in the method embodiment in order to implement. For details, reference can be made to the description in the above method embodiment. To avoid repetition, the detailed description is appropriately omitted here.

[0157] The above describes a training apparatus for a language representation model in an embodiment of the present application. The following will describe an apparatus for searching statements in an embodiment of the present application.

[0158] As Figure 9 shown, an apparatus 900 for searching statements includes: a statement acquisition module 910 and a statement matching module 920.

[0159] In an implementation manner of the present application, an apparatus 900 for searching statements is provided in an embodiment of the present application. The apparatus includes: a statement acquisition module 910 configured to acquire a statement to be matched; a statement matching module 920 configured to input the statement to be matched into a target language representation model obtained by using the implementation manner of the first aspect, and obtain a target statement that matches the statement to be matched through the target language representation model.

[0160] In an implementation manner of the present application, the statement matching module 920 is further configured to: extract a to-be-matched feature vector of the statement to be matched; match the to-be-matched feature vector with at least one group of candidate feature vectors to obtain the target statement, where one group of candidate feature vectors is used to represent a candidate statement, and one group of candidate feature vectors corresponds to one candidate statement.

[0161] In an implementation manner of the present application, the statement matching module 920 is further configured to: calculate a target similarity value between the to-be-matched feature vector and each group of candidate feature vectors in the at least one group of candidate feature vectors based on a weight value, where the target similarity value is used to represent the similarity degree between the to-be-matched feature vector and each group of candidate feature vectors, and the weight value is used to adjust the weight between the extracted word features and the extracted text features; obtain the target statement from the at least one candidate statement through the target similarity value.

[0162] In an implementation manner of the present application, the to-be-matched feature vector includes a to-be-matched text semantic feature sub-vector and a to-be-matched word semantic feature sub-vector, the candidate feature vector corresponding to the Kth candidate statement includes a Kth candidate text semantic feature sub-vector and a Kth candidate word semantic feature sub-vector, the weight value corresponding to the Kth candidate statement includes a Kth text weight value and a Kth word weight value, and the sum of the Kth text weight value and the Kth word weight value is 1; the statement matching module 920 is further configured to: calculate a Kth text similarity value between the to-be-matched text semantic feature sub-vector and the Kth candidate text semantic feature sub-vector, where K is an integer greater than or equal to 1; calculate a Kth word similarity value based on the to-be-matched word semantic feature sub-vector and the Kth candidate word semantic feature sub-vector; calculate the product of the Kth text similarity value and the Kth text weight value to obtain a first product; calculate the product of the Kth word similarity value and the Kth word weight value to obtain a second product; calculate the sum of the first product and the second product to obtain a target similarity value corresponding to the Kth candidate statement.

[0163] In the embodiment of the present application, Figure 9 the modules shown can implement Figure 1 、 Figure 5 、Figure 6 and Figure 7 each process in the method embodiments. Figure 9 The operations and / or functions of each module in are respectively for implementing Figure 1 , Figure 5 , Figure 6 and Figure 7 the corresponding processes in the method embodiments in. For details, please refer to the descriptions in the above method embodiments. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0164] As Figure 10 shown, an electronic device 100 provided in an embodiment of the present application includes: a processor 101, a memory 102, and a bus 103. The processor is connected to the memory through the bus. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the method described in any one of the above all embodiments. For details, please refer to the descriptions in the above method embodiments. To avoid repetition, the detailed descriptions are appropriately omitted here.

[0165] Among them, the bus is used to realize the direct connection and communication of these components. Among them, in the embodiment of the present application, the processor may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0166] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the methods described in the above embodiments can be executed.

[0167] It can be understood that Figure 10 the structure shown is only schematic and may further include more or fewer components than those Figure 10 shown in, or have a configuration different from that Figure 10 shown in. Figure 10 Each component shown in can be implemented by hardware, software, or a combination thereof.

[0168] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a server, it implements the method described in any one of the above-mentioned embodiments. For details, reference can be made to the description in the above-mentioned method embodiments. To avoid repetition, the detailed description is appropriately omitted here.

[0169] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0170] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for training a language representation model, characterized in that, The method includes: Obtaining a target training statement, where the target training statement is obtained by collecting statements in the target domain to which the language representation model is applied; Training a pre-trained language representation model according to the target training statement to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer. The input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, where i and j are integers greater than or equal to 1; the pre-trained language representation model is a BERT model with an altered fully-connected structure; the semantic feature extraction layer includes L layers, and the phrase feature extraction layer includes K layers, where L and K are integers greater than 1. The i-th layer of the semantic feature extraction layer is the first layer of the K layers; the j-th layer of the phrase feature extraction layer is the last layer of the L layers; the first layer of the semantic feature extraction layer includes a first type of node and a second type of node; the second type of node in the semantic feature extraction layer has no connection with the last layer of the syntactic feature extraction layer.

2. The method according to claim 1, wherein Both types of nodes are included in each layer of the pre-trained language representation model, where the first type of node is used to extract text features, and the second type of node is used to extract character features; where The input of the second type of node included in the i-th layer is the output of the second type of node included in the j-th layer; The input of the first type of node included in the i-th layer is the output of the first type of node included in the i-1-th layer; Among them, the text features are used to represent the overall semantic features of the target training statement, and the character features are used to represent the semantic features of a single character in the target training statement.

3. A method for searching statements, characterized in that, The method includes: Obtaining a statement to be matched; Inputting the statement to be matched into the target language representation model obtained by any one of claims 1-2, and obtaining a target statement that matches the statement to be matched through the target language representation model.

4. The method according to claim 3, wherein The obtaining of the target statement that matches the statement to be matched through the target language representation model includes: Extracting a representation vector to be matched of the statement to be matched; Matching the representation vector to be matched with at least one group of candidate representation vectors to obtain the target statement, where one group of candidate representation vectors is used to represent a candidate statement, and one group of candidate representation vectors corresponds to one candidate statement.

5. The method according to claim 4, wherein The matching of the representation vector to be matched with at least one group of candidate representation vectors to obtain the target statement includes: Calculating a target similarity value between the representation vector to be matched and each group of candidate representation vectors in the at least one group of candidate representation vectors based on a weight value, where the target similarity value is used to represent the similarity degree between the representation vector to be matched and each group of candidate representation vectors, and the weight value is used to adjust the weight between the extracted character features and the extracted text features; Obtaining the target statement from at least one candidate statement through the target similarity value.

6. The method according to claim 5, wherein The to-be-matched feature vector includes a to-be-matched text semantic feature sub-vector and a to-be-matched word semantic feature sub-vector. The candidate feature vector corresponding to the K-th candidate statement includes a K-th candidate text semantic feature sub-vector and a K-th candidate word semantic feature sub-vector. The weight value corresponding to the K-th candidate statement includes a K-th text weight value and a K-th word weight value, and the sum of the K-th text weight value and the K-th word weight value is 1; Calculating the target similarity value between the to-be-matched feature vector and each group of candidate feature vectors in the at least one group of candidate feature vectors based on the weight value includes: Calculating a K-th text similarity value between the to-be-matched text semantic feature sub-vector and the K-th candidate text semantic feature sub-vector, where K is an integer greater than or equal to 1; Calculating a K-th word similarity value based on the to-be-matched word semantic feature sub-vector and the K-th candidate word semantic feature sub-vector; Calculating the product of the K-th text similarity value and the K-th text weight value to obtain a first product; Calculating the product of the K-th word similarity value and the K-th word weight value to obtain a second product; Calculating the sum of the first product and the second product to obtain the target similarity value corresponding to the K-th candidate statement.

7. An apparatus for training a language representation model, characterized in that, The device includes: A training statement acquisition module configured to acquire target training statements, where the target training statements are acquired by collecting statements in the target domain to which the language representation model is applied; A model training module configured to train a pre-trained language representation model according to the target training statements to obtain a target language representation model, where the pre-trained language representation model sequentially includes a phrase feature extraction layer, a syntactic feature extraction layer, and a semantic feature extraction layer. The input of some nodes in the i-th layer of the semantic feature extraction layer is the output of the j-th layer of the phrase feature extraction layer, and i and j are integers greater than or equal to 1; the pre-trained language representation model is a BERT model with a changed fully connected structure; the semantic feature extraction layer includes L layers, the phrase feature extraction layer includes K layers, L and K are integers greater than 1, the i-th layer of the semantic feature extraction layer is the first layer of the K layers; the j-th layer of the phrase feature extraction layer is the last layer of the L layers; the first layer of the semantic feature extraction layer includes a first type of node and a second type of node; the second type of node in the semantic feature extraction layer has no connection with the last layer of the syntactic feature extraction layer.

8. A device for searching statements, characterized in that, The device includes: A statement acquisition module configured to acquire a to-be-matched statement; A statement matching module configured to input the to-be-matched statement into the target language representation model obtained by using any one of claims 1-2, and obtain a target statement that matches the to-be-matched statement through the target language representation model.

9. An electronic device, characterized in that, Includes: A processor, a memory, and a bus; The processor is connected to the memory through the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Text emotion classification method and device, electronic device and readable storage medium

    CN110222178A

  • Text extraction method, text extraction system, electronic equipment and storage device

    CN113505218A