Text matching model training method, text query method, device and apparatus

By using a hybrid expert dual-tower model, the shared attention sub-model parameters enable query and document text features to reside in the same semantic space. The independent expert sub-models each possess their own characteristics, thus addressing the issue of poor training performance of the dual-tower model and improving the accuracy of text matching and recall.

CN116028824BActive Publication Date: 2026-02-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211703504.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-02-03
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

Existing dual-tower models suffer from poor training performance in text matching and text retrieval tasks because the features of query text and document text are not in the same semantic space or are not modeled for their respective characteristics.

Method used

A hybrid expert dual-tower model is adopted. By sharing parameters between the first and second attention sub-models, the attention features of the query text and the document text are in the same semantic space. At the same time, the first and second expert sub-models do not share parameters, so that they have their own characteristics. The model is optimized by adjusting the shared and independent gradients.

Benefits of technology

It improves the accuracy of text matching and text recall, thereby enhancing the accuracy of query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028824B_ABST
    Figure CN116028824B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text matching model training method, relating to the technical fields of artificial intelligence, natural language processing and intelligent recommendation. The specific implementation scheme is: determining the attention features of the first sample text and the second sample text respectively; processing the attention features of the first sample text using a first expert sub-model to obtain the semantic features of the first sample text; processing the attention features of the second sample text using a second expert sub-model to obtain the semantic features of the second sample text; calculating the loss of the text matching model according to the similarity between the semantic features of the first sample text and the semantic features of the second sample text; determining the shared attention gradient, the first independent gradient of the first expert sub-model and the second independent gradient of the second expert sub-model according to the loss; and adjusting the text matching model according to the shared attention gradient, the first independent gradient and the second independent gradient. The present disclosure also provides a text query method, device, equipment and medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of natural language processing, deep learning, and intelligent recommendation technology. More specifically, this disclosure provides a method for training a text matching model, a text query method, an apparatus, an electronic device, and a storage medium. Background Technology

[0002] Information retrieval (or information search, information query), as an important branch of natural language processing, includes tasks such as text matching and text recall. For example, given a search term "query," the goal of a text matching or text recall task is to select the most relevant documents from a candidate document library. Summary of the Invention

[0003] This disclosure provides a method for training a text matching model, a text query method, an apparatus, a device, and a storage medium.

[0004] According to the first aspect, a training method for a text matching model is provided, the text matching model including a first expert sub-model and a second expert sub-model; the method includes: determining the attention features of a first sample text and a second sample text respectively; processing the attention features of the first sample text using the first expert sub-model to obtain the semantic features of the first sample text; processing the attention features of the second sample text using the second expert sub-model to obtain the semantic features of the second sample text; calculating the loss of the text matching model based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text; determining the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model based on the loss; and adjusting the text matching model based on the shared attention gradient, the first independent gradient, and the second independent gradient.

[0005] According to the second aspect, a text query method is provided, which includes: obtaining query text; using a text matching model to calculate the similarity between the query text and candidate documents in a candidate document library; and determining, based on the similarity, candidate documents that match the query text from the candidate document library as query results; wherein the text matching model is trained according to the training method of the aforementioned text matching model.

[0006] According to a third aspect, a training apparatus for a text matching model is provided, the apparatus comprising: a first determining module for determining attention features of a first sample text and a second sample text respectively; a first processing module for processing the attention features of the first sample text using a first expert sub-model to obtain semantic features of the first sample text; a second processing module for processing the attention features of the second sample text using a second expert sub-model to obtain semantic features of the second sample text; a first calculation module for calculating the loss of the text matching model based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text; a second determining module for determining, based on the loss, a shared attention gradient of the text matching model, a first independent gradient of the first expert sub-model, and a second independent gradient of the second expert sub-model; and an adjustment module for adjusting the text matching model based on the shared attention gradient, the first independent gradient, and the second independent gradient.

[0007] According to the fourth aspect, a text query device is provided, the device comprising: an acquisition module for acquiring query text; a second calculation module for calculating the similarity between the query text and candidate documents in a candidate document library using a text matching model; and a third determination module for determining, based on the similarity, candidate documents that match the query text from the candidate document library as query results; wherein the text matching model is trained using the aforementioned text matching model training device.

[0008] According to a fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to the present disclosure.

[0009] According to a sixth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods provided in this disclosure.

[0010] According to a seventh aspect, a computer program product is provided, comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method provided in this disclosure when executed by a processor.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0013] Figure 1 This is a schematic diagram of a dual-tower model in related technologies;

[0014] Figure 2 This is a flowchart of a method for training a text matching model according to an embodiment of the present disclosure;

[0015] Figure 3A This is a schematic diagram of a twin-tower model according to an embodiment of the present disclosure;

[0016] Figure 3B This is a schematic diagram of a twin-tower model according to an embodiment of the present disclosure;

[0017] Figure 4 This is a flowchart of a text query method according to an embodiment of the present disclosure;

[0018] Figure 5 This is a block diagram of a training apparatus for a text matching model according to an embodiment of the present disclosure;

[0019] Figure 6 This is a block diagram of a text query device according to an embodiment of the present disclosure;

[0020] Figure 7 This is a block diagram of an electronic device for training a text matching model and / or a text query method according to an embodiment of the present disclosure. Detailed Implementation

[0021] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0022] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0023] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0024] In information retrieval scenarios, tasks such as text matching and text recall can be achieved using a dual-tower model.

[0025] Figure 1 This is a schematic diagram of a dual-tower model in related technologies.

[0026] like Figure 1 As shown, the dual-tower model may include a first model 110 for query text and a second model 120 for document text. Both the first model 110 and the second model 120 may be deep learning models for text feature processing and text semantic feature extraction. The query text may include query terms or search terms, and the document text may include the titles and / or content of candidate documents in the candidate document library.

[0027] For example, the first model 110 may include multiple processing layers (e.g., a transformer processing layer). The query text can be encoded first to obtain the initial features of the query text. The initial features of the query text are then input into the first model 110. After processing by the multiple processing layers of the first model 110, the semantic features 111 of the query text can be obtained.

[0028] Similarly, the second model 120 may also include multiple processing layers (e.g., transformer processing layers). The doc text can be encoded first to obtain the initial features of the doc text. The initial features of the doc text can be input into the second model 120. After processing by multiple processing layers of the second model 120, the semantic features 121 of the doc text can be obtained.

[0029] Next, a similarity assessment value S can be calculated between semantic feature 111 and semantic feature 121. This similarity assessment value reflects the degree of matching between the query text and the document text. The retrieval results (query results), text matching results, and text recall results can be determined based on the similarity assessment value S. For example, candidate documents in the candidate document library with a similarity greater than a threshold (e.g., 90%) can be used as the user's retrieval results.

[0030] In text search, text matching, and text retrieval tasks, the inconsistency between query text and document text is a significant challenge. For example, a query is typically a few to a dozen words long, while the content of a candidate document usually contains hundreds of words. Therefore, designing a dual-tower model to make the feature representations of query text and document text more accurate is an important issue.

[0031] A dual-tower model design approach is proposed, in which the first model 110 and the second model 120 of the dual-tower model completely share parameters, which allows the features of query text and doc text to naturally reside in the same semantic space, but does not model the individual characteristics of query text and doc text.

[0032] A dual-tower model design approach is proposed, in which the first model 110 and the second model 120 of the dual-tower model do not share any parameters, so that the features of the query text and the features of the doc text have their own characteristics. However, this may result in the features of the query text and the features of the doc text not being in the same semantic space.

[0033] The aforementioned dual-tower model, which completely excludes parameter sharing, results in the query text and document text having features outside the same semantic space. Conversely, the dual-tower model that completely shares parameters fails to model the unique characteristics of both query and document texts. Therefore, both completely excluding and completely sharing parameters in the dual-tower model have drawbacks, leading to poor model training performance.

[0034] Figure 2 This is a flowchart of a training method for a text matching model according to an embodiment of the present disclosure.

[0035] like Figure 2 As shown, the training method 200 of the text matching model includes operations S210 to S260.

[0036] For example, a text matching model can be a dual-tower model based on hybrid experts. This model can include a first hybrid expert model for query text and a second hybrid expert model for document text. Both the first and second hybrid expert models can be deep learning models used for text feature processing and semantic feature extraction. The first hybrid expert model can include a first attention sub-model and a first expert sub-model. The first attention sub-model can be a Self-Attention neural network, and the first expert sub-model can be an FFN (Feed-Forward Network). The second hybrid expert model can include a second attention sub-model and a second expert sub-model. Similarly, the second attention sub-model can also be a Self-Attention neural network, and the second expert sub-model can also be an FFN (Feed-Forward Network).

[0037] In operation S210, the attention features of the first sample text and the second sample text are determined respectively.

[0038] For example, the first sample text can be a query text, and the second sample text can be a document text. Encoding the first sample text yields the initial features of the query text. Inputting these initial features into the first attention sub-model yields the attention features of the query text.

[0039] Similarly, encoding the second sample text yields the initial features of the doc text. Inputting these initial features into the second attention sub-model yields the attention features of the doc text.

[0040] In operation S220, the attention features of the first sample text are processed using the first expert sub-model to obtain the semantic features of the first sample text.

[0041] In operation S230, the attention features of the second sample text are processed using the second expert sub-model to obtain the semantic features of the second sample text.

[0042] For example, the attention features of the query text are input into the first expert sub-model FFN to obtain the semantic features of the query text. Similarly, the attention features of the document text are input into the second expert sub-model FFN to obtain the semantic features of the document text.

[0043] Operations S220 and S230 can be performed in parallel.

[0044] In operation S240, the loss of the text matching model is calculated based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text.

[0045] For example, the similarity between the semantic features of the query text and the semantic features of the document text can be calculated as the model's predicted matching degree between the query text and the document text. Based on this predicted matching degree and the differences between the labeled matching degree tags (e.g., cross-entropy loss, mean squared error loss, etc.), the loss of the text matching model can be calculated.

[0046] In operation S250, based on the loss, the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model are determined.

[0047] For example, the first attention sub-model and the second attention sub-model of a text matching model can share parameters, while the first expert sub-model and the second expert sub-model can not share parameters.

[0048] For example, based on the loss of the text matching model and the sum of the input features of the first attention sub-model and the second attention sub-model, the shared parameters of the first attention sub-model and the second attention sub-model can be differentiated to obtain the shared attention gradient of the first attention sub-model and the second attention sub-model.

[0049] For example, based on the loss of the text matching model and the input features of the first and second expert sub-models, the independent parameters of the first and second expert sub-models can be differentiated to obtain the first independent gradient and the second independent gradient of the first expert sub-model.

[0050] In operating S260, the text matching model is adjusted based on the shared attention gradient, the first independent gradient, and the second independent gradient.

[0051] For example, the shared attention gradient can be backpropagated to the first attention sub-model and the second attention sub-model, so that the shared parameters of the first attention sub-model and the second attention sub-model can be updated.

[0052] It should be noted that the parameters of the updated first attention sub-model are the same as those of the updated second attention sub-model. That is, the first and second attention sub-models share parameters, which makes the attention features of the query text and the attention features of the doc text in the same semantic space, resulting in more accurate attention feature representation.

[0053] For example, the first independent gradient of the first expert sub-model can be backpropagated to the first expert sub-model, thereby updating its parameters. Similarly, the second independent parameters of the second expert sub-model can be backpropagated to the second expert sub-model, thereby updating its parameters.

[0054] It should be noted that the parameters of the updated first expert sub-model and the updated second expert sub-model are generally different. That is, the first expert sub-model and the second expert sub-model do not share parameters, so that the semantic features of the query text and the semantic features of the doc text have their own characteristics.

[0055] Compared to related technologies where dual-tower models either don't share parameters during training (the features of the query text and the document text are not in the same semantic space) or share parameters completely (no modeling is done for the individual characteristics of the query text and the document text), resulting in poor model training performance, the dual-tower model based on hybrid experts in this disclosure shares parameters between the first attention sub-model and the second attention sub-model, ensuring that the attention features of the query text and the document text are in the same semantic space, leading to more accurate attention feature representation. Furthermore, the first expert sub-model and the second expert sub-model do not share parameters, allowing the semantic features of the query text and the document text to have their own distinct characteristics. Therefore, this improves the training performance of the dual-tower model, thereby enhancing the accuracy of text matching, text recall, and text querying.

[0056] The following is combined Figures 3A-3BThis further illustrates the difference between the dual-tower model with fully shared parameters and the dual-tower model provided in this disclosure, in which some sub-models share parameters and some sub-models do not share parameters.

[0057] Figures 3A-3B This is a schematic diagram of a twin-tower model according to an embodiment of the present disclosure.

[0058] like Figures 3A-3B As shown, the dual-tower model includes a first attention sub-model 311, a second attention sub-model 312, a first normalization sub-model 321, a second normalization sub-model 322, a first expert sub-model 331, a second expert sub-model 332, a first normalization sub-model 341, and a second normalization sub-model 342. The first attention sub-model 311, the first normalization sub-model 321, the first expert sub-model 331, and the first normalization sub-model 341 constitute the first hybrid expert model for query text. The second attention sub-model 312, the second normalization sub-model 322, the second expert sub-model 332, and the second normalization sub-model 342 constitute the second hybrid expert model for doc text.

[0059] Figure 3A The dual-tower model shown shares parameters completely, that is, the first attention sub-model 311 and the second attention sub-model 312 share parameters, the first normalization sub-model 321 and the second normalization sub-model 322 share parameters, the first expert sub-model 331 and the second expert sub-model 332 share parameters, and the first normalization sub-model 341 and the second normalization sub-model 342 share parameters.

[0060] Figure 3B In the dual-tower model shown, the first expert sub-model 331 and the second expert sub-model 332 do not share parameters, while the remaining sub-models share parameters. Specifically, the first attention sub-model 311 and the second attention sub-model 312 share parameters, the first normalization sub-model 321 and the second normalization sub-model 322 share parameters, and the first normalization sub-model 341 and the second normalization sub-model 342 share parameters.

[0061] like Figure 3AAs shown, the initial feature q0 of the query text is input into the first attention sub-model 311 to obtain attention feature q1. The sum of the initial feature q0 and attention feature q1 (q0+q1) is input into the first normalization sub-model 321 to obtain normalized feature q2. Normalized feature q2 is input into the first expert sub-model 331 to obtain semantic feature q3. The sum of the normalized feature q2 and semantic feature q3 (q2+q3) is input into the first normalization sub-model 341 to obtain normalized feature q4. Similarly, the initial feature d0 of the document text is input into the second attention sub-model 312 to obtain attention feature d1. The sum of the initial feature d0 and attention feature d1 (d0+d1) is input into the second normalization sub-model 322 to obtain normalized feature d2. Normalized feature d2 is input into the second expert sub-model 332 to obtain semantic feature d3. The sum of normalized feature d2 and semantic feature d3 of the doc text (d2+d3) is input into the second normalized sub-model 342 to obtain normalized feature d4.

[0062] Next, the normalized features q4 of the query text and d4 of the doc text can be used to calculate a similarity evaluation value. Based on the similarity evaluation value, the loss of the dual-tower model is calculated. Based on the loss, the shared attention gradients of the first attention sub-model 311 and the second attention sub-model 312, the shared normalized gradients of the first normalized sub-model 321 and the second normalized sub-model 322, the shared expert gradients of the first expert sub-model 331 and the second expert sub-model 332, and the shared normalized gradients of the first normalized sub-model 341 and the second normalized sub-model 342 can be calculated. Then, the shared attention parameters of the first attention sub-model 311 and the second attention sub-model 312, the shared normalized parameters of the first normalized sub-model 321 and the second normalized sub-model 322, the shared expert parameters of the first expert sub-model 331 and the second expert sub-model 332, and the shared normalized parameters of the first normalized sub-model 341 and the second normalized sub-model 342 can be updated.

[0063] like Figure 3BAs shown, the initial feature q0' of the query text is input into the first attention sub-model 311 to obtain the attention feature q1'. The sum of the initial feature q0' and the attention feature q1' (q0'+q1') is input into the first normalization sub-model 321 to obtain the normalized feature q2'. The normalized feature q2' is input into the first expert sub-model 331 to obtain the semantic feature q3'. The sum of the normalized feature q2' and the semantic feature q3' (q2'+q3') is input into the first normalization sub-model 341 to obtain the normalized feature q4'. Similarly, the initial feature d0' of the document text is input into the second attention sub-model 312 to obtain the attention feature d1'. The sum of the initial feature d0' and the attention feature d1' (d0'+d1') is input into the second normalization sub-model 322 to obtain the normalized feature d2'. Normalized feature d2' is input into the second expert sub-model 332 to obtain semantic feature d3'. The sum of normalized feature d2' and semantic feature d3' of the doc text (d2'+d3') is input into the second normalization sub-model 342 to obtain normalized feature d4'.

[0064] Next, the normalized features q4' of the query text and d4' of the doc text can be used to calculate similarity evaluation values. Based on these similarity evaluation values, the loss of the dual-tower model is calculated. Based on the loss, the shared attention gradient of the first attention sub-model 311 and the second attention sub-model 312, the shared normalized gradient of the first normalized sub-model 321 and the second normalized sub-model 322, the independent gradients of the first expert sub-model 331 and the second expert sub-model 332, and the shared normalized gradient of the first normalized sub-model 341 and the second normalized sub-model 342 are calculated. Subsequently, the shared attention parameters of the first attention sub-model 311 and the second attention sub-model 312, the shared normalized parameters of the first normalized sub-model 321 and the second normalized sub-model 322, the independent expert parameters of the first expert sub-model 331, the independent expert parameters of the second expert sub-model 332, and the shared normalized parameters of the first normalized sub-model 341 and the second normalized sub-model 342 can be updated.

[0065] According to embodiments of this disclosure, the dual-tower model includes N processing layers, each processing layer including a first attention sub-model, a first expert sub-model, a second attention sub-model, and a second expert sub-model, where N is an integer greater than 1. For the i-th processing layer, the semantic features of the first sample text output by the first expert sub-model of the i-th processing layer are input into the first attention sub-model of the (i+1)-th processing layer, and the semantic features of the second sample text output by the second expert sub-model of the i-th processing layer are input into the second attention sub-model of the (i+1)-th processing layer, i = 1, ..., N-1; for the (i+1)-th processing layer, the steps of determining the attention features and semantic features of the first and second sample texts respectively are performed until the semantic features of the first sample text output by the first expert sub-model of the N-th processing layer and the semantic features of the second sample text output by the second expert sub-model of the N-th processing layer are obtained.

[0066] For example, the dual-tower model provided in this disclosure includes multiple processing layers (e.g., 6), and the specific structure of each processing layer is as follows: Figure 3B The structure is shown. The normalized feature q4' of the query text output by the first normalized sub-model 341 of the first processing layer is input to the first attention sub-model 311 of the second processing layer. The normalized feature d4' of the doc text output by the second normalized sub-model 342 of the first processing layer is input to the second attention sub-model 312 of the second processing layer, and so on, until the normalized feature q4' of the query text output by the first normalized sub-model 341 and the normalized feature d4' of the doc text output by the second normalized sub-model 342 of the last processing layer are obtained.

[0067] For example, the loss of the text matching model is calculated based on the normalized features q4' of the query text output by the first normalized sub-model 341 of the last processing layer and the normalized features d4' of the doc text output by the second normalized sub-model 342.

[0068] According to embodiments of this disclosure, the dual-tower model includes N processing layers, and the first expert sub-model and the second expert sub-model of each of the N processing layers do not share parameters.

[0069] For example, the dual-tower model includes six processing layers, each of which is as follows: Figure 3B The structure is shown. In this example, the independent gradients of the first expert sub-model 331 and the second expert sub-model 332 in each processing layer are determined according to the loss of the dual-tower model, and then the independent expert parameters of the first expert sub-model 331 and the second expert sub-model 332 are updated.

[0070] According to embodiments of this disclosure, the dual-tower model includes N processing layers. The first expert sub-model and the second expert sub-model of a designated processing layer do not share parameters, while the first expert sub-model and the second expert sub-model of processing layers other than the designated processing layer share parameters.

[0071] For example, the dual-tower model includes six processing layers, with the first and second expert sub-models in each of the six processing layers not sharing parameters. Exemplarily, the first and second expert sub-models in the first, third, and fifth processing layers do not share parameters; that is, the structures of the first, third, and fifth processing layers are as follows: Figure 3B The structure is shown. The first and second expert sub-models of the second, fourth, and sixth processing layers share parameters; that is, the structures of the second, fourth, and sixth processing layers are as follows: Figure 3A The structure shown.

[0072] Alternatively, in the dual-tower model, the first and second expert sub-models of the first, third, and fifth processing layers share parameters; that is, the structures of the first, third, and fifth processing layers are as follows: Figure 3A The structure is shown. The first and second expert sub-models of the second, fourth, and sixth processing layers do not share parameters; that is, the structures of the second, fourth, and sixth processing layers are as follows: Figure 3B The structure shown.

[0073] For example, the dual-tower model includes 6 processing layers. The first expert sub-model and the second expert sub-model of the first K1 processing layers (e.g., K1=3, i.e., the first 3 processing layers) do not share parameters. That is, the structure of the first 3 processing layers is as follows: Figure 3B The structure is shown. The first and second expert sub-models of the last K2 processing layers (e.g., K2=3, i.e., the last 3 processing layers) share parameters; that is, the structure of the last 3 processing layers is as follows: Figure 3A The structure shown.

[0074] Alternatively, the first expert submodel and the second expert submodel of the first three processing layers in the six processing layers share parameters, that is, the structure of the first three processing layers is as follows: Figure 3A The structure is shown. The first and second expert sub-models of the last three processing layers do not share parameters; that is, the structure of the last three processing layers is as follows: Figure 3B The structure shown.

[0075] The embodiments disclosed herein can specify that the first expert sub-model and the second expert sub-model of any part or all of the processing layers in the dual-tower model do not share parameters, thereby improving the flexibility of dual-tower model design.

[0076] Figure 4 This is a flowchart of a text query method according to an embodiment of the present disclosure.

[0077] like Figure 4 As shown, the text query method 400 includes operations S410 to S430.

[0078] In operation S410, retrieve the query text.

[0079] In operation S420, a text matching model is used to calculate the similarity between the query text and candidate documents in the candidate document library.

[0080] In operation S430, candidate documents that match the query text are determined from the candidate document library based on similarity, and these are used as query results.

[0081] For example, the text matching model is trained using the training method described above. The query text includes the query terms, referred to as the query text, and the titles or content of candidate documents in the candidate document library, referred to as the doc text. By inputting the query text and doc text into the trained text matching model, a similarity evaluation value between the query text and the doc text can be obtained.

[0082] For example, candidate texts with a similarity score greater than a threshold (e.g., 90%) can be used as query results.

[0083] The embodiments of this disclosure use a trained hybrid expert-based text matching model to calculate the similarity between the query text and candidate documents, which can improve the accuracy of text matching and thus improve the accuracy of query results.

[0084] Figure 5 This is a block diagram of a training apparatus for a text matching model according to an embodiment of the present disclosure. The text matching model includes a first expert sub-model and a second expert sub-model.

[0085] like Figure 5 As shown, the training device 500 for the text matching model includes a first determining module 501, a first processing module 502, a second processing module 503, a first calculation module 504, a second determining module 505, and an adjustment module 506.

[0086] The first determining module 501 is used to determine the attention features of the first sample text and the second sample text respectively.

[0087] The first processing module 502 is used to process the attention features of the first sample text using the first expert sub-model to obtain the semantic features of the first sample text.

[0088] The second processing module 503 is used to process the attention features of the second sample text using the second expert sub-model to obtain the semantic features of the second sample text.

[0089] The first calculation module 504 is used to calculate the loss of the text matching model based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text.

[0090] The second determining module 505 is used to determine the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model based on the loss.

[0091] The adjustment module 506 is used to adjust the text matching model based on the shared attention gradient, the first independent gradient, and the second independent gradient.

[0092] According to embodiments of this disclosure, the text matching model further includes a first attention sub-model and a second attention sub-model.

[0093] The first determining module 501 includes a first processing unit and a second processing unit.

[0094] The first processing unit is used to process the first sample text using the first attention sub-model to obtain the attention features of the first sample text.

[0095] The second processing unit is used to process the second sample text using the second attention sub-model to obtain the attention features of the second sample text.

[0096] The adjustment module 506 includes a first update unit, a second update unit, and a third update unit.

[0097] The first update unit is used to update the shared parameters of the first attention sub-model and the second attention sub-model based on the shared attention gradient.

[0098] The second update unit is used to update the independent parameters of the first expert sub-model based on the first independent gradient.

[0099] The third update unit is used to update the independent parameters of the second expert sub-model based on the second independent gradient.

[0100] The text matching model comprises N processing layers, each including a first attention sub-model, a first expert sub-model, a second attention sub-model, and a second expert sub-model, where N is an integer greater than 1. The training device 500 for the text matching model also includes an input module and a third processing module.

[0101] The input module is used to input the semantic features of the first sample text output by the first expert sub-model of the i-th processing layer into the first attention sub-model of the (i+1)-th processing layer, and to input the semantic features of the second sample text output by the second expert sub-model of the i-th processing layer into the second attention sub-model of the (i+1)-th processing layer, i = 1, ..., N-1.

[0102] The third processing module is used to perform the steps of determining the attention features and semantic features of the first sample text and the second sample text respectively for the (i+1)th processing layer, until the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer are obtained.

[0103] The first calculation module 504 is used to calculate the loss of the text matching model based on the similarity between the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer.

[0104] The second determining module 505 is used to determine, based on the loss, the first independent gradient of the first expert submodel and the second independent gradient of the second expert submodel in each of the N processing layers.

[0105] The second determining module 505 includes a first determining unit and a second determining unit.

[0106] The first determining unit is used to determine, based on the loss, the first independent gradient of the first expert submodel of the specified processing layer and the second independent gradient of the second expert submodel in the N processing layers.

[0107] The second determining unit is used to determine the shared expert gradient of the first expert submodel and the second expert submodel of the N processing layers, excluding the specified processing layer, based on the loss.

[0108] According to embodiments of this disclosure, the specified processing layer includes one of an odd number of layers, an even number of layers, the first K1 layers, or the last K2 layers, wherein K1 and K2 are both integers greater than or equal to 1 and less than N.

[0109] Figure 6 This is a block diagram of a text query device according to an embodiment of the present disclosure.

[0110] like Figure 6 As shown, the text query device 600 includes an acquisition module 601, a second calculation module 602, and a third determination module 603.

[0111] The acquisition module 601 is used to acquire the query text.

[0112] The second calculation module 602 is used to calculate the similarity between the query text and candidate documents in the candidate document library using a text matching model.

[0113] The third determining module 603 is used to determine candidate documents that match the query text from the candidate document library based on similarity, and use them as query results.

[0114] The text matching model is trained using the training device described above.

[0115] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0118] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0119] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as methods for training text matching models and / or text querying methods. For example, in some embodiments, the methods for training text matching models and / or text querying methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the methods for training text matching models and / or text querying methods described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform a training method for the text matching model and / or a text query method.

[0120] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0124] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0125] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0126] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0127] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for training a text matching model, wherein the text matching model includes a first expert sub-model and a second expert sub-model; the method includes: Determine the attention features of the first and second sample texts respectively; The attention features of the first sample text are processed using the first expert sub-model to obtain the semantic features of the first sample text. The attention features of the second sample text are processed using a second expert sub-model to obtain the semantic features of the second sample text; The loss of the text matching model is calculated based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text. Based on the loss, determine the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model; and The text matching model is adjusted based on the shared attention gradient, the first independent gradient, and the second independent gradient. The text matching model further includes a first attention sub-model and a second attention sub-model; adjusting the text matching model based on the shared attention gradient, the first independent gradient, and the second independent gradient includes: Update the shared parameters of the first attention sub-model and the second attention sub-model based on the shared attention gradient; Update the independent parameters of the first expert sub-model based on the first independent gradient; Update the independent parameters of the second expert sub-model based on the second independent gradient.

2. The method according to claim 1, wherein, The determination of the attention features of the first sample text and the second sample text includes: The first sample text is processed using the first attention sub-model to obtain the attention features of the first sample text; and The second attention sub-model is used to process the second sample text to obtain the attention features of the second sample text.

3. The method according to claim 2, wherein, The text matching model includes N processing layers, each processing layer including the first attention sub-model, the first expert sub-model, the second attention sub-model, and the second expert sub-model, where N is an integer greater than 1; the method further includes: The semantic features of the first sample text output by the first expert sub-model of the i-th processing layer are input into the first attention sub-model of the (i+1)-th processing layer, and the semantic features of the second sample text output by the second expert sub-model of the i-th processing layer are input into the second attention sub-model of the (i+1)-th processing layer, i=1,……N-1; For the (i+1)th processing layer, the steps of determining the attention features and semantic features of the first sample text and the second sample text are performed until the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer are obtained.

4. The method according to claim 3, wherein, The step of calculating the loss of the text matching model based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text includes: The loss of the text matching model is calculated based on the similarity between the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer.

5. The method according to claim 3, wherein, The step of determining the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model based on the loss includes: Based on the loss, determine the first independent gradient of the first expert submodel and the second independent gradient of the second expert submodel in each of the N processing layers.

6. The method according to claim 3, wherein, The step of determining the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model based on the loss includes: Based on the loss, determine the first independent gradient of the first expert submodel and the second independent gradient of the second expert submodel in the specified processing layer of the N processing layers; and Based on the loss, determine the shared expert gradient of the first expert submodel and the second expert submodel of the processing layers other than the specified processing layer in the N processing layers.

7. The method according to claim 6, wherein, The specified processing layer includes one of the following: an odd layer, an even layer, the first K1 layers, or the last K2 layers, wherein K1 and K2 are both integers greater than or equal to 1 and less than N.

8. A text query method, comprising: Get the query text; The similarity between the query text and candidate documents in the candidate document library is calculated using a text matching model. as well as Based on the similarity, candidate documents that match the query text are determined from the candidate document library and used as the query results; The text matching model is trained using the method described in any one of claims 1 to 7.

9. A training apparatus for a text matching model, the text matching model comprising a first expert sub-model and a second expert sub-model; the apparatus comprising: The first determining module is used to determine the attention features of the first sample text and the second sample text respectively; The first processing module is used to process the attention features of the first sample text using a first expert sub-model to obtain the semantic features of the first sample text. The second processing module is used to process the attention features of the second sample text using the second expert sub-model to obtain the semantic features of the second sample text. The first calculation module is used to calculate the loss of the text matching model based on the similarity between the semantic features of the first sample text and the semantic features of the second sample text. The second determining module is configured to determine, based on the loss, the shared attention gradient of the text matching model, the first independent gradient of the first expert sub-model, and the second independent gradient of the second expert sub-model; and An adjustment module is used to adjust the text matching model based on the shared attention gradient, the first independent gradient, and the second independent gradient. The text matching model further includes a first attention sub-model and a second attention sub-model. The adjustment module includes: The first update unit is used to update the shared parameters of the first attention sub-model and the second attention sub-model according to the shared attention gradient. The second update unit is used to update the independent parameters of the first expert sub-model based on the first independent gradient. The third update unit is used to update the independent parameters of the second expert sub-model based on the second independent gradient.

10. The apparatus according to claim 9, wherein, The first determining module includes: The first processing unit is configured to process the first sample text using the first attention sub-model to obtain the attention features of the first sample text; and The second processing unit is used to process the second sample text using the second attention sub-model to obtain the attention features of the second sample text.

11. The apparatus according to claim 10, wherein, The text matching model includes N processing layers, each processing layer including the first attention sub-model, the first expert sub-model, the second attention sub-model, and the second expert sub-model, where N is an integer greater than 1; the device further includes: The input module is used to input the semantic features of the first sample text output by the first expert sub-model of the i-th processing layer into the first attention sub-model of the (i+1)-th processing layer, and to input the semantic features of the second sample text output by the second expert sub-model of the i-th processing layer into the second attention sub-model of the (i+1)-th processing layer, i=1,……N-1; The third processing module is used to perform the steps of determining the attention features and semantic features of the first sample text and the second sample text respectively for the (i+1)th processing layer, until the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer are obtained.

12. The apparatus according to claim 11, wherein, The first calculation module is used to calculate the loss of the text matching model based on the similarity between the semantic features of the first sample text output by the first expert sub-model of the Nth processing layer and the semantic features of the second sample text output by the second expert sub-model of the Nth processing layer.

13. The apparatus according to claim 11, wherein, The second determining module is used to determine, based on the loss, the first independent gradient of the first expert submodel and the second independent gradient of the second expert submodel in each of the N processing layers.

14. The apparatus according to claim 11, wherein, The second determining module includes: The first determining unit is configured to determine, based on the loss, the first independent gradient of the first expert submodel and the second independent gradient of the second expert submodel in a specified processing layer among the N processing layers; and The second determining unit is used to determine, based on the loss, the shared expert gradient of the first expert sub-model and the second expert sub-model of the processing layers other than the specified processing layer among the N processing layers.

15. The apparatus according to claim 14, wherein, The specified processing layer includes one of the following: an odd layer, an even layer, the first K1 layers, or the last K2 layers, wherein K1 and K2 are both integers greater than or equal to 1 and less than N.

16. A text query device, comprising: The retrieval module is used to retrieve the query text; The second calculation module is used to calculate the similarity between the query text and candidate documents in the candidate document library using a text matching model; as well as The third determining module is used to determine, based on the similarity, candidate documents that match the query text from the candidate document library, as the query result; The text matching model is trained using the apparatus according to any one of claims 9 to 15.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.

19. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Text matching method and device, computer equipment and storage medium

    CN111797204A

  • Text matching method and device based on machine learning, equipment and storage medium

    CN113672701A