Text Processing Method, Device, Electronic Device and Storage Medium

By considering word segmentation categories and association relationships in the text matching process, the similarity matrix calculation is enhanced, and the problem of low text matching accuracy is solved, and the efficiency and accuracy of resource query are improved.

CN113821588BActive Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110614403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-02
Publication Date
2025-07-18
Estimated Expiration
2041-06-02

AI Technical Summary

Technical Problem

In the prior art, the accuracy of text matching is low, resulting in inefficient resource query.

Method used

By determining the matching weight matrix between the query text and the text to be matched, combining the word segmentation category and association relationship, enhancing the similarity matrix calculation, and comprehensively considering the word segmentation category and association relationship to calculate the matching score.

Benefits of technology

It improves the accuracy of text matching, ensures that the matching score more accurately reflects the actual matching between texts, and improves the efficiency of resource query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113821588B_ABST
    Figure CN113821588B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and discloses a text processing method, apparatus, electronic device and storage medium. The method includes: determining a matching weight matrix of a query text relative to a text to be matched, where the matching weight matrix includes at least one of a first weight matrix and a second weight matrix; enhancing a similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix, and the similarity matrix of the query text relative to the text to be matched is obtained by calculating the similarity of the word vectors of each word segment in the query text and the word vectors of each word segment in the text to be matched; determining a matching degree score between the query text and the text to be matched according to the first similarity matrix; and determining a target matching text according to the matching degree score between the query text and the text to be matched. Through this solution, the accuracy of text matching can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a text processing method, device, electronic device and storage medium. Background Art

[0002] Text matching is widely used in resource query scenarios, such as news information query, paper query, etc. In practice, it is found that the accuracy of text matching is low and the efficiency of resource query is low. Therefore, how to improve the accuracy of text matching is a technical problem that needs to be solved urgently in the prior art. Summary of the invention

[0003] The embodiments of the present application provide a text processing method, device, electronic device and storage medium to solve the problem of low text matching accuracy.

[0004] According to one aspect of an embodiment of the present application, a text processing method is provided, the method comprising: determining a matching weight matrix of a query text relative to a text to be matched, the matching weight matrix comprising at least one item of a first weight matrix and a second weight matrix, the first weight matrix being determined according to the segmentation category to which each segmentation in the query text belongs and the segmentation category to which each segmentation in the text to be matched belongs, and the second weight matrix being determined according to the association relationship between each segmentation in the query text and each segmentation in the text to be matched; enhancing a similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix, the similarity matrix of the query text relative to the text to be matched being obtained by performing similarity calculation based on word vectors of each segmentation in the query text and word vectors of each segmentation in the text to be matched; determining a matching score between the query text and the text to be matched according to the first similarity matrix; and determining a target matching text according to the matching score between the query text and the text to be matched.

[0005] According to one aspect of the embodiments of the present application, a text processing device is provided, including: a matching weight matrix determination module, configured to determine a matching weight matrix of a query text relative to a text to be matched, where the matching weight matrix includes at least one of a first weight matrix and a second weight matrix; the first weight matrix is determined according to the word segmentation categories to which the word segments in the query text belong and the word segmentation categories to which the word segments in the text to be matched belong; the second weight matrix is determined according to the association relationship between the word segments in the query text and the word segments in the text to be matched; an enhancement module, configured to enhance a similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix; the similarity matrix is obtained by performing a similarity calculation based on the word vectors of the word segments in the query text and the word vectors of the word segments in the text to be matched; a matching degree score determination module, configured to determine a matching degree score between the query text and the text to be matched according to the first similarity matrix; and a target matching text determination module, configured to determine a target matching text according to the matching degree score between the query text and the text to be matched.

[0006] In some embodiments of the present application, based on the foregoing solution, the matching weight matrix includes a first weight matrix; the matching weight matrix determination module includes: a word segmentation category recognition unit, configured to recognize the word segmentation categories to which the word segments in the query text belong; a first weight determination unit, configured to determine a first weight of each word segment in the query text relative to each word segment in the text to be matched according to the word segmentation categories to which the word segments in the query text belong, the word segmentation categories to which the word segments in the text to be matched belong, and weight mapping information; the weight mapping information indicates a first weight associated with any two word segmentation categories; and a first weight matrix determination unit, configured to combine the first weights of all the word segments in the query text relative to all the word segments in the text to be matched to obtain the first weight matrix.

[0007] In some embodiments of the present application, based on the foregoing solution, the word segmentation category recognition unit includes: a first entity link information acquisition unit, configured to acquire first entity link information, where the first entity link information is obtained by performing entity linking on the word segments in the query text in a knowledge graph; and a word segmentation category determination unit, configured to use the word segmentation category to which the first entity to which the word segments in the query text are linked in the knowledge graph belongs as the word segmentation category to which the word segments in the query text belong.

[0008] In some other embodiments of the present application, based on the foregoing solution, the matching weight matrix includes a second weight matrix; the matching weight matrix determination module includes: an association relationship recognition unit, configured to recognize the association relationship between the word segments in the query text and the word segments in the text to be matched according to the knowledge graph; a second weight determination unit, configured to perform weight lookup according to the association relationship to obtain the second weight of the word segments in the query text relative to the word segments in the text to be matched; and a second weight matrix determination unit, configured to combine the second weights of each word segment in the query text relative to each word segment in the text to be matched to obtain the second weight matrix.

[0009] In some embodiments of the present application, based on the foregoing solution, the association relationship recognition unit includes: a first entity link information acquisition unit, configured to acquire first entity link information, where the first entity link information is used to indicate the first entity linked to by the word segment in the query text on the knowledge graph; a second entity link information acquisition unit, configured to acquire second entity link information, where the second entity link information is used to indicate the second entity linked to by the word segment in the text to be matched on the knowledge graph; and an association relationship determination unit, configured to determine the association relationship between the first entity and the second entity in the knowledge graph as the association relationship between the corresponding word segments in the query text and the corresponding word segments in the text to be matched.

[0010] In some embodiments of the present application, based on the foregoing solution, the enhancement module is further configured to: multiply the matching weight matrix by the similarity matrix to obtain the first similarity matrix.

[0011] In some embodiments of the present application, based on the foregoing solution, the matching degree score determination module includes: a pooling processing unit, configured to perform pooling processing on the first similarity matrix to obtain a second similarity matrix; and a matching degree score calculation unit, configured to calculate the matching degree score between the query text and the text to be matched according to the second similarity matrix.

[0012] In some embodiments of the present application, based on the foregoing solution, the matching degree score calculation unit includes: an attention weighting unit, configured to perform attention weighting on the second similarity matrix based on the attention mechanism to obtain a third similarity matrix; and a score prediction unit, configured to perform score prediction according to the third similarity matrix to obtain the matching degree score between the query text and the text to be matched.

[0013] In some embodiments of the present application, based on the foregoing solution, the attention weighting unit includes: a key matrix determination unit, configured to perform a linear transformation on the word vectors of each word segment in the query text according to the key weight vector to obtain the key vectors corresponding to each word segment in the query text; a query matrix determination unit, configured to perform a linear transformation on the semantic feature vector corresponding to the query text according to the query weight vector to obtain a query vector; an attention score determination unit, configured to calculate the attention scores corresponding to each word segment in the query text according to the key vectors corresponding to each word segment in the query text and the query vector; a target similarity vector determination unit, configured to weight the value vectors corresponding to each word segment in the query text according to the attention scores corresponding to each word segment in the query text to obtain the target similarity vectors corresponding to each word segment in the query text; the value vector corresponding to a word segment in the query text is obtained by performing a linear transformation on the similarity vector corresponding to the word segment in the query text according to the value weight vector, and the similarity vector corresponding to a word segment in the query text is obtained by extracting the elements related to the corresponding word segment in the query text from the second similarity matrix and combining the extracted elements; a third similarity matrix determination unit, configured to combine the target similarity vectors corresponding to each word segment in the query text to obtain the third similarity matrix.

[0014] In some embodiments of the present application, based on the foregoing solution, the second similarity matrix includes a first pooling matrix and a second pooling matrix; the pooling processing unit includes: a first pooling processing unit, configured to perform pooling processing on the first similarity matrix along the horizontal direction of the first similarity matrix to obtain the first pooling matrix; a second pooling processing unit, configured to perform pooling processing on the first similarity matrix along the vertical direction of the first similarity matrix to obtain the second pooling matrix.

[0015] In some embodiments of the present application, based on the foregoing solution, the target matching text determination module includes: a sorting unit, configured to sort a plurality of texts to be matched in descending order of the matching degree scores; a target matching text determination unit, configured to determine the texts to be matched in the top set number in the sorting as the target matching texts.

[0016] In some embodiments of the present application, based on the foregoing solution, the text processing device further includes: a service query request receiving module, configured to receive a service query request sent by a client, where the service query request indicates the query text; and further includes: an application information obtaining module, configured to obtain application information of a target service application, where the target service application refers to the service application corresponding to the target matching text; a query result generating module, configured to generate a query result according to the application information; and a query result returning module, configured to return the query result to the client, so that the client displays a service entry of the target service application according to the query result.

[0017] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the text processing method as described above is implemented.

[0018] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the text processing method as described above is implemented.

[0019] In the solution of the present application, the similarity matrix of the query text relative to the text to be matched is enhanced by the matching weight matrix of the query text relative to the text to be matched. The matching weight matrix includes at least one of a first weight matrix and a second weight matrix. The first weight matrix is determined according to the token categories to which the tokens in the query text belong and the token categories to which the tokens in the text to be matched belong. The second weight matrix is determined according to the association relationship between the tokens in the query text and the tokens in the text to be matched. Thus, in the process of calculating the matching degree score between the query text and the text to be matched, in addition to referring to the similarity calculated based on the word vectors, the factors of the association relationship between the tokens in the query text and the tokens in the text to be matched, and / or the factors of the token categories to which the tokens in the query text belong and the token categories to which the tokens in the text to be matched belong are also referred to. Compared with the prior art that only refers to the similarity calculated based on the word vectors to calculate the matching score between two texts, the solution of the present application refers to more dimensions of factors that affect the matching of two texts to comprehensively calculate the matching degree score between two texts, which can ensure that the calculated matching degree score between two texts accurately reflects the actual matching degree between the two texts and effectively improves the accuracy of text matching. Description of the Drawings

[0020] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1A A schematic diagram showing an exemplary system architecture to which the technical solution of the embodiment of this application can be applied.

[0022] Figure 1B A schematic diagram of a query interface shown according to a specific embodiment of this application.

[0023] Figure 1C A schematic diagram of a secondary query interface for the "Music" option shown according to a specific embodiment of this application.

[0024] Figure 2 A flowchart of a text processing method shown according to an embodiment of this application.

[0025] Figure 3 Shown according to an embodiment of this application Figure 2 A flowchart corresponding to step 210 in the corresponding embodiment.

[0026] Figure 4 A schematic diagram showing the identification of word segmentation categories and the association relationship between two word segmentations based on a knowledge graph according to a specific embodiment of this application.

[0027] Figure 5 Shown according to another embodiment of this application Figure 2 A flowchart corresponding to step 210 in the corresponding embodiment.

[0028] Figure 6 Shown according to an embodiment of this application Figure 2 A flowchart corresponding to step 230 in the corresponding embodiment.

[0029] Figure 7 Shown according to an embodiment of this application Figure 6 A flowchart corresponding to step 620 in the corresponding embodiment.

[0030] Figure 8 A schematic diagram showing attention weighting based on an attention mechanism according to a specific embodiment of this application.

[0031] Figure 9 A flowchart of a text processing method shown according to another embodiment of this application.

[0032] Figure 10It is a schematic diagram of a text matching model shown according to an embodiment of the present application.

[0033] Figure 11 It is a block diagram of a text processing device shown according to an embodiment of the present application.

[0034] Figure 12 It shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0036] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0037] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0039] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0040] Before further describing the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained and described.

[0041] Vertical search: Also known as vertical search, it refers to a professional search engine for a specific industry. It is a subdivision and extension of the search engine, which integrates a certain type of specialized information in the database, extracts the required data by fields in a targeted manner, processes it, and then returns it to the user in a certain form. Vertical search can provide valuable information and related services for a specific field, a specific group of people, or a specific need. The public account search and mini-program search provided in social application platforms such as WeChat can be regarded as a type of vertical search.

[0042] Service search: It refers to a search conducted with services as the retrieval target. Through service search, services that match the query text entered by the user can be directly displayed to the user. For example, when searching for a nanny, service search can directly provide the nanny-finding service menu; another example is that when searching for sending express delivery, service search can directly provide the service entry for sending express delivery, and the user can directly trigger this service entry to enter the page for sending express delivery. Service search is a type of vertical search.

[0043] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0044] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0045] Natural Language Processing (NLP) is a discipline that takes language as its object and uses computer technology to analyze, understand, and process natural language. It uses the computer as a powerful tool for language research, conducts quantitative research on language information with the support of the computer, and provides a language description that can be commonly used between humans and computers. It includes two parts: Natural Language Understanding (NLU) and Natural Language Generation (NLG).

[0046] Text matching is widely used in scenarios such as resource query (or resource retrieval), content recommendation, and intelligent question answering. An important link in text matching is natural language understanding. Through natural language understanding technology, computer devices can understand the semantics of text, and based on this, text matching is carried out.

[0047] In related technologies, there is a problem of low text matching accuracy. In the scenario of resource retrieval, a low text matching degree will cause users to spend more time further screening resources from the retrieval results, or they cannot obtain the required resources from the retrieval results. Based on this, in order to solve the problem of low text matching accuracy in related technologies, the solution of this application is proposed.

[0048] Figure 1A The figure shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of this application can be applied. As Figure 1A shown, the system architecture may include terminal devices (such as Figure 1A one or more of the smart phone 101, tablet computer 102, and portable computer 103 shown in the figure. Of course, it can also be a desktop computer, etc.), network 104, and server 105. The network 104 is used to provide a medium for the communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired communication links, wireless communication links, and so on.

[0049] It should be understood that Figure 1A the numbers of terminal devices, networks, and servers in the figure are only illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. For example, the server 105 can be a server cluster composed of multiple servers, etc.

[0050] In some embodiments of the present application, the text processing method may be executed by the server 105. A user may input a query text through a terminal device and initiate a query request to the server 105. After receiving the query request, the server 105 extracts the query text from the query request, determines the matching degree scores between the query text and each text to be matched according to the method of the present application, and further filters out the texts to be matched with relatively high matching degree scores with the query text as the target matching texts, so as to generate a query result based on the target matching texts and return the query result to the terminal device.

[0051] It can be understood that the terminal device provides a query entry for the user to input the query text. After detecting that the query text is input at the query entry, the terminal device may initiate a query request to the server based on the query text.

[0052] In some embodiments of the present application, in a scenario where multiple data sources are provided in the user interface of the terminal device, the user may further select a data source, and the selected data source is used as the target data source. After receiving the query request sent by the terminal device, the server 105 determines the matching degree scores between each text to be matched in the target data source selected by the user and the query text, and further determines the query result corresponding to the query text based on the texts to be matched in the target data source.

[0053] In an application scenario of the present application, the query text may be a text input for video query, image query, audio query, text content (such as news, blogs, papers, official account articles, etc.) query, service application query, commodity query, mini-program query, official account query, advertisement query, emoji query. The query text may include one or more words.

[0054] In each query scenario, the data source correspondingly provides resource description texts corresponding to each resource, and the resource description text is the text to be matched in the solution of the present application. The resources are, for example, the videos, images, audios, text contents, mini-programs, service applications, official accounts, emojis, etc. listed above. The resource description text refers to the text used to describe the corresponding resource. For example, the video description corresponding to the video, the image description text corresponding to the image, the audio description text corresponding to the audio, the lyrics and / or song name corresponding to the song, the service description text of the service application, etc.; the resource description text may also be other aspects of text content, such as one or more of the text tags marked for the resource, the name of the resource, the evaluation content of other users for the resource, etc., which are not specifically limited herein.

[0055] Of course, since the types of resources to be queried are different, after the server 105 determines the target matching text corresponding to the query text, the query results returned to the terminal device also differ accordingly. For example, if the query request is initiated to query service applications, then after determining the target matching text corresponding to the query text, a query result is generated based on the service application associated with the target matching text, and this query result is used to indicate the service application associated with the target matching text; if the query request is initiated to query videos, then the query result is used to indicate the video associated with the target matching text; if the query request is initiated to query products, then the query result is used to indicate the product associated with the target matching text.

[0056] In some embodiments of the present application, the query entry provided in the query interface can be used to query a single type of resource or multiple types of resources. For example, based on the query text input in the query entry, queries for mini-programs, service applications, emojis, and official accounts can be performed.

[0057] Figure 1B is a schematic diagram of a query interface shown according to a specific embodiment. As Figure 1B shown, the query interface provides a query entry 111, and query text can be input in the query entry 111. The query text can be obtained by directly inputting text, or by converting the input voice into text through speech-to-text conversion, or by performing optical character recognition on the text in the input image.

[0058] Figure 1B The interface shown also includes a resource type selection area 112, and a plurality of selectable resource type options are provided in the resource type selection area. Figure 1B The resource type options in the resource type selection area 112 shown include an "article" option, an "official account" option, a "mini-program" option, a "music" option, an "emoji" option, and a "service" option. It is worth mentioning that Figure 1B the resource type options shown in the resource type selection area are merely exemplary. In other embodiments, the resource type selection area may also include more or fewer other resource type options, which can be specifically set according to actual needs. By providing resource type options, users can select the resource type according to actual query needs, and then perform targeted queries in the resource library corresponding to the selected resource type. The query results will not include resources in the resource libraries corresponding to unselected resource types, which facilitates users to quickly obtain the required resources from the query results and improves the resource query efficiency. Performing a resource query in the resource library corresponding to the resource type selected by the user is a vertical search.

[0059] A resource library is set for each resource type option. For example, for the "article" option, an article resource library is set, which includes multiple articles available for query and the description text corresponding to the articles. Another example is that for the "service" option, a service application resource library is set, which includes the call interfaces of multiple service applications that users can enter and the description text corresponding to the service applications. Based on the call interfaces of the service applications, the application entry of the service application can be displayed in the query result interface, and the user can enter the interface of the service application by triggering the application entry of a service application.

[0060] If the user selects a resource type option in the resource type selection area 112, the target matching text corresponding to the query text is determined in the resource library corresponding to the selected resource type option, and then the target resource for the target matching text is determined. For example, if the user selects the "emoji" option, the target matching text corresponding to the query text is determined in the resource library corresponding to the "emoji" option, and then the emoji image corresponding to the target matching text is determined.

[0061] In some embodiments of the present application, if the user selects a resource type option in the resource type selection, it is also possible to Figure 1B jump from the shown query interface to the secondary query interface for the selected resource type option. Figure 1C The secondary query interface for the "music" option is shown. If the user enters the query text in the query entry 111 of the Figure 1C shown secondary query interface, a corresponding request is made to perform text matching based on the query text in the resource library corresponding to the "music" option (i.e., the music resource library), to determine the target matching text that matches the query text, and then the music corresponding to the target matching text is determined. Further, as Figure 1C shown, popular search content is further provided in this secondary query interface for the user to select.

[0062] In some embodiments of the present application, if the user does not select a resource type in the Figure 1B shown query interface, query text matching can be performed in multiple resource libraries based on the query text. Thus, the query results obtained can include resources of multiple resource types. For example, if the query text is "home appliance repair" and the user does not select a resource type in the Figure 1B shown query interface, the query results can include resources such as the matched articles, service applications, mini-programs, emojis, etc.

[0063] In other application scenarios, Figure 1BThe query interface shown can also be a query entry provided for resource queries for one or more specified resource types, so that users do not need to additionally select resource type options. Correspondingly, the resource type selection area may not be set either.

[0064] In other embodiments, the method of the present application may also be executed by a terminal device with computing and processing capabilities, or by a system composed of a terminal device and a server, which is not specifically limited herein.

[0065] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below.

[0066] Figure 2 is a flowchart of a text processing method shown according to an embodiment of the present application. The method can be executed by a computer device with processing capabilities, such as a server or a terminal device, which is not specifically limited herein. Refer to Figure 2 As shown, the method at least includes steps 210 to 240, which are introduced in detail as follows:

[0067] Step 210, determining a matching weight matrix of the query text with respect to the text to be matched. The matching weight matrix includes at least one of a first weight matrix and a second weight matrix. The first weight matrix is determined according to the word segmentation categories to which the word segments in the query text belong and the word segmentation categories to which the word segments in the text to be matched belong. The second weight matrix is determined according to the association relationship between the word segments in the query text and the word segments in the text to be matched.

[0068] The query text refers to the text input for resource query. The query text may include one or more words; the query text may be a sentence or a paragraph that conforms to grammar rules, or may also be a text formed by combining several independent phrases or words. The query text may be directly input by the user, or may be obtained by converting the user's input voice into text, or may also be obtained by performing character recognition on the picture input by the user.

[0069] To enable users to perform resource queries, a resource library is correspondingly provided. The resource library includes multiple resources. The text to be matched refers to the resource description text used to describe the resources in the resource library. Each resource in the resource library corresponds to a resource description text. The resource description text corresponding to the resource may include one or more of a resource name, a resource introduction, a resource source, a resource label, and evaluation content of other users for the resource.

[0070] Resources of different types may have different corresponding resource description texts. For example, if the resource type is a service, the service description text corresponding to the service application may include the name of the service application and the brief introduction information of the service application; if the resource type is an official account article, the resource description text corresponding to the official account article may be one or more of the brief introduction of the official account article (such as the abstract of the official account article), the author of the official account article, the tags of the official account article (such as keywords), the article title, etc.

[0071] The matching weight matrix of the query text relative to the text to be matched is used to indicate the matching weights of the word segments in the query text relative to the word segments in the text to be matched. Among them, an element in the matching weight matrix is used to represent the matching weight of a word segment in the query text relative to a word segment in the text to be matched.

[0072] In some embodiments of the present application, the first weight matrix can be used alone as the matching weight matrix, the second weight matrix can be used alone as the matching weight matrix, or both the first weight matrix and the second weight matrix can be used as the matching weight matrix.

[0073] In order to determine the first weight matrix or the second weight matrix of the query text relative to the text to be matched, it is necessary to first segment the query text and the text to be matched to determine the word segments in the query text and the word segments in the text to be matched.

[0074] In some embodiments of the present application, the query text and the text to be matched can be segmented according to a dictionary. In some scenarios, since resources in some professional fields involve many professional terms, such as the medical field, the material field, the chemical engineering field, etc., these professional terms are quite different from common general terms in daily life. Therefore, in order to ensure the accuracy of word segmentation, in the retrieval of resources related to a professional field, a dictionary related to the professional field can be pre-constructed for the professional terms that may be involved in the professional field, and then the query text and the text to be matched related to the professional field can be segmented with the constructed dictionary related to the professional field.

[0075] An element in the first weight matrix represents a matching weight of a word segment in the query text relative to a word segment in the text to be matched (for the sake of distinction, the matching weight represented in the first weight matrix is called the first weight); an element in the second weight matrix represents another matching weight of a word segment in the query text relative to a word segment in the text to be matched (the matching weight represented by the element in the second weight matrix is called the second weight).

[0076] In some embodiments of the present application, token pairs can be constructed first based on the tokens in the query text and the tokens in the text to be matched. Among them, one token in the token pair comes from the query text, and the other token comes from the text to be matched; on this basis, the elements in the first weight matrix are determined based on the token categories to which the two tokens in the token pair belong respectively, and the elements in the second weight matrix are determined based on the association relationship between the two tokens in the token pair.

[0077] In some embodiments of the present application, the mapping relationship between tokens and token categories can be preset to obtain a token-token category mapping data set; then, the tokens in the query text are searched in this token-token category mapping data set, and the token category corresponding to the searched token is used as the token category to which the token in the query text belongs. The token categories to which the tokens in the text to be matched belong can also be determined according to this method, which will not be elaborated here.

[0078] The token categories in the token-token category data set can be set according to actual needs. In some embodiments of the present application, the token categories can include entity words, state words, action words, brand words, compound words, etc. Among them, entity words refer to words used to represent entities, state words refer to words used to represent states, action words refer to words used to represent actions and behaviors, brand words refer to words used to represent brand names, and compound words refer to phrases formed by combining at least two words. For example, under the above setting of token categories, if the tokenization result of a text is: Brand A / air conditioner / clean / arrive home / service, in this tokenization result, "Brand A" is a brand word, "air conditioner" is an entity word, "clean" is an action word, "arrive home" is a state word, and "service" is an action word; if "arrive home service" is regarded as one token during the tokenization process, then "arrive home service" is a compound word.

[0079] The above is only an exemplary example of token categories. In other embodiments, more or fewer token categories can also be set according to other classification principles. For example, in the chemical field, entity words can be further divided into chemical substance name words, instrument name words, person name words, etc.

[0080] In some embodiments of the present application, on the basis of setting token categories, the first weight corresponding to the token category pair formed by any two token categories is further set. On this basis, the token category pair formed by the token categories to which the two tokens in the token pair belong can be formed, and then the first weight corresponding to this token category pair is searched. The searched first weight is the first weight corresponding to the token pair. Repeating this process can correspondingly determine each element in the first weight matrix.

[0081] In some embodiments of the present application, a second weight corresponding to each association relationship is preset. After determining the association relationship between the two words in each word pair, the second weight corresponding to the association relationship is used as the second weight corresponding to the word pair. The second weights corresponding to all the word pairs constructed from the words in the query text and the text to be matched are combined to obtain a second weight matrix.

[0082] The association relationship between two words can be a relationship determined based on the semantics of the two words. For example, it can be a hypernym (hyponym) relationship, a synonym relationship, an antonym relationship, etc. The association relationship between two words can also be a relationship determined based on the general collocation of the two words, such as an extended word relationship. The extended word relationship means that the two words have no semantic association, but the two words are usually used in combination. For example, in "home appliance cleaning", the two words "home appliance" and "cleaning" have no semantic association, but the two words are usually used in combination. Therefore, the two words have an extended word relationship. If the two words have neither a semantic relationship nor a general collocation, the association relationship between the two words can be set as other relationships. The association relationships listed above are only exemplary examples and should not be considered as a limitation on the scope of use of the present application. In other embodiments, more or fewer association relationships can also be set.

[0083] In some embodiments of the present application, for the convenience of matrix calculation, the elements in the first weight matrix and the elements in the second weight matrix can be arranged in the same order. For example, in the first weight matrix, the horizontal order is arranged according to the order of the words corresponding to the elements in the text to be matched, and the vertical order is arranged according to the order of the words corresponding to the elements in the query text. Then, the elements in the second weight matrix are also arranged in the same order.

[0084] Step 220: According to the matching weight matrix, enhance the similarity matrix of the query text relative to the text to be matched to obtain a first similarity matrix. The similarity matrix of the query text relative to the text to be matched is obtained by calculating the similarity between the word vectors of the words in the query text and the word vectors of the words in the text to be matched.

[0085] A word vector (Word embedding) is a vector obtained by mapping a word or phrase to a real number. This kind of mapping can be generated by a neural network, a probability model, or an interpretable knowledge base method.

[0086] Before step 220, after the query text and the text to be matched are respectively segmented, the words in the query text and the words in the text to be matched are respectively mapped to a real number vector space to obtain the word vectors of the words in the query text and the word vectors of the words in the text to be matched.

[0087] In some embodiments of the present application, the similarity matrix of the query text relative to the text to be matched can be determined in units of token pairs (where one token is from the query text and the other token is from the text to be matched). An element in the similarity matrix is used to reflect the similarity between the word vectors corresponding to the two tokens in a token pair. Combining the similarities between the word vectors corresponding to the two tokens in each token pair gives the similarity matrix of the query text relative to the text to be matched. Among them, the similarity between two word vectors can be cosine similarity, Euclidean distance, etc., which is not specifically limited here.

[0088] In some embodiments of the present application, step 220 further includes: multiplying the matching weight matrix by the similarity matrix to obtain a first similarity matrix.

[0089] Step 230, determining the matching degree score between the query text and the text to be matched according to the first similarity matrix.

[0090] Similar to the similarity matrix, each element in the first similarity matrix reflects the enhanced similarity between the two tokens in a token pair. After obtaining the first similarity matrix, the matching degree score can be calculated by comprehensively considering the enhanced similarities corresponding to all the tokens in all the token pairs, so that the matching degree score directly comprehensively reflects the matching degree between the query text and the text to be matched.

[0091] In some embodiments of the present application, a fully connected network can be used to perform classification prediction according to the first similarity matrix and output the matching degree score between the query text and the text to be matched.

[0092] Step 240, determining the target matching text according to the matching degree score between the query text and the text to be matched.

[0093] In some embodiments of the present application, in a scenario where there are multiple texts to be matched, the target matching text corresponding to the query text can be determined according to the following process: sorting the multiple texts to be matched in descending order of the matching degree score; determining the texts to be matched in the top set number in the sorting as the target matching text. The set number can be set according to actual needs, such as 5, 10, 20, 30, etc., which is not specifically limited here.

[0094] In some embodiments of the present application, a matching degree score threshold can also be set. If the matching degree score between a text to be matched and the query text is not lower than the matching degree score threshold, then the text to be matched is determined as the target matching text.

[0095] In some embodiments of the present application, after determining the target matching text, a query result can be determined based on the target matching text. The query result is used to indicate the resource corresponding to the target matching text, and the query result is returned to the user. If the query text is the text input for service application query, the resource corresponding to the target matching text is the service application associated with the target matching text; if the query text is the text input for advertisement query, the resource corresponding to the target matching text is the advertisement associated with the target matching text.

[0096] In text matching in the related art, generally, the matching degree score between two texts is calculated directly based on the similarity of each pair of segmented words in the two texts (the query text and a text to be matched). That is to say, the matching degree score only refers to the similarity between the word vectors of the two segmented words in the pair of segmented words.

[0097] The inventors of the present application found in practice that: if the co-occurrence frequency of two segmented words for which similarity calculation is required is relatively high, for example, they are conventional collocations, the existing word vector generation model will regard the two segmented words as segmented words with high semantic similarity due to the high co-occurrence frequency of the two segmented words. Therefore, the similarity of the word vectors generated for the two segmented words is also relatively high. In this case, if only the similarity between the word vectors corresponding to the two segmented words in the pair of segmented words is considered to calculate the matching degree score between the two texts, the relatively high similarity between the word vectors corresponding to the segmented words with high co-occurrence frequency may lead to a deviation between the matching degree reflected by the calculated matching degree score and the actual matching degree between the two texts, resulting in misidentifying texts that are actually not matched as matched texts. At the same time, the inventors of the present application also found that if the association relationship between two segmented words during the text matching process is a hypernym (hyponym) relationship, this situation will also introduce noise to the text matching. Generally speaking, the association relationship between the two segmented words in the pair of segmented words will affect the semantic matching degree of the two segmented words, and this kind of influence cannot be reflected from the similarity obtained by calculating the similarity of the word vectors of the two segmented words.

[0098] The inventors of the present application also found that in the related art, only the similarity of each pair of segmented words in the two texts is referred to for calculating the matching degree score between the two texts, without introducing other factors for calculating the matching degree score. This method is equivalent to treating the contribution degrees of each pair of segmented words to the text semantics equally, while in practice, the contribution degrees of different types of segmented words to the text semantics are different. Correspondingly, the contribution degrees of different pairs of segmented word types to the judgment of the matching degree between the two texts are also different.

[0099] Based on the association relationship between the segmented words discovered by the inventors of the present application and the influence of different types of segmented words on text matching noise, the solution of the present application is proposed.

[0100] In the solution of this application, the similarity matrix of the query text relative to the text to be matched is enhanced by querying the matching weight matrix of the query text relative to the text to be matched. The matching weight matrix includes at least one of a first weight matrix and a second weight matrix. The first weight matrix is determined according to the token categories to which the tokens in the query text belong and the token categories to which the tokens in the text to be matched belong. The second weight matrix is determined according to the association relationship between the tokens in the query text and the tokens in the text to be matched. Thus, in the process of calculating the matching degree score between the query text and the text to be matched, in addition to referring to the similarity calculated based on the word vectors, the factors of the association relationship between the tokens in the query text and the tokens in the text to be matched, and / or the factors of the token categories to which the tokens in the query text belong and the token categories to which the tokens in the text to be matched belong are also referred to. Compared with the prior art that only calculates the matching score between two texts by referring to the similarity calculated based on word vectors, the solution of this application refers to more dimensions of factors that affect the matching of two texts to comprehensively calculate the matching degree score between the two texts, which can ensure that the calculated matching degree score between the two texts accurately reflects the actual matching degree between the two texts and effectively improves the accuracy of text matching.

[0101] In some embodiments of this application, the matching weight matrix includes the first weight matrix. In this embodiment, as Figure 3 shown, step 210 includes:

[0102] Step 310, identify the token categories to which the tokens in the query text belong.

[0103] In some embodiments of this application, the token categories to which the tokens in the query text belong can be identified according to the token-token category data set constructed above.

[0104] In some other embodiments of this application, please continue to refer to Figure 3 , step 310 further includes steps 311 and 312. Among them, step 311, obtain the first entity link information, and the first entity link information is obtained by performing entity linking on the tokens in the query text in the knowledge graph.

[0105] The knowledge graph includes a plurality of interrelated graph nodes, and each graph node represents a word. The relationship between the graph nodes is used to indicate the association relationship between the two words located at the two graph nodes. In the solution of this application, the token category to which the word located at the graph node belongs is further set. The word located at the graph node can also be called an entity.

[0106] Entity linking a token in a knowledge graph means finding and determining the entity corresponding to the token in the knowledge graph, and the determined entity is the entity linked to by the token in the knowledge graph. For ease of distinction, the entity linked to by a token in the query text in the knowledge graph is called the first entity.

[0107] Due to the existence of cases of homonyms with different entities (the same word has multiple different meanings in different contexts) or the same entity with different names (the same thing can have multiple ways of referring, for example, "potato" and "sweet potato" refer to the same object), therefore, in order to avoid ambiguity, entity linking is performed on the tokens in the query text to clarify the entities referred to by each token in the query text.

[0108] In some embodiments of the present application, due to the knowledge of the association relationships between entities and the token categories to which the entities belong in the knowledge graph, therefore, the token categories to which each token in the query text and the text to be matched belong can be identified through the knowledge graph, and the association relationship between a token in the query text and a token in the text to be matched can be identified based on the knowledge graph.

[0109] Figure 4 FIG. shows a schematic diagram of identifying token categories and the association relationship between two tokens through a knowledge graph according to a specific embodiment of the present application. Among them, Figure 4 The knowledge graph shown in FIG. is constructed for the "housekeeping category". In other embodiments, knowledge graphs can also be constructed for each category respectively based on the category division performed, for example, knowledge graphs are constructed for the beauty category and the health care category respectively.

[0110] Figure 4 The shown knowledge graph can be used for text matching in service application queries. In this knowledge graph, the token categories include: service entity words, service behavior words, service status words, and service brands. Among them, service entity words refer to entity words related to services; service behavior words refer to behavior words related to services; service status words refer to status words related to services; service brands refer to brand words related to services.

[0111] Please continue to refer to Figure 4 , Figure 4 FIG. shows some entities in the knowledge graph (the circled part in Figure 4 ), including "household appliances", "air conditioner", "cleaning", "cleanliness", "Brand A", "home delivery"; among them, "household appliances" and "air conditioner" are service entity words; "cleaning" and "cleanliness" are service status words; "home delivery" is a service status word; "Brand A" is a service brand. Further, as Figure 4As shown, the relationship between the entity "household appliances" and the entity "air conditioner" is a hypernym relationship; the relationship between the entity "cleaning" and the entity "cleanliness" is a synonym relationship.

[0112] Based on Figure 4 the knowledge graph shown, if the two texts for which the matching degree score needs to be calculated include Text I and Text II, where the word segmentation result of Text I is: Brand A / household appliances / cleaning; the word segmentation result of Text II is: air conditioner / cleaning, / door-to-door / service. Link the word segments in Text I to entities in the knowledge graph, as Figure 4 shown, the word segment "Brand A" in Text I is linked to the entity "Brand A" in this knowledge graph; the word segment "household appliances" in Text I is linked to the entity "household appliances" in the knowledge graph; the word segment "cleaning" in Text I is linked to the entity "cleaning" in the knowledge graph.

[0113] Link the word segments in Text II to entities in the knowledge graph, as Figure 4 shown, the word segment "air conditioner" in Text II is linked to the entity "air conditioner" in the knowledge graph; the word segment "cleaning" in Text II is linked to the entity "cleaning" in the knowledge graph; the word segment "door-to-door" in Text II is linked to the entity "door-to-door" in the knowledge graph. Figure 4 Only some entities in the knowledge graph are shown. Therefore, Figure 4 the entities linked to in the knowledge graph for Text II are not shown in

[0114] After linking the word segments in Text I and Text II to entities in the knowledge graph, the word segment category to which the word segment belongs and the relationship between the two word segments can be determined based on the entities linked to by the word segments. For example, the word segment category to which the entity linked to by the word segment in the knowledge graph belongs can be determined as the word segment category to which the word segment belongs; the relationship between the entities linked to by the two word segments in the knowledge graph can also be determined as the relationship between the two word segments.

[0115] In some embodiments of the present application, since the entity linked to by the word segment in the query text in the knowledge graph is the same or similar in semantics to the word segment, the entity linked to by the word segment in the knowledge graph can be used as reference information for generating the word vector of the word segment. In other words, the word vector of the word segment is determined by combining the word segment and the entity linked to by the word segment in the knowledge graph. For the sake of distinction, the entity linked to by the word segment in the query text in the knowledge graph is referred to as the first entity.

[0116] In some embodiments of the present application, in order to facilitate the construction of the word vectors of the word segments in the query text, the word vectors of the entities linked to by the word segments in the query text in the knowledge graph can be used as the word vectors of the word segments.

[0117] In some embodiments of the present application, since there is a relatively close association relationship among the neighbor entities of an entity in the knowledge graph, in order to provide more reference information in the process of generating the word vector of a word segment, the neighbor entities of the entity linked to by the word segment in the knowledge graph can also be used as the reference information for generating the word vector of the word segment.

[0118] In some embodiments of the present application, since the closer the distance between entities in the knowledge graph, the closer the relationship between the entities, such as similar semantics, collocation, etc., and the more reference information, the greater the computational complexity in the process of generating word vectors. Therefore, in order to balance the accuracy of word vectors and the computational complexity, the neighbor entities used as the reference for generating word vectors can be the entities directly adjacent to the first entity, and the entities directly adjacent to the first entity can also be referred to as the one-hop neighborhood of the first entity.

[0119] In some embodiments of the present application, after determining the first entity linked to by the word segment in the query text in the knowledge graph, the first entity and the one-hop neighborhood of the first entity can be spliced to the corresponding word segment in the query text. Thus, the obtained spliced text (for the sake of distinction, this spliced text is referred to as the first spliced text) is input into the word vector generation model, and the word vector generation model determines the word vectors of each word segment in the query text according to the first spliced text.

[0120] In some embodiments of the application, since the word segment in the query text is first entity-linked in the knowledge graph, and the first entity linked to is semantically the same or similar to the corresponding word segment in the query text, therefore, the first entity linked to by the word segment in the knowledge graph can be used as the data basis for generating the word vector of the corresponding word segment. Specifically, in the order of each word segment in the query text, the corresponding linked first entities are spliced, and the neighbor entities (such as one-hop neighbors) of the first entity are spliced to the corresponding first entity. The spliced text obtained from this splicing process (for the sake of distinction, this spliced text is referred to as the second spliced text) is input into the neural network, and the neural network outputs the word vectors of each word segment in the query text according to the second spliced text.

[0121] In the scenario of generating the word vectors of each word segment in the query text according to the first spliced text or the second spliced text as above, the first entity link information can be the first spliced text or the second spliced text obtained based on entity linking. Of course, the word vectors of each word segment in the text to be matched can be generated according to the above method for generating the word vectors of each word segment in the query text, which will not be elaborated here.

[0122] Step 312: Use the token category to which the token in the query text linked to the first entity belongs in the knowledge graph as the token category to which the token in the query text belongs.

[0123] After determining the first entity to which each token in the query text is linked in the knowledge graph, since the linked first entity has the same or similar semantics as the token, based on the token categories to which each entity belongs set in the knowledge graph, the token category to which the token in the query text linked to the first entity belongs can be directly used as the token category to which the corresponding token in the query text belongs.

[0124] Through the process of steps 311 - 312 above, it is achieved to identify the token categories to which each token in the query text belongs based on the knowledge graph, and at the same time, the token categories to which each token in the text to be matched belongs can also be identified in a similar way.

[0125] Step 320: Determine the first weight of each token in the query text relative to each token in the text to be matched according to the token categories to which each token in the query text belongs, the token categories to which each token in the text to be matched belongs, and the weight mapping information; the weight mapping information indicates the first weight associated with any two token categories.

[0126] After determining the token categories to which each token in the query text belongs and the token categories to which each token in the text to be matched belongs, the token categories to which the two tokens in each token pair (one token comes from the query text and the other token comes from the text to be matched) belong can be correspondingly determined, obtaining the token category pair corresponding to each token pair, so as to find the token category pair corresponding to the token pair in the weight mapping information, and further determine the first weight associated with the token category pair, that is, the first weight of the corresponding token in the query text relative to the corresponding token in the text to be matched.

[0127] Step 330: Combine the first weights of all tokens in the query text relative to all tokens in the text to be matched to obtain the first weight matrix.

[0128] In some embodiments of the present application, the first weights of each token in the query text relative to each token in the text to be matched can be combined in a set order. The set order is, for example, the horizontal order exemplified above is sorted according to the positions of the tokens corresponding to the elements in the text to be matched, and the vertical order is sorted according to the positions of the tokens corresponding to the elements in the query text.

[0129] Of course, for the convenience of calculation, the arrangement order of the elements in the first weight matrix is the same as the arrangement order of the elements in the similarity matrix.

[0130] In other embodiments of the present application, the matching weight matrix includes a second weight matrix; as Figure 5 shown, step 210 includes:

[0131] Step 510, identify the association relationship between the word segments in the query text and the word segments in the text to be matched according to the knowledge graph.

[0132] In some embodiments of the present application, step 510 includes: obtaining first entity link information, where the first entity link information is used to indicate the first entity linked to by the word segment in the query text on the knowledge graph; obtaining second entity link information, where the second entity link information is used to indicate the second entity linked to by the word segment in the text to be matched on the knowledge graph; and determining the association relationship between the first entity and the second entity in the knowledge graph as the association relationship between the corresponding word segments in the query text and the corresponding word segments in the text to be matched.

[0133] The second entity refers to the entity linked to by the word segment in the text to be matched on the knowledge graph. In some embodiments of the present application, similar to the first entity link information, the second entity link information may be the spliced text obtained by splicing the second entities linked to by each word segment in the text to be matched to the corresponding word segments in the text to be matched and splicing the neighbor entities of the second entity to the corresponding second entity, or may be the spliced text obtained by splicing the second entities linked to by each word segment according to the positions of the word segments in the text to be matched and splicing the neighbor entities of the second entity to the corresponding second entity.

[0134] Since the connection relationship between two entities in the knowledge graph indicates the association relationship between the two entities, therefore, after determining the first entity (assumed to be entity P1) linked to by the word segment (assumed to be word segment A) in the query text and the second entity (assumed to be entity P2) linked to by the word segment (assumed to be word segment B) in the text to be matched, because the entity P1 linked to by the word segment A on the knowledge graph is highly matched, for example, the semantics are the same, the part of speech is the same or similar, and the entity P2 linked to by the word segment B on the knowledge graph is also highly matched. Therefore, in this sense, the word segment A can be replaced by the linked entity P1, the word segment B can be replaced by the linked entity P2, and further the association relationship between the entity P1 and the entity P2 on the knowledge graph can represent the association relationship between the word segment A and the word segment B. Therefore, the association relationship between the entity P1 and the entity P2 on the knowledge graph can be determined as the association relationship between the word segment A in the query text and the word segment B in the text to be matched.

[0135] Step 520, perform a weight search according to the association relationship to obtain the second weight of the word segment in the query text relative to the word segment in the text to be matched.

[0136] In some embodiments of the present application, a second weight corresponding to each association relationship is preset. After determining the association relationship between the word segments in the query text and the word segments in the text to be matched, the second weight corresponding to the association relationship is obtained accordingly. The second weight of each word segment in the query text relative to each word segment in the text to be matched can be determined according to this process.

[0137] Step 530: Combine the second weights of each word segment in the query text relative to each word segment in the text to be matched to obtain a second weight matrix.

[0138] In one embodiment, for the convenience of calculation, the arrangement order of the elements in the second weight matrix is the same as that of the elements in the similarity matrix.

[0139] Through the above steps 510 - 530, the association relationship between the word segments in the query text relative to the word segments in the text to be matched is determined based on the knowledge graph, and the second weight corresponding to the determined association relationship is found, obtaining the second weight of the word segments in the query text relative to the word segments in the text to be matched.

[0140] If the matching weight matrix includes a first weight matrix and a second weight matrix, then the first weight matrix can be determined corresponding to the Figure 3 shown embodiment, and the second weight matrix can be determined corresponding to the Figure 5 shown embodiment, and then the similarity matrix between the query text and the text to be matched is enhanced using the first weight matrix and the second weight matrix.

[0141] In some embodiments of the present application, as Figure 6 shown, step 230 includes:

[0142] Step 610: Perform pooling processing on the first similarity matrix to obtain a second similarity matrix.

[0143] Performing pooling processing on the first similarity matrix means extracting important feature information from the first similarity matrix and downsampling the first similarity features. Therefore, the dimension of the second similarity matrix obtained by the pooling processing is smaller than that of the first similarity matrix, the number of parameters is reduced, and it is convenient for subsequent calculation of the matching degree score.

[0144] The pooling processing can be maximum pooling processing, average pooling processing, global average pooling, global adaptive pooling, etc., and is not specifically limited herein.

[0145] In some embodiments of the present application, the second similarity matrix includes a first pooling matrix and a second pooling matrix; step 610 further includes: performing pooling processing on the first similarity matrix along the horizontal direction of the first similarity matrix to obtain the first pooling matrix; performing pooling processing on the first similarity matrix along the vertical direction of the first similarity matrix to obtain the second pooling matrix.

[0146] When one direction dimension of the first similarity matrix represents the word segments in the query text and the other direction dimension represents the word segments in the text to be matched, by performing pooling processing on the first similarity matrix along the two direction dimensions (i.e., the horizontal direction and the vertical direction) respectively, a pooling matrix for reflecting the hit situation of the word segments in the query text in the text to be matched and a pooling matrix for reflecting the hit situation of the word segments in the text to be matched in the query text can be obtained respectively. Therefore, the obtained first pooling matrix and second pooling matrix reflect the matching situation between the query text and the text to be matched with the query text as the benchmark and the text to be matched as the benchmark respectively. Combining the first pooling matrix and the second pooling matrix can reflect the matching situation between the query text and the text to be matched from different perspectives.

[0147] Step 620, calculate the matching degree score between the query text and the text to be matched according to the second similarity matrix.

[0148] In some embodiments of the present application, the prediction of the matching degree score can be performed through a fully connected layer. Input the second similarity matrix into the fully connected layer, and the fully connected layer performs full connection on the second similarity matrix and predicts the matching degree score between the query text and the text to be matched.

[0149] In some embodiments of the present application, as Figure 7 shown, step 620 includes:

[0150] Step 710, perform attention weighting on the second similarity matrix based on the attention mechanism to obtain a third similarity matrix.

[0151] In a specific embodiment of the present application, step 710 includes: performing a linear transformation on the word vectors of each word segment in the query text according to the key weight vector to obtain the key vectors corresponding to each word segment in the query text; performing a linear transformation on the semantic feature vector corresponding to the query text according to the query weight vector to obtain the query vector; calculating the attention scores corresponding to each word segment in the query text based on the key vectors corresponding to each word segment in the query text and the query vector; weighting the value vectors corresponding to each word segment in the query text respectively according to the attention scores corresponding to each word segment in the query text to obtain the target similarity vectors corresponding to each word segment in the query text; the value vector corresponding to a word segment in the query text is obtained by performing a linear transformation on the similarity vector corresponding to the word segment in the query text according to the value weight vector, and the similarity vector corresponding to a word segment in the query text is obtained by extracting the elements related to the corresponding word segment in the query text from the second similarity matrix and combining the extracted elements; combining the target similarity vectors corresponding to each word segment in the query text to obtain the third similarity matrix.

[0152] For ease of description, the word segments in the query text are referred to as the first word segments. In this embodiment, for each first word segment, the attention score of the first word segment is calculated first, and then the similarity vector corresponding to the first word segment is weighted by the attention score of the first word segment to obtain the target similarity vector corresponding to the first word segment. The similarity vector corresponding to the first word segment is obtained by extracting the elements related to the first word segment from the second similarity matrix and combining the extracted elements.

[0153] Suppose the query text includes n word segments, and the word vector corresponding to the th word segment is , then the key vector corresponding to the th word segment in the query text is:

[0154] ; (Formula 1)

[0155] where is the key weight vector;

[0156] The query vector corresponding to the query text is:

[0157] ; (Formula 2)

[0158] where is the semantic feature vector corresponding to the query text, is the query weight vector; in a specific embodiment, the semantic feature vector corresponding to the query text can be obtained by fusing the word vectors of all word segments in the query text.

[0159] According to the scaled dot product model, the The attention score corresponding to the word segmentation is:

[0160] ; (Formula 3)

[0161] Among them, represents the transpose of the key vector corresponding to the -th word segmentation ; is a scaling factor, among which, can be the dimension of the key vector corresponding to the -th word segmentation. In other embodiments, the attention score can also be calculated based on the key vector and the query vector according to the additive model, dot product model or bilinear model.

[0162] The value vector corresponding to the -th word segmentation in the query text is:

[0163] ; (Formula 4)

[0164] Among them, is the similarity vector corresponding to the -th word segmentation, is the value weight matrix. In the above formula, , , are parameters that can be learned by the model.

[0165] The target similarity vector corresponding to the -th word segmentation in the query text is:

[0166] ; (Formula 5)

[0167] Or by combining Formula (3) and Formula (4), the target similarity vector corresponding to the -th word segmentation in the query text is obtained as:

[0168] ; (Formula 6)

[0169] Figure 8 is a schematic diagram showing the calculation of the target similarity vector corresponding to each word segmentation in the query text according to an embodiment of the present application. As Figure 8 shown, the semantic feature vector corresponding to the query text is linearly transformed according to the query weight matrix to obtain the query vector ; the word vectors of each word segmentation in the query text are linearly transformed according to the key weight vector to obtain the key vectors corresponding to each word segmentation in the query text ; perform a linear transformation on the similarity vector corresponding to the word segmentation in the query text according to the value weight vector to obtain the value vector corresponding to the word segmentation in the query text .

[0170] Then, steps 810-830 are sequentially executed. Among them, in step 810, the attention scores corresponding to each word segmentation in the query text are calculated through the MatMul function. Specifically, the attention scores are calculated according to the above formula (3). The MatMul function is used to multiply two matrices. In step 820, the obtained attention scores are normalized through the Softmax function. In step 830, the target similarity vectors corresponding to each word segmentation in the query text are calculated through the MatMul function. In step 830, the normalized attention scores are multiplied by the value vectors corresponding to the word segmentations in the query text through the MatMul function to obtain the corresponding target similarity vectors.

[0171] Step 720, perform score prediction according to the third similarity matrix to obtain the matching score between the query text and the text to be matched.

[0172] In some embodiments of the present application, the matching score between the query text and the text to be matched can be calculated according to the third similarity matrix through a fully connected layer (FC). The input of the trained fully connected layer is the third similarity matrix, and the third similarity matrix is transformed through this fully connected layer to output the matching score.

[0173] The meaning expressed by a word in a text is usually related to the context of the word. Therefore, the context information of the word helps to enhance its semantic representation. At the same time, the roles played by different words in the context in semantic representation are often different. Therefore, the attention mechanism can be used to calculate the attention scores corresponding to each word segmentation in the query text. The attention score corresponding to the word segmentation reflects the contribution degree of the word segmentation to the overall semantics of the query text. It can be understood that if the contribution degree of a word segmentation in the query text to the overall semantics of the query text is higher, then in the text matching process, the similarity related to the word segmentation needs to be focused on.

[0174] In this embodiment, by using the attention scores corresponding to each word segmentation in the query text to enhance the matching degree related to the word segmentation in the second similarity matrix, and introducing the factor of the contribution degree of the word segmentation to the overall semantics of the query text to enhance the matching degree corresponding to the word segmentation, the matching accuracy can be further improved.

[0175] Furthermore, through practical analysis, the inventors of this application found that if the neighbor entities of the entities linked to the knowledge graph by a participle are referred to during the process of constructing the word vectors of the participles in the query text, the referred neighbor entities may introduce noise to the text matching. For example, if the participle result of the query text is: air conditioner / water leakage, and a neighbor entity of the entity linked to the knowledge graph by the participle "water leakage" is "pipe repair", where "water leakage" and "pipe repair" are in an extended word relationship. In this case, if a text to be matched is "toilet pipe repair", without considering the semantic contribution degree of each participle in the query text to the query text, during the process of text matching through the model, the model will consider that the query text "air conditioner water leakage" has a high matching degree with the text to be matched "toilet pipe repair", but in fact, the two are irrelevant. The reason is that the input neighbor entity "pipe repair" introduces noise, resulting in an incorrect matching result.

[0176] For this problem, according to the solution of this embodiment, the matching strength of the participles can be weighted based on the attention scores calculated by the attention mechanism, so as to reduce the influence of the noise introduced by the neighbor entities on the text matching result.

[0177] In some embodiments of this application, as Figure 9 shown, before step 210, the method further includes: step 910, receiving a service query request sent by the client, where the service query request indicates a query text; in this embodiment, after step 240, it further includes:

[0178] step 920, obtaining the application information of the target service application, where the target service application refers to the service application corresponding to the target matching text.

[0179] step 930, generating a query result according to the application information.

[0180] step 940, returning the query result to the client, so that the client can display the service entry of the target service application according to the query result.

[0181] Among them, the service query request is initiated for querying service applications. The application information of the determined target service application may include the application name of the target service application, the call interface related to the service that the current user needs to call, etc.

[0182] The query result indicates the call interfaces of the target service applications. Thus, after receiving the query interface, the client can display the application entrances of the target service applications according to the call interfaces of the target service applications in the query result, and the user can enter the corresponding service page by triggering the application entrance. For example, if the user needs to query service applications related to "sending express delivery", after obtaining the query result, the client displays an application entrance for entering the service page providing "sending express delivery", and the user can trigger the application entrance to enter the service page of "sending express delivery".

[0183] In some embodiments of the present application, the process of calculating the matching degree score between the query text and the text to be matched during the text processing can be implemented through a text matching model, and the text matching model can be constructed by a convolutional neural network, a pooling neural network, a fully connected neural network, a recurrent neural network, etc.

[0184] Figure 10 It is a schematic diagram of the text matching model shown according to a specific embodiment of the present application.

[0185] As Figure 10 shown, the text matching model includes an input layer 1010, a similarity matrix layer 1020, a pooling layer (including a first pooling layer 1031 and a second pooling layer 1032), an attention layer 1040, and a fully connected layer 1050. Among them, the input layer 1010 is used to receive the input text. Specifically, in the present application, the input text includes the input text for the query text and the input text for the text to be matched.

[0186] The input text for the query text refers to the text obtained by concatenating the entities linked to each word segment on the knowledge graph in the order of the word segments in the query text after entity linking of the word segments in the query text on the knowledge graph, and also splicing the one-hop neighbor entities corresponding to the entities linked to on the knowledge graph to the corresponding entities.

[0187] The input text for the text to be matched refers to the text obtained by concatenating the entities linked to each word segment on the knowledge graph in the order of the word segments in the text to be matched after entity linking of the word segments in the text to be matched on the knowledge graph, and also splicing the one-hop neighbor entities corresponding to the entities linked to on the knowledge graph to the corresponding entities.

[0188] In the solution of this application, the word segmentation categories to which each entity belongs and the association relationships between entities are set in the knowledge graph. Thus, in the subsequent process, the text matching model can determine the word segmentation categories to which each word segmentation in the query text belongs, the word segmentation categories to which each word segmentation in the text to be matched belongs, and the association relationships between the word segmentations in the query text and the word segmentations in the text to be matched based on the knowledge graph, the input text for the query text, and the input text for the text to be matched.

[0189] The similarity matrix layer 1020 first calculates the similarity matrix of the query text relative to the text to be matched based on the word vectors of each word segmentation in the query text and the word vectors of each word segmentation in the text to be matched; and determines the matching weight matrix of the query text relative to the text to be matched according to the input text for the query text, the input text for the text to be matched, and the knowledge graph; then enhances the similarity matrix through the matching weight matrix to obtain the first similarity matrix. The determination of the similarity matrix and the matching weight matrix refers to the above description and will not be elaborated here. The matching weight matrix can be a matrix obtained by fusing the first weight matrix and the second weight matrix, or the first weight matrix, or the second weight matrix.

[0190] The first pooling layer 1031 and the second pooling layer 1032 in the pooling layer perform pooling processing on the first similarity matrix layer in two directions respectively. Specifically, Figure 10 in it, the first pooling layer 1031 performs pooling by moving the pooling kernel along the horizontal direction of the first similarity matrix to obtain the first pooling matrix; the second pooling layer 1032 performs pooling by moving the pooling kernel along the vertical direction of the first similarity matrix to obtain the second pooling matrix.

[0191] The first pooling matrix and the second pooling matrix obtained by the pooling layer through pooling processing are input into the attention layer 1040. The attention layer performs attention weighting on the first pooling matrix and the second pooling matrix based on the attention mechanism and outputs the second similarity matrix. As described above, the attention layer also needs to calculate the target similarity vector corresponding to each word segmentation by means of the semantic feature vector of the query text and the vectors of each word segmentation in the query text.

[0192] The fully connected layers (FC) 1050 act as a classifier, which is used to perform classification prediction according to the input second similarity matrix and output a matching degree score representing the matching degree between the query text and the text to be matched.

[0193] Next, a specific embodiment will be combined to Figure 10The text matching effect of the shown text matching model is further described. If the word segmentation result of the query text is: provident fund / loan; and the word segmentation result of the text to be matched is: provident fund / query / provident fund / contribution. The similarity of the word segmentation in the query text relative to the word segmentation in the text to be matched calculated based on the word vectors of each word segmentation in the query text and the word vectors of each word segmentation in the text to be matched is shown in Table 1 below. In Table 1, "Query" represents the query text, and "Doc" represents the text to be matched.

[0194] Table 1

[0195]

[0196] Referring to the similarities in Table 1, the similarity matrix of the query text relative to the text to be matched is: . If direct pooling processing is performed on this similarity matrix, where pooling is performed along the vertical direction of this similarity matrix, the resulting pooling matrix is ; and pooling is performed along the horizontal direction of this similarity matrix, the resulting pooling matrix is ; On this basis, the fully connected layer performs classification prediction based on these two pooling matrices and outputs a matching degree score indicating that the query text "provident fund loan" matches the text to be matched "provident fund query provident fund contribution". Obviously, "provident fund loan" and "provident fund query provident fund contribution" actually do not have the same semantics. Therefore, the matching degree score obtained by directly predicting the matching degree score based on the similarity matrix does not match the actual situation, and the matching result indicated by the output matching degree score is incorrect.

[0197] Table 2 shows the first weight associated with the two word segmentation categories indicated in the weight mapping information.

[0198] Table 2

[0199]

[0200] After determining the word segmentation categories to which each word segmentation in the query text belongs and the word segmentation categories to which each word segmentation in the text to be matched belongs, based on the first weight indicated by the weight mapping information, the first weight matrix of the query text relative to the text to be matched can be correspondingly determined.

[0201] Table 3 shows the second weights corresponding to some association relationships.

[0202] Table 3

[0203]

[0204] After determining the association relationship between each word segment in the query text and each word segment in the text to be matched, based on the second weight corresponding to the preset association relationship, the second weight matrix of the query text relative to the text to be matched can be determined.

[0205] After determining the first weight matrix and the second weight matrix of the query text relative to the text to be matched, multiply these three matrices: the first weight matrix, the second weight matrix, and the similarity matrix of the determined query text relative to the text to be matched, to obtain the first similarity matrix. The similarity of each word segment in the enhanced query text relative to the text to be matched indicated by the first similarity matrix is shown in Table 4 below:

[0206] Table 4

[0207]

[0208] The obtained first similarity matrix is: ; Perform pooling along the vertical direction of the first similarity matrix, and the obtained second pooling matrix is: ; Perform pooling along the horizontal direction of the first similarity matrix, and the obtained first pooling matrix is: ; The fully connected layer performs classification prediction based on the first pooling matrix and the second pooling matrix, and outputs a matching degree score indicating that the query text "provident fund loan" does not match the text to be matched "provident fund query provident fund deposit". It can be seen from this that after enhancing the similarity matrix based on the matching weight matrix, the matching degree indicated by the matching degree score predicted according to the enhanced first similarity matrix is consistent with the actual matching situation of "provident fund loan" relative to "provident fund query provident fund deposit". This proves that enhancing the similarity matrix through the matching weight matrix of the query text relative to the text to be matched can improve the accuracy of text matching.

[0209] In some embodiments of the present application, Figure 10 the text matching model shown can be obtained by improving an existing text matching model. By improving the existing text matching model, the improved text matching model can predict the matching degree score between the query text and the text to be matched according to the text matching process in the solution of the present application.

[0210] In some embodiments of the present application, it may be to improve the K-NRM model (Kernel-based Neural Rank Model), to obtain Figure 10The text matching model shown above enables the improved model to perform text matching according to the process in the solution of this application. After improving the K-NRM model, a comparative test was conducted on the original K-NRM model and the improved K-NRM model. Specifically, the original K-NRM model and the improved K-NRM model were respectively used for text matching in the information query scenario, and the query results were determined based on the matching degree scores obtained from the text matching. Based on the query results, the precision, recall, and F1 score (also known as F1 Score) of the query results of the original K-NRM model and the improved K-NRM model were statistically calculated. The comparison of the effects between the improved model and the original K-NRM model is shown in Table 5 below.

[0211] Table 5

[0212]

[0213] Among them, precision = TP / (TP + FP), which represents the proportion of correct entries among the retrieved entries. Recall = TP / (TP + FN), which represents the proportion of retrieved correct entries among all correct entries.

[0214] The F1 score (F1 Score), also known as the balanced F Score, is defined as the harmonic mean of precision and recall. It is an index used in statistics to measure the accuracy of a binary classification model. It takes into account both the accuracy and recall of the classification model. The F1 score can be regarded as a weighted average of the model's accuracy and recall. Its maximum value is 1 and its minimum value is 0. Among them, F1 score = precision * recall / 2(precision + recall).

[0215] Among them, TP (True Positive, true positive example) refers to the number of entries that make a positive determination and the determination is correct. FP (False Positive, false positive example) refers to the number of entries that make a positive determination and the determination is incorrect. FN (False Negative, false negative example) refers to the number of entries that make a negative determination and the determination is incorrect. Specifically, in the scenario of resource query, a positive determination can be that the model determines that a text to be matched is matched with the query text; a negative determination can be that the model determines that a text to be matched is not matched with the query text.

[0216] As can be seen from Table 5 above, after improving the K-NRM model, the improved K-NRM model performs text matching according to the method of this application, and the precision, recall, and F1 score of the text matching are all effectively improved, indicating that the improved K-NRM model performing text matching according to the method of this application effectively improves the accuracy of text matching.

[0217] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.

[0218] Figure 11 is a block diagram of a text processing device shown according to an embodiment, as Figure 11 shown, the text processing device includes: a matching weight matrix determination module 1110, configured to determine a matching weight matrix of a query text relative to a text to be matched, where the matching weight matrix includes at least one of a first weight matrix and a second weight matrix; the first weight matrix is determined according to the word segmentation categories to which the word segments in the query text belong and the word segmentation categories to which the word segments in the text to be matched belong; the second weight matrix is determined according to the association relationship between the word segments in the query text and the word segments in the text to be matched. An enhancement module 1120, configured to enhance a similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix; the similarity matrix is obtained by performing a similarity calculation based on the word vectors of the word segments in the query text and the word vectors of the word segments in the text to be matched. A matching degree score determination module 1130, configured to determine a matching degree score between the query text and the text to be matched according to the first similarity matrix. A target matching text determination module 1140, configured to determine a target matching text according to the matching degree score between the query text and the text to be matched.

[0219] In some embodiments of the present application, the matching weight matrix includes the first weight matrix; the matching weight matrix determination module 1110 includes: a word segmentation category recognition unit, configured to recognize the word segmentation categories to which the word segments in the query text belong; a first weight determination unit, configured to determine a first weight of each word segment in the query text relative to each word segment in the text to be matched according to the word segmentation categories to which the word segments in the query text belong, the word segmentation categories to which the word segments in the text to be matched belong, and weight mapping information; the weight mapping information indicates a first weight associated with any two word segmentation categories; a first weight matrix determination unit, configured to combine the first weights of all the word segments in the query text relative to all the word segments in the text to be matched to obtain the first weight matrix.

[0220] In some embodiments of the present application, the word segmentation category recognition unit includes: a first entity link information acquisition unit, configured to acquire first entity link information, where the first entity link information is obtained by performing entity linking on the word segments in the query text in a knowledge graph; a word segmentation category determination unit, configured to use the word segmentation category to which the first entity linked to the word segment in the query text in the knowledge graph belongs as the word segmentation category to which the word segment in the query text belongs.

[0221] In some other embodiments of the present application, the matching weight matrix includes a second weight matrix; the matching weight matrix determination module 1110 includes: an association relationship recognition unit, configured to recognize the association relationship between the word segments in the query text and the word segments in the text to be matched according to the knowledge graph; a second weight determination unit, configured to perform weight lookup according to the association relationship to obtain the second weight of the word segments in the query text relative to the word segments in the text to be matched; and a second weight matrix determination unit, configured to combine the second weights of each word segment in the query text relative to each word segment in the text to be matched to obtain the second weight matrix.

[0222] In some embodiments of the present application, the association relationship recognition unit includes: a first entity link information acquisition unit, configured to acquire first entity link information, where the first entity link information is used to indicate the first entity linked to by the word segment in the query text on the knowledge graph; a second entity link information acquisition unit, configured to acquire second entity link information, where the second entity link information is used to indicate the second entity linked to by the word segment in the text to be matched on the knowledge graph; and an association relationship determination unit, configured to determine the association relationship between the first entity and the second entity in the knowledge graph as the association relationship between the corresponding word segments in the query text and the corresponding word segments in the text to be matched.

[0223] In some embodiments of the present application, the enhancement module is further configured to: multiply the matching weight matrix by the similarity matrix to obtain a first similarity matrix.

[0224] In some embodiments of the present application, the matching degree score determination module 1130 includes: a pooling processing unit, configured to perform pooling processing on the first similarity matrix to obtain a second similarity matrix; and a matching degree score calculation unit, configured to calculate the matching degree score between the query text and the text to be matched according to the second similarity matrix.

[0225] In some embodiments of the present application, the matching degree score calculation unit includes: an attention weighting unit, configured to perform attention weighting on the second similarity matrix based on the attention mechanism to obtain a third similarity matrix; and a score prediction unit, configured to perform score prediction according to the third similarity matrix to obtain the matching degree score between the query text and the text to be matched.

[0226] In some embodiments of the present application, the attention weighting unit includes: a key matrix determination unit, configured to perform a linear transformation on the word vectors of each word segment in the query text according to a key weight vector to obtain key vectors corresponding to each word segment in the query text; a query matrix determination unit, configured to perform a linear transformation on the semantic feature vector corresponding to the query text according to a query weight vector to obtain a query vector; an attention score determination unit, configured to calculate attention scores corresponding to each word segment in the query text according to the key vectors and the query vector corresponding to each word segment in the query text; a target similarity vector determination unit, configured to weight the value vectors corresponding to each word segment in the query text according to the attention scores corresponding to each word segment in the query text to obtain target similarity vectors corresponding to each word segment in the query text; the value vector corresponding to a word segment in the query text is obtained by performing a linear transformation on the similarity vector corresponding to the word segment in the query text according to a value weight vector, and the similarity vector corresponding to a word segment in the query text is obtained by extracting elements related to the corresponding word segment in the query text from a second similarity matrix and combining the extracted elements; a third similarity matrix determination unit, configured to combine the target similarity vectors corresponding to each word segment in the query text to obtain a third similarity matrix.

[0227] In some embodiments of the present application, the second similarity matrix includes a first pooling matrix and a second pooling matrix; the pooling processing unit includes: a first pooling processing unit, configured to perform pooling processing on the first similarity matrix along the horizontal direction of the first similarity matrix to obtain the first pooling matrix; a second pooling processing unit, configured to perform pooling processing on the first similarity matrix along the vertical direction of the first similarity matrix to obtain the second pooling matrix.

[0228] In some embodiments of the present application, the target matching text determination module includes: a sorting unit, configured to sort a plurality of texts to be matched in descending order of matching degree scores; a target matching text determination unit, configured to determine the texts to be matched at the top of the preset number in the sorting as the target matching texts.

[0229] In some embodiments of the present application, the text processing device further includes: a service query request receiving module, configured to receive a service query request sent by a client, the service query request indicating a query text; and further includes: an application information acquisition module, configured to acquire application information of a target service application, the target service application being the service application corresponding to the target matching text; a query result generation module, configured to generate a query result according to the application information; a query result return module, configured to return the query result to the client, so that the client can display a service entry of the target service application according to the query result.

[0230] Figure 12 The structure diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.

[0231] It should be noted that Figure 12 the computer system 1200 of the illustrated electronic device is only an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.

[0232] As Figure 12 shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the methods in the above embodiments. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0233] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as required so that a computer program read from it can be installed into the storage section 1208 as required.

[0234] In particular, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.

[0235] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0236] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in a flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and the combination of blocks in a block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0237] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.

[0238] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the text processing method in any of the above embodiments is implemented.

[0239] According to one aspect of the present application, there is also provided an electronic device, which includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, the text processing method in any of the above embodiments is implemented.

[0240] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text processing method in any of the above embodiments.

[0241] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0242] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented in software or in a manner of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0243] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0244] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A text processing method, characterized in that, Including: Determine a matching weight matrix of the query text relative to the text to be matched. The matching weight matrix includes at least one of a first weight matrix and a second weight matrix. The first weight matrix is determined according to the word segmentation categories to which the word segments in the query text belong and the word segmentation categories to which the word segments in the text to be matched belong. The second weight matrix is determined according to the association relationship between the word segments in the query text and the word segments in the text to be matched; Enhance the similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix. The similarity matrix of the query text relative to the text to be matched is obtained by performing a similarity calculation based on the word vectors of the word segments in the query text and the word vectors of the word segments in the text to be matched; Determine a matching degree score between the query text and the text to be matched according to the first similarity matrix; Determine a target matching text according to the matching degree score between the query text and the text to be matched; Wherein, the first weight matrix is determined according to the following process: identify the word segmentation categories to which the word segments in the query text belong; determine the first weight of each word segment in the query text relative to each word segment in the text to be matched according to the word segmentation categories to which the word segments in the query text belong, the word segmentation categories to which the word segments in the text to be matched belong, and weight mapping information. The weight mapping information indicates the first weight associated with any two word segmentation categories; combine the first weights of all the word segments in the query text relative to all the word segments in the text to be matched to obtain the first weight matrix; The second weight matrix is determined according to the following process: identify the association relationship between the word segments in the query text and the word segments in the text to be matched according to the knowledge graph; perform a weight search according to the association relationship to obtain the second weight of the word segments in the query text relative to the word segments in the text to be matched; combine the second weights of the word segments in the query text relative to the word segments in the text to be matched to obtain the second weight matrix.

2. The method according to claim 1, characterized in that The identifying the word segmentation categories to which the word segments in the query text belong includes: Obtain first entity link information, where the first entity link information is obtained by performing entity linking on the word segments in the query text in the knowledge graph; Use the word segmentation category to which the first entity linked to by the word segments in the query text in the knowledge graph belongs as the word segmentation category to which the word segments in the query text belong.

3. The method according to claim 1, wherein The identifying the association relationship between the word segments in the query text and the word segments in the text to be matched according to the knowledge graph includes: Obtain first entity link information, where the first entity link information is used to indicate the first entity linked to by the word segments in the query text on the knowledge graph; Obtain second entity link information, where the second entity link information is used to indicate the second entity linked to by the word segments in the text to be matched on the knowledge graph; Determine the association relationship between the first entity and the second entity in the knowledge graph as the association relationship between the corresponding word segments in the query text and the corresponding word segments in the text to be matched.

4. The method according to claim 1, wherein The enhancing the similarity matrix of the query text with respect to the text to be matched according to the matching weight matrix to obtain a first similarity matrix includes: Multiplying the matching weight matrix by the similarity matrix to obtain the first similarity matrix.

5. The method according to claim 1, characterized in that, The determining the matching degree score between the query text and the text to be matched according to the first similarity matrix includes: Performing a pooling process on the first similarity matrix to obtain a second similarity matrix; Calculating the matching degree score between the query text and the text to be matched according to the second similarity matrix.

6. The method according to claim 5, wherein The calculating the matching degree score between the query text and the text to be matched according to the second similarity matrix includes: Performing attention weighting on the second similarity matrix based on an attention mechanism to obtain a third similarity matrix; Performing score prediction according to the third similarity matrix to obtain the matching degree score between the query text and the text to be matched.

7. The method according to claim 6, wherein The performing attention weighting on the second similarity matrix based on an attention mechanism to obtain a third similarity matrix includes: Performing a linear transformation on the word vectors of each word segment in the query text according to a key weight vector to obtain key vectors corresponding to each word segment in the query text; Performing a linear transformation on the semantic feature vector corresponding to the query text according to a query weight vector to obtain a query vector; Calculating attention scores corresponding to each word segment in the query text according to the key vectors corresponding to each word segment in the query text and the query vector; Weighting the value vectors corresponding to each word segment in the query text according to the attention scores corresponding to each word segment in the query text respectively to obtain target similarity vectors corresponding to each word segment in the query text; the value vectors corresponding to the word segments in the query text are obtained by performing a linear transformation on the similarity vectors corresponding to the word segments in the query text according to a value weight vector, and the similarity vectors corresponding to the word segments in the query text are obtained by extracting elements related to the corresponding word segments in the query text from the second similarity matrix and combining the extracted elements; Combining the target similarity vectors corresponding to each word segment in the query text to obtain the third similarity matrix.

8. The method according to claim 5, wherein The second similarity matrix includes a first pooling matrix and a second pooling matrix; The performing a pooling process on the first similarity matrix to obtain a second similarity matrix includes: Performing a pooling process on the first similarity matrix along the horizontal direction of the first similarity matrix to obtain a first pooling matrix; Performing a pooling process on the first similarity matrix along the vertical direction of the first similarity matrix to obtain a second pooling matrix.

9. The method according to claim 1, wherein The determining the target matching text according to the matching degree score between the query text and the text to be matched includes: Sorting multiple texts to be matched in descending order of the matching degree score; Determining the texts to be matched in the top set number in the sorting as the target matching texts.

10. The method according to claim 1, characterized in that Before determining the matching weight matrix of the query text relative to the text to be matched, the method further includes: Receiving a service query request sent by a client, where the service query request indicates the query text; After determining the target matching text according to the matching degree score between the query text and the text to be matched, the method further includes: Obtaining application information of a target service application, where the target service application refers to the service application corresponding to the target matching text; Generating a query result according to the application information; Returning the query result to the client so that the client can display a service entry of the target service application according to the query result.

11. A text processing device, characterized in that, including: A matching weight matrix determination module, configured to determine a matching weight matrix of the query text relative to the text to be matched, where the matching weight matrix includes at least one of a first weight matrix and a second weight matrix; the first weight matrix is determined according to the word segmentation categories to which the word segments in the query text belong and the word segmentation categories to which the word segments in the text to be matched belong; the second weight matrix is determined according to the association relationship between the word segments in the query text and the word segments in the text to be matched; An enhancement module, configured to enhance the similarity matrix of the query text relative to the text to be matched according to the matching weight matrix to obtain a first similarity matrix; The similarity matrix is obtained by performing a similarity calculation based on the word vectors of the word segments in the query text and the word vectors of the word segments in the text to be matched; A matching degree score determination module, configured to determine the matching degree score between the query text and the text to be matched according to the first similarity matrix; A target matching text determination module, configured to determine a target matching text according to the matching degree score between the query text and the text to be matched; Among them, the first weight matrix is determined according to the following process: identifying the word segmentation categories to which the word segments in the query text belong; determining the first weight of each word segment in the query text relative to each word segment in the text to be matched according to the word segmentation categories to which the word segments in the query text belong, the word segmentation categories to which the word segments in the text to be matched belong, and weight mapping information; the weight mapping information indicates the first weight associated with any two word segmentation categories; combining the first weights of all the word segments in the query text relative to all the word segments in the text to be matched to obtain the first weight matrix; The second weight matrix is determined according to the following process: identifying the association relationship between the word segments in the query text and the word segments in the text to be matched according to a knowledge graph; performing weight lookup according to the association relationship to obtain the second weight of the word segments in the query text relative to the word segments in the text to be matched; combining the second weights of the word segments in the query text relative to the word segments in the text to be matched to obtain the second weight matrix.

12. The device according to claim 11, wherein The matching weight matrix determination module is further configured to: Obtain first entity link information, where the first entity link information is obtained by performing entity linking on the word segments in the query text in a knowledge graph; The word segmentation category to which the word segmentation in the query text is linked to the first entity in the knowledge graph is used as the word segmentation category to which the word segmentation in the query text belongs.

13. The apparatus according to claim 11, wherein The matching weight matrix determination module is further configured to: obtain first entity link information, where the first entity link information is used to indicate the first entity to which the word segmentation in the query text is linked on the knowledge graph; obtain second entity link information, where the second entity link information is used to indicate the second entity to which the word segmentation in the text to be matched is linked on the knowledge graph; Determine the association relationship between the first entity and the second entity in the knowledge graph as the association relationship between the corresponding word segmentations in the query text and the corresponding word segmentations in the text to be matched.

14. The device according to claim 11, characterized in that, The enhancement module is further configured to: multiply the matching weight matrix by the similarity matrix to obtain the first similarity matrix.

15. The device according to claim 11, characterized in that, The matching degree score determination module includes: A pooling processing unit, configured to perform pooling processing on the first similarity matrix to obtain a second similarity matrix; A matching degree score calculation unit, configured to calculate the matching degree score between the query text and the text to be matched according to the second similarity matrix.

16. The device according to claim 15, wherein, The matching degree score calculation unit includes: An attention weighting unit, configured to perform attention weighting on the second similarity matrix based on the attention mechanism to obtain a third similarity matrix; A score prediction unit, configured to perform score prediction according to the third similarity matrix to obtain the matching degree score between the query text and the text to be matched.

17. The device according to claim 16, characterized in that, The attention weighting unit includes: A key matrix determination unit, configured to perform a linear transformation on the word vectors of each word segmentation in the query text according to the key weight vector to obtain the key vectors corresponding to each word segmentation in the query text; A query matrix determination unit, configured to perform a linear transformation on the semantic feature vector corresponding to the query text according to the query weight vector to obtain a query vector; An attention score determination unit, configured to calculate the attention scores corresponding to each word segmentation in the query text according to the key vectors corresponding to each word segmentation in the query text and the query vector; A target similarity vector determination unit, configured to weight the value vectors corresponding to each word segmentation in the query text according to the attention scores corresponding to each word segmentation in the query text to obtain the target similarity vectors corresponding to each word segmentation in the query text; the value vectors corresponding to the word segmentations in the query text are obtained by performing a linear transformation on the similarity vectors corresponding to the word segmentations in the query text according to the value weight vector, and the similarity vectors corresponding to the word segmentations in the query text are obtained by extracting the elements related to the corresponding word segmentations in the query text from the second similarity matrix and combining the extracted elements; a third similarity matrix determination unit, configured to combine the target similarity vectors corresponding to each word segmentation in the query text to obtain the third similarity matrix.

18. The device according to claim 15, characterized in that, The second similarity matrix includes a first pooling matrix and a second pooling matrix; the pooling processing unit includes: The first pooling processing unit is configured to perform pooling processing on the first similarity matrix along the horizontal direction of the first similarity matrix to obtain a first pooling matrix; The second pooling processing unit is configured to perform pooling processing on the first similarity matrix along the vertical direction of the first similarity matrix to obtain a second pooling matrix.

19. The device according to claim 11, wherein The target matching text determination module includes: The sorting unit is configured to sort multiple texts to be matched in descending order of the matching degree score; The target matching text determination unit is configured to determine the texts to be matched at the top of the sorting with a preset number as the target matching texts.

20. The device according to claim 11, wherein, The text processing device further includes: The service query request receiving module is configured to receive a service query request sent by a client, and the service query request indicates the query text; The application information obtaining module is configured to obtain application information of a target service application, where the target service application refers to the service application corresponding to the target matching text; The query result generating module is configured to generate a query result according to the application information; the query result returning module is configured to return the query result to the client, so that the client displays a service entry of the target service application according to the query result.

21. An electronic device, characterized in that, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1-10 is implemented.

22. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, the method according to any one of claims 1-10 is implemented.

23. A computer program product, characterized in that, including a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-10 is implemented.

Citation Information

Patent Citations

  • Data processing method and device and computer readable storage medium

    CN110110035A

  • Short text similarity calculation method and device and readable storage medium

    CN110929498A

  • Service discovery method based on convolutional neural network under BERT

    CN110941698A