Medical title matching method, device and equipment and storage medium

By determining the vectors of medical titles and search terms, and combining feature similarity and intent matching, the problem of inaccurate title matching in medical searches is solved using the BERT model and a similarity recognition model, thus achieving more accurate medical content retrieval.

CN113569124BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-01-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In the field of medical search, existing technologies cannot accurately determine the matching degree between medical search terms and medical titles, resulting in users being unable to accurately find the medical content they need.

Method used

By determining the title vector of medical titles and the statement vector of medical search queries, and combining feature similarity and intent matching results, the BERT model and similarity recognition model are used to analyze the degree of matching between medical titles and search queries, thereby achieving more accurate matching.

Benefits of technology

It improves the accuracy of matching medical titles with search terms, ensuring that users can retrieve the medical content they need more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569124B_ABST
    Figure CN113569124B_ABST
Patent Text Reader

Abstract

The application provides a medical title matching method and device, equipment and a storage medium. The application relates to the application of artificial intelligence technology. After obtaining a medical search statement, the application determines the feature similarity in semantics between a medical title and the medical search statement, and determines whether the medical intention between the medical title and the medical search statement is the same, for each medical title, in combination with the title vector of the medical title and the statement vector of the medical search statement. On this basis, the matching degree between the medical title and the medical search statement can be more comprehensively analyzed in combination with the feature similarity in semantics between the medical title and the medical search statement and the intention matching result of the medical intention, so that the medical title that is more matched with the medical search statement can be more accurately determined from multiple medical titles, and then the medical text content pointed to by the medical title can be more accurately searched based on the medical search statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of search technology, and in particular to a medical title matching method, apparatus, device, and storage medium. Background Technology

[0002] With the development of internet-based healthcare, users can search for medical knowledge through browsers or medical dictionary applications. For example, users can use medical encyclopedias to find more professional and authoritative medical information.

[0003] In a medical content search scenario, each piece of medical content has a title. Based on this, after receiving the user's medical-related search query, the search engine matches the search query with the titles of each piece of medical content and retrieves at least one medical document whose title matches the search query.

[0004] However, in the field of medical search, it is not possible to accurately determine the matching degree between medical search terms and medical titles, making it impossible for users to accurately find the medical content they need through medical encyclopedias. Summary of the Invention

[0005] In view of this, this application provides a medical title matching method, apparatus, device and storage medium to achieve more accurate matching of medical content represented by medical titles using medical search statements.

[0006] To achieve the above objectives, this application provides the following technical solution:

[0007] On the one hand, this application provides a medical title matching method, including:

[0008] Obtain medical search terms;

[0009] For each of the multiple medical titles to be matched, determine the title vector of the medical title;

[0010] Determine the statement vector of the medical search statement;

[0011] For each medical title, the feature similarity between the medical title and the medical search statement is determined based on the title vector of the medical title and the statement vector of the medical search statement.

[0012] For each medical title, based on the title vector of the medical title and the statement vector of the medical search statement, an intent recognition model is used to determine the intent matching result between the medical title and the medical search statement. The intent matching result is used to characterize whether the medical intent between the medical title and the medical search statement is the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled by themselves, and using the vectors of the medical title samples and medical search statement samples in each first sample pair.

[0013] By combining the feature similarity and intent matching results between each medical title and the medical search statement, the matching degree ranking of the multiple medical titles is determined.

[0014] In one possible scenario, prior to determining the feature similarity between the medical title and the medical search query, the method further includes:

[0015] The difference vector is obtained by determining the vector difference between the statement vector of the medical search query and the title vector of the medical title;

[0016] The determination of feature similarity between the medical title and the medical search query based on the title vector of the medical title and the query vector of the medical search query includes:

[0017] Based on the title vector of the medical title, the statement vector of the medical search query, and the difference vector, the feature similarity between the medical title and the medical search query is determined.

[0018] In another possible scenario, determining the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector includes:

[0019] Based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, the feature similarity between the medical title and the medical search statement is determined using a similarity recognition model.

[0020] The similarity recognition model is trained based on the feature similarity of multiple second sample pairs labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within the second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample.

[0021] In another possible scenario, before determining the feature similarity and intent matching results between the medical title and the medical search statement, the process further includes:

[0022] The dimensionality reduction of the medical search statement vector is performed by a vector dimensionality reduction model, which is obtained by training the similarity recognition model using the title vector of the medical title sample and the statement vector of the medical search statement sample in the second sample pair.

[0023] The dimensionality of the medical title vector is reduced using the vector dimensionality reduction model described above.

[0024] The step of determining the vector difference between the statement vector of the medical search query and the title vector of the medical title to obtain the difference vector includes:

[0025] The difference vector is obtained by determining the vector difference between the statement vector of the medical search statement after dimensionality reduction and the title vector of the medical title after dimensionality reduction.

[0026] In yet another possible scenario, determining the title vector of the medical title includes:

[0027] The title vector of the medical title is determined using a vector transformation model;

[0028] The process of determining the statement vector of the medical search statement includes:

[0029] The vector transformation model is used to determine the statement vector of the medical search statement;

[0030] The vector transformation model is a bidirectional encoding representation BERT model based on a transformer, and the vector transformation model is trained by using masked word sequences corresponding to multiple medical corpus samples and predicting the masked words in the masked word sequences as the training target.

[0031] The medical corpus sample consists of a medical title sample and the medical text content represented by the medical title sample, and the masked word sequence is a word sequence obtained by masking at least one word contained in the medical corpus sample.

[0032] In another aspect, this application also provides a medical title matching device, comprising:

[0033] The statement acquisition unit is used to obtain medical search statements;

[0034] The first vector determination unit is used to determine the title vector of each medical title among multiple medical titles to be matched;

[0035] The second vector determination unit is used to determine the statement vector of the medical search statement;

[0036] The feature determination unit is used to determine the feature similarity between the medical title and the medical search statement for each medical title, based on the title vector of the medical title and the statement vector of the medical search statement;

[0037] An intent determination unit is used to determine the intent matching result between the medical title and the medical search statement for each medical title based on the title vector of the medical title and the statement vector of the medical search statement, and using an intent recognition model. The intent matching result is used to characterize whether the medical intent between the medical title and the medical search statement is the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled by themselves, and using the vectors of the medical title samples and medical search statement samples in each first sample pair.

[0038] The matching determination unit is used to determine the matching degree ranking of the multiple medical titles by combining the feature similarity and intent matching results between each medical title and the medical search statement.

[0039] One possible implementation also includes:

[0040] The difference determination unit is used to determine the vector difference between the statement vector of the medical search statement and the title vector of the medical title before the feature determination unit determines the feature similarity between the medical title and the medical search statement, and obtain the difference vector.

[0041] The feature determination unit is specifically used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector.

[0042] In yet another possible implementation, the feature determination unit includes:

[0043] The feature determination subunit is used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, and using a similarity recognition model.

[0044] The similarity recognition model is trained based on the feature similarity of multiple second sample pairs labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within the second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample.

[0045] In another aspect, this application also provides a server, including a memory and a processor;

[0046] The memory is used to store programs;

[0047] The processor is used to execute the program, which, when executed, is specifically used to implement the medical title matching method as described above.

[0048] In another aspect, this application also provides a storage medium for storing a program, which, when executed, implements the medical title matching method as described in any of the above claims.

[0049] As described above, after obtaining the medical search query, this application, for each medical title, not only combines the title vector of the medical title and the statement vector of the medical search query to determine the semantic feature similarity between the medical title and the medical search query, but also determines whether the medical intents of the medical title and the medical search query are the same. Based on this, by combining the semantic feature similarity and intent matching results between the medical title and the medical search query, a more comprehensive analysis of the matching degree between the medical title and the medical search query can be performed. This facilitates a more accurate identification of the medical title that best matches the medical search query from multiple medical titles, and further enables a more accurate search for the medical text content pointed to by the medical title based on the medical search query. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0051] Figure 1 A schematic diagram of a system architecture to which the solution of this application applies is shown;

[0052] Figure 2 This invention provides a schematic flowchart of one embodiment of a medical title matching method.

[0053] Figure 3 This paper illustrates a flowchart of yet another embodiment of a medical title matching method according to this application;

[0054] Figure 4 This diagram illustrates a schematic representation of the implementation principle framework of a medical title matching method according to this application.

[0055] Figure 5 This paper presents a schematic diagram illustrating the principle framework for training a BERT model according to this application.

[0056] Figure 6This diagram illustrates an implementation flow of training a similarity recognition model and an intent recognition model according to this application.

[0057] Figure 7 A schematic diagram illustrating an application scenario to which the medical title matching method of this application is applicable is shown;

[0058] Figure 8 This is a schematic diagram of an interface on which the terminal displays matched medical encyclopedia articles;

[0059] Figure 9 This invention provides a schematic diagram illustrating the structural composition of one embodiment of a medical title matching device.

[0060] Figure 10 A schematic diagram of the structural composition of an embodiment of an electronic device according to this application is shown. Detailed Implementation

[0061] The solution proposed in this application is applicable to any medical content search scenario. In this scenario, the medical titles of various medical content items can be matched based on the medical search query, and at least one medical content item under a medical title that matches the medical search query can be identified.

[0062] To facilitate understanding, the medical search system to which the solution in this application is applicable will be introduced first.

[0063] like Figure 1 As shown, it illustrates a schematic diagram of the composition architecture of a medical search system to which this application applies.

[0064] The medical search system may include: a medical search platform 100 and a terminal 200.

[0065] The medical search platform can store multiple pieces of medical content, each with a medical title. Since each piece of medical content only deals with one disease, the medical title represents the disease topic covered by the medical text.

[0066] Each piece of medical content can include information on disease symptoms, causes, diagnosis and treatment, or health care.

[0067] Medical content can take many forms. For example, it can be medical text content, such as articles or short texts introducing medical information. For instance, medical text content and medical titles could be a medical question and its answer, respectively.

[0068] Of course, the medical content can also be medical video content, etc., and there are no restrictions on this.

[0069] The terminal 200 can access the medical search platform through a browser or a medical search application that matches the medical search platform, and send a search request to the medical search platform. The search request can carry medical search terms.

[0070] Accordingly, the medical search platform 100 may include at least one server 101.

[0071] The server can match the medical titles of multiple medical contents in the medical retrieval platform according to the medical search statement sent by the terminal, and search for at least one medical title with a high degree of matching with the medical search statement, so as to return the medical content pointed to by the at least one searched medical title to the terminal.

[0072] For example, a medical search platform can be a medical encyclopedia platform. Based on this, the terminal can request various medical encyclopedia knowledge and related information such as disease symptoms from the medical encyclopedia platform.

[0073] It is understandable that medical search platforms can store medical content and its titles on the aforementioned servers, or they can set up a database within the medical search platform itself. Figure 1 (Not shown in the image), and stores multiple medical contents and their associated medical titles in a database without restriction.

[0074] In this application, the server of the medical search platform can combine artificial intelligence technology to perform medical title matching and related processing.

[0075] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0076] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0077] In order to match medical search statements with medical titles, the medical search platform in this application will involve at least several artificial intelligence technologies such as natural language processing and machine learning.

[0078] Natural Language Processing (NLP) is an important area within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. NLP technologies typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0079] For example, in order to train vector models and other models, this application may involve text processing such as word segmentation of medical title samples and search statement samples, and may also involve semantic understanding of medical titles and medical search statements, etc.

[0080] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0081] The following section, using a flowchart, explains the medical title matching method of this application and the artificial intelligence and other technologies involved in the medical title matching method.

[0082] like Figure 2 The diagram illustrates a flowchart of a medical title matching method according to this application. The method of this embodiment can be applied to the server mentioned above, and may include:

[0083] S201, Obtain medical search terms.

[0084] For example, the server can receive a search request sent by a terminal, which may carry a medical search query. This search request is used to search for at least one piece of medical content that matches the medical search query. Accordingly, the server can obtain the medical search query carried in the search request.

[0085] The medical search statement is a search statement (also known as a query statement) related to the requested medical content. For ease of distinction, this search statement is referred to as a medical search statement.

[0086] The medical search query can be a string containing at least one character. For example, the medical search query can be words or sentences such as "how to care for children with anorexia" or "fever".

[0087] S202, for each of the multiple medical titles to be matched, determine the title vector of the medical title.

[0088] It is understandable that a medical retrieval platform can store multiple medical contents, each with a medical title. Therefore, the platform will store multiple medical titles. To determine the medical content that matches the medical search query, the degree of matching between the search query and the medical title needs to be determined. Therefore, this application specifies the relevant operations in steps S202 to S205 for each medical title.

[0089] Of course, in practical applications, you can also first filter the medical titles in the medical search platform based on the keywords contained in the medical query, and select the medical titles that may match the medical query, and use the results as the medical titles to be matched.

[0090] For ease of distinction, the vector derived from medical titles is called a title vector. The title vector represents the semantic features of the medical title.

[0091] Understandably, there are multiple ways to determine the title vector for a medical headline. For example, existing word vector models can be used to determine the title vector.

[0092] In one possible implementation, the title vector of the medical title can be determined using a pre-trained vector transformation model.

[0093] Understandably, considering that general word vector models may not be able to accurately determine the vectors of medical texts in the medical field, this application can also fine-tune existing word vector models or encoder models in advance using medical corpus samples to train a vector transformation model suitable for determining the vectors of medical texts in the medical field.

[0094] The medical corpus sample can consist of medical title samples and the medical text content represented by the medical title samples. Of course, the medical corpus sample can also be other medical samples related to medical content, and there are no restrictions on this.

[0095] For example, a word2vec model can be trained using medical corpus samples to obtain a word2vec model suitable for determining word vectors of words related to the medical field.

[0096] For example, this vector transformation model can be a trained Bidirectional Encoder Representations from Transformers (BERT) model. Correspondingly, this vector transformation model can be trained using masked word sequences corresponding to multiple medical corpus samples, with the training objective being to predict the masked words in these sequences. Here, the masked word sequence corresponding to the medical corpus samples is the word sequence obtained after masking at least one word contained in the medical corpus sample.

[0097] Masking of medical corpus samples containing at least one character refers to using set masking rules to replace or change some or all of the words in the medical corpus samples so that the at least one word in the masked medical corpus samples changes.

[0098] For example, assuming a medical corpus sample can be segmented into 100 words, 80 of the 100 words can be kept unchanged, while of the remaining 20 words, 80% can be replaced with mask identifiers, 10% can be replaced with other characters, and 10% can remain unchanged. Of course, the same applies if each character in the medical corpus sample is identified as a word.

[0099] Understandably, because the BERT model employs a multi-layered transformer to learn bidirectionally from the text (such as the medical title in this application), it can more accurately learn the contextual relationships between words in the text (e.g., the words in the medical title), thereby enabling the accurate extraction of semantic features from each word in the text. Therefore, the trained BERT model can more accurately extract semantic features reflecting medical meaning from medical titles.

[0100] S203, Determine the statement vector of the medical search statement.

[0101] For ease of distinction, the vector derived from a medical search statement is called a statement vector. The statement vector of a medical search statement is used to represent the semantic features of that medical search statement.

[0102] Similar to step S202, there are multiple ways to determine the statement vector. For example, a general word vector model can be used to determine the statement vector of a medical search statement.

[0103] Alternatively, a vector transformation model can be pre-trained using medical corpus text. Correspondingly, this vector transformation model can be used to determine the statement vectors of medical search queries. This vector transformation model can be the same as the vector transformation model in step S202, such as a trained BERT model.

[0104] It is understandable that the order of steps S202 and S203 can be interchanged or performed simultaneously, without any restrictions.

[0105] S204. For each medical title, determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title and the statement vector of the medical search statement.

[0106] The feature similarity is used to characterize the semantic similarity between medical titles and medical search statements.

[0107] It is understandable that the feature similarity can be a similarity score or a similarity level that represents the degree of similarity. For example, five similarity levels can be set from completely identical to completely different, and different similarity levels can represent different degrees of feature similarity.

[0108] There are several ways to determine the similarity of this feature.

[0109] For example, in one possible scenario, the cosine similarity between the title vector of the medical title and the statement vector of the medical search query can be calculated, and the calculated cosine similarity can be determined as the feature similarity between the medical title and the medical search query.

[0110] In another possible scenario, this application can pre-train a similarity recognition model. This model can be trained based on the feature similarity of multiple sample pairs labeled with their respective features, and using the vectors corresponding to the medical title sample and the medical search statement sample within each sample pair. Each sample pair includes one medical title sample and one medical search statement sample, where the medical title sample is the medical title used as the training sample, and the medical search statement sample is the medical search statement used as the training sample.

[0111] For example, the similarity recognition model can be a classification model such as normalized softmax. Based on this, the similarity category corresponding to each pair of samples is pre-labeled. For instance, the similarity categories can include five categories: similarity category 1 to similarity category 5, where these five categories represent, in order: completely identical, similarity greater than 80%, similarity greater than 50% but less than 80%, similarity less than 50%, and completely different. Based on this, the title vectors of the medical title samples and the statement vectors of the medical search statement samples from each of the multiple sample pairs labeled with similarity categories are used to train a classification model such as softmax. The specific training process is not restricted.

[0112] In this case, the feature similarity between medical titles and medical search statements can be determined based on the title vector of the medical title and the statement vector of the medical search statement, using this similarity recognition model.

[0113] S205, for each medical title, based on the title vector of the medical title and the statement vector of the medical search statement, and using an intent recognition model, determine the intent matching result between the medical title and the medical search statement.

[0114] The intent matching result is used to characterize whether the medical intent is the same between the medical title and the medical search statement.

[0115] For example, intent matching results can be divided into two categories: same intent and different intent. Based on this, for each medical title, the intent matching result between the medical title and the medical search statement can be the same intent or different intent.

[0116] Medical intent can represent the direction of medical knowledge that is desired or requested. Specifically, the medical intent of a medical search query can determine the category of medical knowledge expected to be obtained based on that search query; while the medical intent of a medical title can represent the category of medical knowledge reflected in the medical content pointed to by that title.

[0117] For example, medical intent can be categorized into several types, such as symptoms, causes, medical treatment, medications, treatment, and prevention. For instance, if the medical intent of a medical search query is "symptoms," it means that the search aims to find medical information related to the symptoms of a disease.

[0118] In this application, the trained intent recognition model can be used to analyze whether the medical intent of the medical search statement and the medical title are the same, thereby obtaining the intent recognition result.

[0119] The intent recognition model is trained based on the intent matching results of multiple sample pairs labeled with their respective intents, and using the vectors of medical title samples and medical search statement samples within each sample pair.

[0120] Each sample pair can include a medical title sample and a medical search statement sample, and each sample pair is labeled with an intent matching result. Based on this, the title vector of the medical title sample and the statement vector of the medical search statement in each sample pair can be sequentially input into the trained intent recognition model. The intent recognition result predicted by the intent recognition model for each sample pair is compared with the actual labeled intent recognition result of that sample pair to indicate that the prediction accuracy of the intent recognition model meets the requirements.

[0121] The sample pairs used to train the intent recognition model can be the same as or different from the sample pairs used to train the similarity recognition model; there are no restrictions on this.

[0122] For ease of distinction, the sample pairs used to train the intent recognition model are referred to as the first sample pair in the claims of this application, while the sample pairs used to train the similarity recognition model are referred to as the second sample pair. Of course, the first sample pair and the second sample pair are merely for distinguishing sample pairs used to train different models and have no other meaning. In subsequent embodiments, the sample pairs used to train the similarity recognition model may also be referred to as the first sample pair, while the sample pairs used to train the intent recognition model may be referred to as the second sample pair, as needed.

[0123] S206. Based on the feature similarity and intent matching results between each medical title and the medical search statement, determine the matching degree ranking of multiple medical titles.

[0124] For example, medical titles with higher feature similarity to medical search statements are ranked higher. If the feature similarity is the same, then based on the intent matching results of the medical titles, medical titles with the same intent as the medical search statements are ranked higher.

[0125] It is understandable that this embodiment analyzes the matching degree between medical titles and medical search statements from two dimensions: feature similarity and the similarity of medical intent. Since medical intent reflects the type of medical knowledge requested by the search statement, and the medical intent of a medical title reflects the type of medical content it points to, determining the matching degree ranking of medical titles by combining feature similarity with intent matching, based on determining the feature similarity between medical titles and search statements, allows medical titles that better match the medical content requested by the search statement to be ranked higher, thus facilitating more accurate retrieval of medical content.

[0126] Understandably, in practical applications, after determining the matching degree ranking of multiple medical titles, the medical content corresponding to at least one of the top-ranked medical titles can be returned to the terminal according to the matching degree ranking of these multiple medical titles.

[0127] It is understood that this application can determine the degree of matching between the medical title and the medical search statement by combining the feature similarity and intent matching results between the medical title and the medical search statement.

[0128] For example, a first weight corresponding to feature similarity and a second weight for intent matching results can be set. The sum of the first weight and the second weight can be zero. Based on this, for each medical title, if the intent matching results of the medical title and the medical search statement determine that the intent of the medical title and the medical search statement are the same, then the value of the intent matching result is set to 1; correspondingly, if the intent of the medical title and the medical search statement are not the same, the intent matching result is set to zero.

[0129] Accordingly, given the feature similarity and intent matching results between the medical title and the medical search query, the degree of matching between the medical title and the medical search query can be determined as follows:

[0130] Calculate the first product of the feature similarity between the medical title and the medical search query and the first weight;

[0131] Calculate the value of the intent match between the medical title and the medical search statement, and the second product of the second weight;

[0132] The sum of the first and second products is used to determine the degree of match between the medical title and the medical search query.

[0133] After obtaining the match degree between the medical titles and the medical search query, the match degree of each medical title can be ranked. Alternatively, the system can directly return at least one medical content with a high match degree to the terminal based on the match degree of each medical title.

[0134] As can be seen, after obtaining the medical search query, this application, for each medical title, not only combines the title vector of the medical title and the statement vector of the medical search query to determine the semantic feature similarity between the medical title and the medical search query, but also determines whether the medical intents of the medical title and the medical search query are the same. Based on this, by combining the semantic feature similarity and intent matching results between the medical title and the medical search query, a more comprehensive analysis of the matching degree between the medical title and the medical search query can be performed. This facilitates a more accurate identification of the medical title that best matches the medical search query from multiple medical titles, and further enables a more accurate search for the medical text content pointed to by the medical title based on the medical search query.

[0135] It is understandable that the vector difference between the statement vector of a medical search query and the title vector of a medical title can also reflect the difference in semantic features between the medical title and the medical search query, and thus reflect the degree of similarity between the two in terms of semantic features on one dimension.

[0136] Based on this, before determining the feature similarity between medical titles and medical search statements, this application can also determine the vector difference between the statement vector of the medical search statement and the title vector of the medical title, obtaining a difference vector. Accordingly, for each medical title, this application can...

[0137] In one alternative approach, for each medical title, the feature similarity between the medical title and the medical search query can be determined using a similarity recognition model based on the title vector of the medical title, the query vector of the medical search query, and the difference vector.

[0138] In this case, the similarity recognition model can be trained based on the feature similarity of multiple sample pairs labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within each second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample within that sample pair.

[0139] For example, a softmax model trained using multiple sample pairs labeled with feature similarity can be used as the similarity recognition model.

[0140] It is understood that, in the above embodiments of this application, considering that the vectors converted from medical search statements and medical title statements have high dimensionality, this application can also reduce the dimensionality of the statement vectors of medical search statements and the title vectors of medical titles, and then determine the feature similarity and intent matching results mentioned above based on the dimensionality-reduced statement vectors and title vectors.

[0141] The following explanation uses one implementation method as an example, such as Figure 3The diagram illustrates a flowchart of yet another embodiment of a medical title matching method according to this application. The method of this embodiment may include:

[0142] S301, Obtain medical search terms.

[0143] S302, use the BERT model to determine the statement vector of the medical search statement.

[0144] S303: For each medical title among multiple medical titles to be matched, the BERT model is used to determine the title vector of that medical title.

[0145] The BERT model is trained using masked word sequences corresponding to multiple medical corpus samples, with the training objective being to predict the masked words in the masked word sequences. The masked word sequence is a sequence of words obtained by masking at least one word in the medical corpus samples.

[0146] It should be noted that this embodiment is for ease of understanding, using the BERT model trained using medical corpus samples as an example for illustration. However, it is understood that this embodiment is also applicable to other vector conversion models or other methods for determining the sentence vector and title vector.

[0147] S304. The dimension reduction of the statement vector of the medical search statement is performed using a vector dimension reduction model to obtain the dimension-reduced statement vector.

[0148] Specifically, the vector dimensionality reduction model is obtained by training the similarity recognition model using the statement vectors of medical search statement samples and the title vectors of medical title samples from each sample pair used in training the similarity recognition model. In other words, the vector dimensionality reduction model can be trained together with the similarity recognition model.

[0149] For example, in one alternative approach, the vector dimensionality reduction model can be a pooling model.

[0150] S305. For each medical title, the title vector of the medical title is reduced in dimensionality using this vector dimensionality reduction model to obtain the dimensionality-reduced title vector.

[0151] S306, for each medical title, determine the vector difference between the dimension-reduced sentence vector and the dimension-reduced title vector corresponding to that medical title, and obtain the difference vector.

[0152] S307, For each medical title, input the dimensionality-reduced title vector, the dimensionality-reduced sentence vector, and the difference vector into the trained similarity recognition model to obtain the feature similarity between the medical title and the medical search sentence output by the similarity recognition model.

[0153] For example, for each medical title, the dimensionality-reduced title vector, the dimensionality-reduced sentence vector, and the difference vector between the title vector and the sentence vector of the medical title can be reconstructed into a single vector. The reconstructed vector is then input into the similarity recognition model to obtain the feature similarity output by the similarity recognition model.

[0154] In one possible scenario, the feature similarity model can be a trained first softmax model, which can determine the feature similarity category between the medical title and the medical search statement, and use the feature similarity category to characterize the degree of similarity between the medical title and the medical search statement.

[0155] S308: For each medical title, the dimensionality-reduced title vector and the dimensionality-reduced sentence vector are input into the intent recognition model to obtain the intent matching results between the medical title and the medical search sentence.

[0156] The intent matching result is used to characterize whether the medical intent is the same between the medical title and the medical search statement.

[0157] For example, the intent recognition model can be a trained second softmax model, which can determine the intent matching results between medical titles and medical search statements.

[0158] S309, combining the feature similarity and intent matching results between each medical title and the medical search statement, determine the matching degree ranking of multiple medical titles.

[0159] This step S309 can be found in the relevant description of the previous embodiments, and will not be repeated here.

[0160] For ease of understanding Figure 3 For an example, see [link to example]. Figure 4 It shows a block diagram illustrating one implementation principle of medical title matching in this application. Figure 4 Taking the BERT model as an example, the vector transformation model is the pooling model, the similarity recognition model is the first softmax classification model, and the intent recognition model is the second softmax classification model.

[0161] Depend on Figure 4 As can be seen, the medical search query can be processed by the BERT model to obtain a query vector u, and the medical title can be processed by the BERT model to obtain a title vector v. Based on this, the query vector u will undergo dimensionality reduction through a pooling layer to obtain a dimensionality-reduced query vector u; similarly, the medical title vector v will also undergo dimensionality reduction through a pooling layer to obtain a dimensionality-reduced title vector v.

[0162] It should be noted that, Figure 4 To facilitate understanding of the process of obtaining the dimensionality-reduced statement vector u and the dimensionality-reduced title vector v, the diagram shows the branches of medical search statements through the BERT model and pooling layer, as well as the branches of medical titles through the BERT model and pooling layer. However, in practical applications, the BERT model used to process medical search statements and medical titles can be the same, and correspondingly, the pooling layer can also be the same.

[0163] Based on the above, this application calculates the difference vector uv between the dimension-reduced sentence vector u and the dimension-reduced title vector v. Then, the dimension-reduced sentence vector u, the dimension-reduced title vector v, and the difference vector uv are input into the first softmax classification model, which serves as the similarity recognition model, to obtain the similarity category between medical titles and medical search statements.

[0164] Meanwhile, the dimensionality-reduced sentence vector u and the dimensionality-reduced title vector v are also input into the second softmax classification model, which serves as the intent recognition model, to obtain the intent matching results between medical titles and medical search statements.

[0165] Based on this, by combining the similarity category and intent matching results between each medical title and the medical search statement, the matching degree ranking of each medical title and the medical search statement can be determined.

[0166] In the embodiments of this application, the vector transformation model can be trained separately. After the BERT model is trained, the intent recognition model and the similarity recognition model can be trained simultaneously, or the intent recognition model and the similarity recognition model can be trained separately.

[0167] To facilitate understanding, the following sections will introduce the possible scenarios for training each of the above models in this application.

[0168] First, we will introduce the training process of the vector transformation model. For ease of explanation, we will still use the BERT model as the vector transformation model, and take the medical content pointed to by the medical title as the medical text content as an example.

[0169] like Figure 5 As shown, it illustrates one implementation principle of training the BERT model in this application.

[0170] This application can obtain multiple medical corpus samples, each of which includes a medical title sample and the corresponding medical text content.

[0171] For each medical corpus sample, a word sequence can be obtained, consisting of multiple words from the medical title sample and the medical text content that constitute that medical corpus sample. Based on this, the words in the word sequence can be masked, resulting in a partially masked word sequence for each medical corpus sample.

[0172] like Figure 5 As shown, some words in the word sequence corresponding to the medical title sample in the medical corpus are masked; similarly, some words in the word sequence corresponding to the medical text content in the same medical corpus are also masked. Based on this, after inputting the masked word sequence with masking labels into the BERT model to be trained, the BERT model can determine the word vector of each word in the masked word sequence based on the contextual relationship between each word in the masked word sequence.

[0173] The word vectors of each word in the masked word sequence output by the BERT model can be input into a fully connected network layer. Through the fully connected network layer, the masking probability of each word in the masked word sequence belonging to the masked word can be obtained. Based on this, according to the masking probability of each word in the masked word sequence, the predicted masked words in the masked word sequence can be obtained.

[0174] Accordingly, by combining the actual masked words and the predicted masked words in the masked word sequences corresponding to each medical corpus sample, the prediction accuracy of the BERT model and the fully connected network layer is analyzed. If the prediction accuracy does not meet the requirements, the parameters in the BERT model and the fully connected network layer are adjusted, and the BERT model is retrained using each medical corpus sample until the prediction accuracy meets the requirements.

[0175] After training the BERT model, it can be combined with... Figure 4 The architecture diagram shown illustrates the training of the similarity recognition model and the intent recognition model. The example uses a first-normalized softmax classification model for the similarity recognition model and a second-normalized softmax classification model for the intent recognition model. Figure 4 Please provide an explanation. For example... Figure 6 The diagram illustrates an implementation flow of training a similarity recognition model and an intent recognition model according to this application. This embodiment may include:

[0176] S601, obtain multiple first sample pairs labeled with similarity categories and multiple second sample pairs labeled with intention matching results.

[0177] Each of the first and second sample pairs consists of a medical title sample and a medical search statement sample.

[0178] The similarity categories can include at least two categories: feature similarity and feature dissimilarity. Additionally, a feature similarity category that falls between feature similarity and feature dissimilarity can be set as needed.

[0179] It is understood that this embodiment uses the training of a similarity recognition model with a first sample pair and the training of an intent recognition model with a second sample pair as an example. In practical applications, multiple sample pairs for training the similarity recognition model and the intent recognition model can also accomplish the same task. In this case, each sample pair can be labeled with both the similarity category and the intent matching result.

[0180] S602, For any one of the first sample pair and the second sample pair, use the trained BERT model to determine the title vector corresponding to the medical title sample and the statement vector corresponding to the medical search statement sample in the sample pair, and then execute step S603.

[0181] S603, using the pooling model to be trained, the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample are pooled respectively to obtain the pooled title vector and the pooled statement vector.

[0182] S604. For each first sample pair, calculate the difference vector between the pooled statement vector and the pooled title vector corresponding to the first sample pair, and input the title vector of the medical title sample, the statement vector corresponding to the medical search statement sample, and the difference vector into the first normalized classification model to obtain the similarity category predicted by the first normalized classification model.

[0183] S605, for each second sample pair, input the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample in the second sample pair into the second normalized classification model to obtain the intent recognition result predicted by the second normalized classification model.

[0184] S606, based on the similarity category of each first sample to the corresponding actual annotation and the predicted similarity category, and the intent recognition result of each second sample to the corresponding actual annotation and the predicted intent recognition result, check whether the training termination condition is met. If yes, training ends; if no, adjust the parameters in the first normalized classification model, the second normalized classification model and the pooling model, and return to step S603 until the training termination condition is met.

[0185] The training termination condition can be set as needed.

[0186] For example, the loss function value can be calculated according to the set loss function; if the loss function value is found to converge, then the training termination condition is met.

[0187] For the first normalized classification model, the objective function is to optimize the matching classification function: softmax(u,v,|uv|).

[0188] For example, cross-entropy loss can be used to optimize this objective function, and the cross-entropy loss can be expressed as follows:

[0189]

[0190] Where n corresponds to the number of the first sample pairs, m is the set number of similarity categories, and y ij This represents the label of the i-th sample pair belonging to similarity category j. If the predicted similarity category corresponding to the i-th sample pair belongs to at least one of the set similarity categories, then the label is 1; otherwise, it is 0. For a single classification task, since there is only one classification, only one category's label is non-zero. f(x) ij The loss represents the probability that sample pair i is predicted to be of similarity category j. The magnitude of the loss depends entirely on the probability of being classified into the correct label category. When all samples are classified correctly, the loss = 0; otherwise, it is greater than 0.

[0191] exist Figure 6 The first and second normalized classification models can be viewed as a multi-task model training process. Based on this, the overall objective function for training this multi-task model can be expressed as:

[0192] Obj_total=alpha*softmax(u,v,|uv|)+(1-alpha)*ObjFuntion(class(u),class(v))

[0193] Here, alpha represents the importance ratio of softmax(u,v,|uv|) corresponding to the first normalized classification model, and its value ranges from 0 to 1. Generally, alpha can be set to a value greater than 0.5 and less than 1.

[0194] For example, if the first sample pair and the second sample pair are the same, the purpose of training the multi-task model is to make the values ​​determined for each sample pair based on the overall objective function converge.

[0195] Understandably, after training, the first normalized classification model is the similarity recognition model mentioned earlier, while the second normalized classification model is the intent recognition model, and the pooling model can be a vector dimensionality reduction model.

[0196] Understandable Figure 6This is merely one implementation of training the intent recognition model and similarity recognition model in this application. In practical applications, the similarity recognition model can be trained first, and the pooling model can be trained simultaneously during the training process. Based on the completed similarity recognition model, the intent recognition model can be trained separately using multiple second sample pairs, building upon the already trained BERT model and pooling model.

[0197] The following section will introduce an application scenario, using a medical search platform as an example, where the medical encyclopedia is located.

[0198] like Figure 7 As shown, the medical dictionary platform 710 may include multiple servers 711 that provide medical encyclopedia dictionary services.

[0199] The terminal 720 can install a medical encyclopedia application corresponding to the medical encyclopedia.

[0200] Terminal 720 can send a medical search request to the server 711 of the medical dictionary platform through the medical encyclopedia application. This medical search request carries a medical search query.

[0201] The server 711 of the medical dictionary platform obtains the medical search statement from the medical search request, and, according to the scheme of any of the above embodiments of this application, determines the feature similarity and intent matching results between each medical text content and the medical search statement; for each medical title, it determines the matching degree between the medical text content and the medical search statement by combining the feature similarity and intent matching results between the medical title and the medical search statement. Based on this, the server of the medical dictionary platform can return medical science articles corresponding to at least one medical title with a high matching degree to the terminal according to the matching degree corresponding to each medical title.

[0202] Correspondingly, the terminal can display various medical science articles returned by the medical dictionary platform's server. For example... Figure 8 The image shown is a schematic diagram of an interface displaying the searched medical text content on the terminal. Figure 8 As shown, after entering the medical search term "childhood anorexia" in the search bar on the terminal, the server returns medical science articles in the following order: articles titled "How to care for children with anorexia" and "How to treat symptoms of childhood anorexia," etc. Clicking on a specific article allows users to view its detailed content.

[0203] Corresponding to the medical title matching method of this application, this application also provides a medical title matching device. For example... Figure 9The diagram illustrates the structural composition of an embodiment of a medical title matching device according to this application. The device may include:

[0204] Statement acquisition unit 901 is used to acquire medical search statements;

[0205] The first vector determination unit 902 is used to determine the title vector of each medical title among multiple medical titles to be matched.

[0206] The second vector determination unit 903 is used to determine the statement vector of the medical search statement;

[0207] The feature determination unit 904 is used to determine the feature similarity between the medical title and the medical search statement for each medical title, based on the title vector of the medical title and the statement vector of the medical search statement.

[0208] The intent determination unit 905 is used to determine the intent matching result between the medical title and the medical search statement for each medical title based on the title vector of the medical title and the statement vector of the medical search statement, and using the intent recognition model. The intent matching result is used to characterize whether the medical intent between the medical title and the medical search statement is the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled by themselves, and using the vectors of the medical title samples and medical search statement samples in each first sample pair.

[0209] The matching determination unit 906 is used to determine the matching degree ranking of the multiple medical titles by combining the feature similarity and intent matching results between each medical title and the medical search statement.

[0210] In one possible implementation, the device may further include:

[0211] The difference determination unit is used to determine the vector difference between the statement vector of the medical search statement and the title vector of the medical title before the feature determination unit determines the feature similarity between the medical title and the medical search statement, and obtain the difference vector.

[0212] Accordingly, the feature determination unit is specifically used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector.

[0213] As an alternative approach, the feature determining unit includes:

[0214] The feature determination subunit is used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, and using a similarity recognition model.

[0215] The similarity recognition model is trained based on the feature similarity of multiple second samples labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within each second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample.

[0216] In one alternative embodiment, the device further includes:

[0217] The first vector dimensionality reduction unit is used to reduce the dimensionality of the statement vector of the medical search statement by means of a vector dimensionality reduction model before the feature determination unit and the intent determination unit determine the feature similarity and intent matching results between the medical title and the medical search statement. The vector dimensionality reduction model is trained by using the title vector of the medical title sample and the statement vector of the medical search statement sample in the second sample pair during the training of the similarity recognition model.

[0218] The second vector dimensionality reduction unit is used to reduce the dimensionality of the title vector of the medical title using the vector dimensionality reduction model;

[0219] The difference determination unit is specifically used to determine the vector difference between the statement vector of the medical search statement after dimensionality reduction and the title vector of the medical title after dimensionality reduction, and obtain the difference vector.

[0220] In one possible implementation, the first vector determination unit is specifically used to determine the title vector of the medical title using a vector transformation model. The vector transformation model is a transformer-based bidirectional encoding representation BERT model, and it is trained using masked word sequences corresponding to multiple medical corpus samples, with the prediction of masked words in the masked word sequences as the training objective. The medical corpus samples consist of medical title samples and the medical text content represented by the medical title samples, and the masked word sequence is a word sequence obtained by masking at least one word contained in the medical corpus sample.

[0221] The second vector determination unit is specifically used to determine the statement vector of the medical search statement using the vector transformation model.

[0222] Furthermore, this application also provides a server, which is a server within a medical search platform. For example... Figure 10 This illustrates a schematic diagram of the server architecture provided in this application. Figure 10 In this context, the server 1000 may include a processor 1001 and a memory 1002.

[0223] Optionally, the server may also include: a communication interface 1003, an input unit 1004, a display 1005, and a communication bus 1006.

[0224] The processor 1001, memory 1002, communication interface 1003, input unit 1004 and display 1005 communicate with each other through communication bus 1006.

[0225] In this embodiment of the application, the processor 1001 may be a central processing unit, an application-specific integrated circuit, etc.

[0226] The processor can call programs stored in memory 1002. Specifically, the processor can execute the server-side operations described in the above embodiments.

[0227] The memory 1002 is used to store one or more programs, which may include program code, and the program code includes computer operation instructions. In the embodiments of this application, the memory stores at least the medical title matching method for implementing any of the above embodiments.

[0228] In one possible implementation, the memory 1002 may include a program storage area and a data storage area, wherein the program storage area may store the operating system, the programs mentioned above, etc.; and the data storage area may store data created during the use of the server.

[0229] The communication interface 1003 can be used as an interface for a communication module.

[0230] This application may also include an input unit 1004, which may include a touch sensing unit, a keyboard, etc.

[0231] The display 1005 includes a display panel, such as a touch display panel.

[0232] certainly, Figure 10 The server structure shown does not constitute a limitation on the server in the embodiments of this application. In practical applications, the server may include more than [other components]. Figure 10 More or fewer components as shown, or combinations of certain components.

[0233] On the other hand, this application also provides a storage medium storing computer-executable instructions, which, when loaded and executed by a processor, implement the medical title matching method as described in any of the above embodiments.

[0234] This application also proposes a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the above-described medical title matching method or medical title matching apparatus. Specific implementation processes can be referred to the descriptions of the corresponding embodiments above, and will not be repeated here.

[0235] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. Furthermore, the features described in the various embodiments of this specification can be substituted or combined with each other, enabling those skilled in the art to implement or use this application. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0236] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0237] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0238] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A medical title matching method, characterized in that, include: Obtain medical search terms; For each of the multiple medical titles to be matched, determine the title vector of the medical title; Determine the statement vector of the medical search statement; The difference vector is obtained by determining the vector difference between the statement vector of the medical search query and the title vector of the medical title; Based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, a similarity recognition model is used to determine the feature similarity between the medical title and the medical search statement; wherein, the similarity recognition model is trained based on the feature similarity of multiple second samples labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within the second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample. For each medical title, based on the title vector of the medical title and the statement vector of the medical search query, an intent recognition model is used to determine the intent matching result between the medical title and the medical search query. The intent matching result is used to characterize whether the medical intent between the medical title and the medical search query is the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled by themselves, and using the vectors of the medical title samples and medical search query samples within each first sample pair. The similarity recognition model and the intent recognition model share a vector transformation model and a vector dimensionality reduction model, and are trained synchronously through a joint loss function. The vector transformation model is a BERT model, and the vector dimensionality reduction model is a pooling model. By combining the feature similarity and intent matching results between each medical title and the medical search statement, the matching degree ranking of the multiple medical titles is determined; The training process for the similarity recognition model and the intent recognition model includes: Obtain multiple second sample pairs labeled with similarity categories and multiple first sample pairs labeled with intention matching results; For any one of the first sample pair and the second sample pair, the trained BERT model is used to determine the title vector corresponding to the medical title sample and the statement vector corresponding to the medical search statement sample in the sample pair, respectively. The title vector of the medical title sample and the statement vector corresponding to the medical search statement sample are pooled using the pooling model to be trained, respectively, to obtain the pooled title vector and the pooled statement vector. For each second sample pair, calculate the difference vector between the pooled statement vector and the pooled title vector corresponding to the second sample pair, and input the title vector of the medical title sample, the statement vector corresponding to the medical search statement sample, and the difference vector into the first normalized classification model to obtain the similarity category predicted by the first normalized classification model. For each first sample pair, the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample in the first sample pair are input into the second normalized classification model to obtain the intent recognition result predicted by the second normalized classification model. Based on the similarity category and predicted similarity category of each second sample to the corresponding actual annotation, and the intent recognition result and predicted intent recognition result of each first sample to the corresponding actual annotation, it is checked whether the training termination condition is met. If yes, training ends; if no, the parameters in the first normalized classification model, the second normalized classification model, and the pooling model are adjusted, and the process returns to using the pooling model to be trained to pool the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample, respectively, to obtain the pooled title vector and the pooled statement vector, until the training termination condition is met.

2. The method according to claim 1, characterized in that, Before determining the feature similarity and intent matching results between the medical title and the medical search statement, the process also includes: The dimensionality reduction of the medical search statement vector is performed by a vector dimensionality reduction model, which is obtained by training the similarity recognition model using the title vector of the medical title sample and the statement vector of the medical search statement sample in the second sample pair. The dimensionality of the medical title vector is reduced using the vector dimensionality reduction model described above. The step of determining the vector difference between the statement vector of the medical search query and the title vector of the medical title to obtain the difference vector includes: The difference vector is obtained by determining the vector difference between the statement vector of the medical search statement after dimensionality reduction and the title vector of the medical title after dimensionality reduction.

3. The method according to claim 1, characterized in that, Determining the title vector of the medical title includes: The title vector of the medical title is determined using a vector transformation model; The process of determining the statement vector of the medical search statement includes: The vector transformation model is used to determine the statement vector of the medical search statement; The vector transformation model is a bidirectional encoding representation BERT model based on a transformer, and the vector transformation model is trained by using masked word sequences corresponding to multiple medical corpus samples and predicting the masked words in the masked word sequences as the training target. The medical corpus sample consists of a medical title sample and the medical text content represented by the medical title sample, and the masked word sequence is a word sequence obtained by masking at least one word contained in the medical corpus sample.

4. A method for medical title matching, characterized in that, include: The difference vector is obtained by determining the vector difference between the statement vector of the medical search query and the title vector of each medical title among the multiple medical titles to be matched; Based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, a similarity recognition model is used to determine the feature similarity between the medical title and the medical search statement; wherein, the similarity recognition model is trained based on the feature similarity of multiple second samples labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within the second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample. For each medical title, based on the title vector of the medical title and the statement vector of the medical search query, an intent matching result is determined using an intent recognition model. The intent matching result characterizes whether the medical intents of the medical title and the medical search query are the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled with their respective vectors, using the vectors of the medical title samples and medical search query samples within each first sample pair. The similarity recognition model and the intent recognition model share a vector transformation model and a vector dimensionality reduction model, and are trained synchronously using a joint loss function. The vector transformation model is a BERT model, and the vector dimensionality reduction model is a pooling model. By combining the feature similarity and intent matching results of each medical title and the medical search statement, the matching result of the medical search statement is obtained; The training process for the similarity recognition model and the intent recognition model includes: Obtain multiple second sample pairs labeled with similarity categories and multiple first sample pairs labeled with intention matching results; For any one of the first sample pair and the second sample pair, the trained BERT model is used to determine the title vector corresponding to the medical title sample and the statement vector corresponding to the medical search statement sample in the sample pair, respectively. The title vector of the medical title sample and the statement vector corresponding to the medical search statement sample are pooled using the pooling model to be trained, respectively, to obtain the pooled title vector and the pooled statement vector. For each second sample pair, calculate the difference vector between the pooled statement vector and the pooled title vector corresponding to the second sample pair, and input the title vector of the medical title sample, the statement vector corresponding to the medical search statement sample, and the difference vector into the first normalized classification model to obtain the similarity category predicted by the first normalized classification model. For each first sample pair, the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample in the first sample pair are input into the second normalized classification model to obtain the intent recognition result predicted by the second normalized classification model. Based on the similarity category and predicted similarity category of each second sample to the corresponding actual annotation, and the intent recognition result and predicted intent recognition result of each first sample to the corresponding actual annotation, it is checked whether the training termination condition is met. If yes, training ends; if no, the parameters in the first normalized classification model, the second normalized classification model, and the pooling model are adjusted, and the process returns to using the pooling model to be trained to pool the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample, respectively, to obtain the pooled title vector and the pooled statement vector, until the training termination condition is met.

5. A medical title matching device, characterized in that, include: The statement acquisition unit is used to obtain medical search statements; The first vector determination unit is used to determine the title vector of each medical title among multiple medical titles to be matched; The second vector determination unit is used to determine the statement vector of the medical search statement; The feature determination unit is used to determine the feature similarity between the medical title and the medical search statement for each medical title, based on the title vector of the medical title and the statement vector of the medical search statement; An intent determination unit is used to determine the intent matching result between a medical title and a medical search statement for each medical title, based on the title vector of the medical title and the statement vector of the medical search statement, and using an intent recognition model. The intent matching result is used to characterize whether the medical intents of the medical title and the medical search statement are the same. The intent recognition model is trained based on the intent matching results of multiple first samples labeled by themselves, and using the vectors of the medical title samples and medical search statement samples within each first sample pair. The similarity recognition model and the intent recognition model share a vector transformation model and a vector dimensionality reduction model, and are trained synchronously through a joint loss function. The vector transformation model is a BERT model, and the vector dimensionality reduction model is a pooling model. The matching determination unit is used to determine the matching degree ranking of the multiple medical titles by combining the feature similarity and intent matching results between each medical title and the medical search statement; The difference determination unit is used to determine the vector difference between the statement vector of the medical search statement and the title vector of the medical title before the feature determination unit determines the feature similarity between the medical title and the medical search statement, and obtain the difference vector. The feature determination unit is specifically used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector. The feature determination unit includes: The feature determination subunit is used to determine the feature similarity between the medical title and the medical search statement based on the title vector of the medical title, the statement vector of the medical search statement, and the difference vector, and using a similarity recognition model. The similarity recognition model is trained based on the feature similarity of multiple second samples labeled with each other, and using the vectors corresponding to the medical title sample and the medical search statement sample within the second sample pair, as well as the difference vector between the vectors of the medical title sample and the medical search statement sample. The device is also used for: Obtain multiple second sample pairs labeled with similarity categories and multiple first sample pairs labeled with intention matching results; For any one of the first sample pair and the second sample pair, the trained BERT model is used to determine the title vector corresponding to the medical title sample and the statement vector corresponding to the medical search statement sample in the sample pair, respectively. The title vector of the medical title sample and the statement vector corresponding to the medical search statement sample are pooled using the pooling model to be trained, respectively, to obtain the pooled title vector and the pooled statement vector. For each second sample pair, calculate the difference vector between the pooled statement vector and the pooled title vector corresponding to the second sample pair, and input the title vector of the medical title sample, the statement vector corresponding to the medical search statement sample, and the difference vector into the first normalized classification model to obtain the similarity category predicted by the first normalized classification model. For each first sample pair, the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample in the first sample pair are input into the second normalized classification model to obtain the intent recognition result predicted by the second normalized classification model. Based on the similarity category and predicted similarity category of each second sample to the corresponding actual annotation, and the intent recognition result and predicted intent recognition result of each first sample to the corresponding actual annotation, it is checked whether the training termination condition is met. If yes, training ends; if no, the parameters in the first normalized classification model, the second normalized classification model, and the pooling model are adjusted, and the process returns to using the pooling model to be trained to pool the title vector of the medical title sample and the statement vector corresponding to the medical search statement sample, respectively, to obtain the pooled title vector and the pooled statement vector, until the training termination condition is met.

6. The apparatus according to claim 5, characterized in that, The device further includes: The first vector dimensionality reduction unit is used to reduce the dimensionality of the statement vector of the medical search statement by a vector dimensionality reduction model before determining the feature similarity and intent matching results between the medical title and the medical search statement. The vector dimensionality reduction model is obtained by training the title vector of the medical title sample and the statement vector of the medical search statement sample in the second sample pair during the training of the similarity recognition model. The second vector dimensionality reduction unit is used to reduce the dimensionality of the title vector of the medical title using the vector dimensionality reduction model; The difference determination unit is specifically used to determine the vector difference between the statement vector of the medical search statement after dimensionality reduction and the title vector of the medical title after dimensionality reduction, and obtain the difference vector.

7. The apparatus according to claim 5, characterized in that, The first vector determination unit is specifically used to determine the title vector of the medical title using a vector transformation model, wherein the vector transformation model is a transformer-based bidirectional encoding representation BERT model, and the vector transformation model is trained using masked word sequences corresponding to multiple medical corpus samples, with the prediction of the masked words in the masked word sequences as the training target; the medical corpus samples consist of medical title samples and medical text content represented by the medical title samples, and the masked word sequence is a word sequence obtained after masking at least one word in the medical corpus samples; The second vector determination unit is specifically used to determine the statement vector of the medical search statement using the vector transformation model.

8. A server, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program, which, when executed, is specifically used to implement the medical title matching method as described in any one of claims 1 to 4.

9. A storage medium, characterized in that, Used to store a program, which, when executed, is used to implement the medical title matching method as described in any one of claims 1 to 4.

10. A computer program product, characterized in that, The computer program product includes instructions that, when executed on a computer device, cause the computer device to perform the medical title matching method as described in any one of claims 1 to 4.