Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By constructing a vector representation model of entity words in the co-occurrence graph in the medical business field, the problem of low search accuracy in the medical business field is solved, and more accurate and comprehensive synonym mining and search results are achieved.

CN115204154BActive Publication Date: 2025-06-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210794411.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-06-17
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

In the search of the medical business field, the search accuracy is not high due to colloquial and diverse language forms, the dictionary of annotation synonyms is not covered in complete coverage and the synonym mining effect is poor.

Method used

By constructing a co-occurrence graph, using the entity words extracted from the question text, reply text and the description text of the reply account, vector representation is performed based on the target graph embedding model to obtain entity word pairs with similar meanings.

Benefits of technology

The accuracy of feature vector representation of low-frequency words and the adequacy of vector representation learning of colloquial entity words is improved, thereby improving the accuracy of synonym mining and comprehensive coverage, and improving the search accuracy of Q&A text data in the medical business field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115204154B_ABST
    Figure CN115204154B_ABST
Patent Text Reader

Abstract

The present application relates to a data processing method, apparatus, electronic device, and storage medium. The method includes: obtaining a set of entity words to be processed; performing vector representation processing on the entity words in the set of entity words based on a target graph embedding model to obtain word vectors of each entity word; the graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, and the co-occurrence graph is constructed based on entity words respectively extracted from question texts, reply texts, and description texts of reply accounts included in multiple sample Q&A text corpora in the medical business field, and the description text is used to indicate the sub-business field corresponding to the reply account in the medical business field; obtaining entity word pairs with similar meanings in the set of entity words according to the similarity between the word vectors. According to the technical solution of the present application, the accuracy of synonym mining and the search accuracy of Q&A texts can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet application technologies, and in particular, to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of the Internet, there are more and more rich online interactions, such as medical Q&A interactions. This can facilitate users to ask questions to get timely responses. When there are similar searches later, these Q&A data can be displayed. For these more professional fields, the words used in searches are generally more colloquial and the language forms are diverse, resulting in low search accuracy. In related technologies, a labeled synonym dictionary is used to improve the search effect, or synonym mining is used to improve the search accuracy. However, the former requires a large amount of labeling work, and the covered synonyms are more formal, and the coverage of colloquial words is incomplete; the latter has the phenomenon of uneven colloquial words due to diverse languages, resulting in poor synonym mining effects. Summary of the Invention

[0003] In view of the above technical problems, this application provides a data processing method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of this application, a data processing method is provided. The method includes:

[0005] Obtain a set of entity words to be processed;

[0006] Perform vector representation processing on the entity words in the set of entity words based on a target graph embedding model to obtain word vectors of each entity word; the target graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, and the co-occurrence graph is constructed based on entity words respectively extracted from question texts, answer texts, and description texts of answer accounts included in multiple sample Q&A text corpora in the medical business field. The description text is used to indicate the sub-business field corresponding to the answer account in the medical business field; the edges in the co-occurrence graph have edge weights, and the edge weight of each edge is obtained based on at least one of the co-occurrence probability of the nodes connected by each edge, the node type, and the semantic similarity information between the nodes. The nodes are the entity words, and the node type represents a preset dimension belonging to the corresponding entity word and used to describe the medical business field;

[0007] Obtain pairs of entity words with similar meanings in the set of entity words according to the similarity between the word vectors.

[0008] According to another aspect of this application, a data processing apparatus is provided, including:

[0009] An obtaining module, configured to obtain a set of entity words to be processed;

[0010] A vector representation module, configured to perform vector representation processing on the entity words in the entity word set based on a target graph embedding model to obtain word vectors of each entity word; the target graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, and the co-occurrence graph is constructed based on entity words respectively extracted from question texts, answer texts, and description texts of answer accounts included in multiple sample Q&A text corpora in the medical business field, and the description text is used to indicate a sub-business field corresponding to the answer account in the medical business field; edges in the co-occurrence graph have edge weights, and the edge weight of each edge is obtained based on at least one of the co-occurrence probability of nodes connected by each edge, node types, and semantic similarity information between nodes, the nodes are the entity words, and the node types represent preset dimensions to which the corresponding entity words belong and are used to describe the medical business field;

[0011] An entity word pair acquisition module, configured to acquire entity word pairs with similar meanings in the entity word set according to the similarity between the word vectors.

[0012] According to another aspect of the present application, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the above method.

[0013] According to another aspect of the present application, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein, when the computer program instructions are executed by a processor, the above method is implemented.

[0014] By introducing entity words in three types of texts, namely entity words respectively extracted from question texts, answer texts, and description texts of answer accounts, into the co-occurrence graph, the co-occurrence times of low-frequency word entities in the co-occurrence graph can be effectively balanced, the accuracy learning of the feature vector representation of low-frequency words can be improved, so that the vector representation of entity words by the target graph embedding model trained according to the co-occurrence graph is more accurate, and further the accuracy of synonym mining of low-frequency words can be improved; and due to the introduction of the description text, the colloquial entity words from the question text in the co-occurrence graph can have an association relationship with the entity words indicating the sub-business field, making the vector representation learning of the colloquial entity words more sufficient, so that the synonym mining based on the target graph embedding model can improve the accuracy and comprehensiveness of synonym mining of colloquial entity words; that is, by performing synonym mining according to the target graph embedding model trained according to the co-occurrence graph, the accuracy and coverage comprehensiveness of synonym mining can be improved, and further the search accuracy of Q&A text data in the medical business field can be improved, and even in the case of colloquial search terms or low-frequency search terms in the medical business field, the search accuracy is relatively high.

[0015] Other features and aspects of the present application will become apparent from the following detailed description of the exemplary embodiments with reference to the accompanying drawings. Description of the Drawings

[0016] The drawings included in and constituting a part of the specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present application and are used to explain the principles of the present application.

[0017] Figure 1 A schematic diagram showing an application system provided according to an embodiment of the present application.

[0018] Figure 2 A flowchart showing a data processing method provided according to an embodiment of the present application.

[0019] Figure 3a and Figure 3b A schematic diagram showing a sample question-and-answer text corpus provided according to an embodiment of the present application.

[0020] Figure 4a A schematic diagram showing a co-occurrence graph provided according to an embodiment of the present application.

[0021] Figure 4b A schematic diagram showing the embedding vector representation of each entity word in a co-occurrence graph provided according to an embodiment of the present application.

[0022] Figure 5 A flowchart showing a method for obtaining a co-occurrence graph corresponding to multiple sample question-and-answer text corpora in the medical business field provided according to an embodiment of the present application.

[0023] Figure 6 A schematic diagram showing a text sequence corresponding to a node provided according to an embodiment of the present application.

[0024] Figure 7 A schematic diagram showing the training process of a skip-gram language model provided according to an embodiment of the present application.

[0025] Figure 8 A block diagram showing a data processing device provided according to an embodiment of the present application.

[0026] Figure 9 A block diagram showing an electronic device for data processing provided according to an embodiment of the present application. Detailed Embodiments

[0027] The following will detail various exemplary embodiments, features, and aspects of the present application with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0028] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior to or better than other embodiments.

[0029] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present application.

[0030] Please refer to Figure 1 , Figure 1 which shows a schematic diagram of an application system provided according to an embodiment of the present application. The application system can be used for the data processing method of the present application. As Figure 1 shown, the application system can at least include a server 01 and a terminal 02.

[0031] In an embodiment of the present application, the server 01 can be used for data processing. The server 01 can include an independent physical server, or can be a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0032] In an embodiment of the present application, the terminal 02 can be used to provide a search application or page, enabling a user to perform data search, and can receive and display target text data. The terminal 02 can include entity devices of types such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. The entity devices can also include software running on the entity devices, such as application programs. The operating system running on the terminal 02 in an embodiment of the present application can include, but is not limited to, Android system, IOS system, Linux, Windows, etc.

[0033] In an embodiment of this specification, the above terminal 02 and server 01 can be directly or indirectly connected through wired or wireless communication means, and the present application does not limit this.

[0034] In a specific embodiment, when the server 02 is a distributed system, the distributed system can be a blockchain system. When the distributed system is a blockchain system, it can be formed by multiple nodes (any form of computing device connected to the network, such as a server or a user terminal). A peer-to-peer (P2P) network is formed among the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server or a terminal can join and become a node. A node includes a hardware layer, a middle layer, an operating system layer, and an application layer. Specifically, the functions of each node in the blockchain system may involve the following functions:

[0035] 1) Routing, which is a basic function of a node and is used to support communication between nodes.

[0036] In addition to the routing function, a node may also have the following functions:

[0037] 2) Application, which is used to be deployed in the blockchain, implement specific services according to actual business needs, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes successfully verify the source and integrity of the record data, the record data is added to the temporary block.

[0038] It should be noted that in the specific implementation of this application, when it comes to user-related data, when the following embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0039] Figure 2 The flowchart of a data processing method provided according to an embodiment of the present application is shown. As Figure 2 shown, the data processing method may include:

[0040] S201, obtain a set of entity words to be processed.

[0041] In the embodiments of this specification, the set of entity words to be processed may be a set of a large number of entity words that need to perform synonym mining. These a large number of entity words may be entity words in the medical business field, that is, medical entity words.

[0042] S203, perform vector representation processing on the entity words in the set of entity words based on the target graph embedding model to obtain the word vectors of each entity word.

[0043] Among them, the graph embedding model can be obtained by training a preset graph embedding model according to a co-occurrence graph. The co-occurrence graph is constructed from entity words respectively extracted from question texts, answer texts, and description texts of answer accounts included in multiple sample Q&A text corpora in the medical business domain. Here, the description text can be used to indicate the sub-business domain corresponding to the answer account in the medical business domain. Optionally, the edges in the co-occurrence graph may not have edge weights; or the edges in the co-occurrence graph may have edge weights. In the case of having edge weights, the edge weight of each edge can be obtained based on at least one of the co-occurrence probability of the nodes connected by each edge, the node type, and the semantic similarity information between the nodes. The nodes can be entity words. The node type can represent a preset dimension belonging to the corresponding entity word and used to describe the medical business domain. The preset dimension can include dimensions such as departments, diseases, symptoms, drugs, and examination methods. The present disclosure does not limit this. Here, the co-occurrence probability can be the ratio of the number of times the two nodes connected by each edge appear to the sum of the number of times each of the two nodes appears; the semantic similarity information can be the semantic similarity between the two entity words corresponding to the two nodes connected by each edge, such as the distance in the semantic vector representation. As an example, the edge weight can be positively correlated with the semantic similarity information, the co-occurrence probability, and the node type, that is, the higher the semantic similarity information represents the semantic similarity, the higher the co-occurrence probability, and the same node type, the higher the corresponding edge weight. Here, the same node type can represent positive, and different node types can represent negative.

[0044] In the embodiments of this specification, the target graph embedding model can be Node2Vec, MetaPath2Vec, etc. obtained by pre-training a preset Node2Vec or a preset MetaPath2Vec, etc. according to the co-occurrence graph. The present disclosure does not limit this. Based on this, taking Node2Vec as an example, each entity word in the entity word set can be input into Node2Vec for vector representation processing to obtain the word vectors of each entity word.

[0045] In the embodiments of this specification, the corresponding sub-business domain can refer to the sub-business domain to which the answer account belongs in the medical business domain and / or the sub-business domain in which the answer account is proficient in the medical business domain. In the medical business domain, the sub-business domain to which the answer account belongs can refer to the sub-business domain pre-divided for the answer account, such as pediatrics; the sub-business domain in which the answer account is proficient can be the department of respiratory medicine and upper respiratory tract infection. This application does not limit these.

[0046] As an example, the sub-business area can be a subdivision of the medical business area, which can include primary subdivisions, secondary subdivisions, etc. The secondary subdivision can be a further subdivision of the primary subdivision, and the present application does not limit this. Taking the medical field as an example, the sub-business area can be a primary subdivision under the medical field, such as a medical department, such as pediatrics, internal medicine, etc. Or the sub-business area can be a multi-level subdivision under the medical field. For example, the primary subdivision can refer to a medical department. In the case where the primary subdivision is internal medicine, the secondary subdivision can be gastroenterology, respiratory medicine, etc., and a tertiary subdivision can also be carried out, etc., which can be regarded as a subdivision corresponding to specific diseases, and the present disclosure does not limit this.

[0047] In the embodiments of this specification, multiple sample Q&A text corpora can be obtained from the Q&A text data associated with the medical business area. Among them, the Q&A text data associated with the medical business area can be the Q&A text data stored in the application program or website corresponding to the medical business area. The Q&A text data can refer to the text data including questions and answers. Thus, multiple sample Q&A text corpora can be obtained from the Q&A text data. A sample Q&A text corpus can include a question text and a corresponding answer data. The answer data can include an answer text and a description text of the answer account that gives feedback on the answer text. As Figure 3a shown, the question text can be like Figure 3a 301 shown, and the answer data corresponding to the 301 can be like Figure 3a 302 shown. Among them, the answer text can be like Figure 3a 3021 shown, and the description text can be like Figure 3a 3032 shown. Here, the description text can be the sub-business area that the answer account is good at. Optionally, the answer data can also include the avatar, interaction data, etc. of the answer account, and the present application does not limit this.

[0048] Furthermore, a co-occurrence graph can be constructed based on the entity words respectively extracted from the question texts, answer texts, and description texts of the answer accounts included in the multiple sample Q&A text corpora. For example, these entity words can be used as nodes, and edges can be constructed between the entity words from the question text and the entity words from the answer data that belong to the same sample Q&A text corpus to obtain a co-occurrence graph. This can correspond to the situation where the co-occurrence graph does not have edge weights.

[0049] In a possible implementation manner, Figure 5 shows a flowchart of a method for obtaining a co-occurrence graph corresponding to multiple sample Q&A text corpora in the medical business area according to an embodiment of the present application. As Figure 5 shown, the construction process of the co-occurrence graph can include:

[0050] S501. Obtain multiple sample Q&A text corpora from the Q&A text data associated with the medical business domain. Each sample Q&A text corpus may include a question text, a reply text, and a description text of the reply account that provides feedback on the reply text, such as Figure 3a and Figure 3b shown.

[0051] Optionally, the multiple sample Q&A text corpora may be filtered. Based on this, in one example, multiple initial sample Q&A text corpora may be obtained from the Q&A text data associated with the medical business domain; determine the content ratio information between the question text and the reply text in each initial sample Q&A text corpus, and / or the number of views of each initial sample Q&A text corpus. As an example, the content ratio information may refer to the ratio between the number of words in the question text and the number of words in the reply text. In this way, the lower the content ratio information, the more detailed the content of the reply text. The number of views can be queried from the log records, and this application does not limit it.

[0052] Furthermore, sample Q&A text corpora with content ratio information greater than a ratio threshold and / or the number of views less than a number threshold may be filtered out from the multiple initial sample Q&A text corpora to obtain multiple sample Q&A text corpora. Among them, the number threshold may be 5, and this application does not limit it. By filtering the initial sample Q&A text corpora based on the content ratio information and / or the number of views, sample Q&A text corpora with a certain amount of text information and page quality can be effectively selected, making the co-occurrence relationships in the co-occurrence graph more effective and providing a basis for the accurate learning of word vectors. Optionally, initial sample Q&A text corpora with the lengths of the question text and the reply text less than a preset length may also be directly filtered out. This application does not limit the specific filtering method, as long as the quality of the sample Q&A text corpora used to construct the co-occurrence graph can be guaranteed.

[0053] S503. Extract a first entity word from the question text, a second entity word from the reply text, and a third entity word from the description text.

[0054] In the embodiments of this specification, based on a first preset entity word dictionary, such as a preset medical entity word dictionary, a first entity word may be extracted from the question text and a second entity word may be extracted from the reply text respectively. For example, the question text and the reply text may be segmented to obtain multiple segments, and then based on the preset entity word dictionary, matching entity words may be extracted from the multiple segments to obtain the first entity word and the second entity word. The number of each of the first entity word and the second entity word may be one or more. Taking the preset medical entity word dictionary as an example, it may include entity words in multiple dimensions such as diseases, symptoms, drugs, examinations, treatments, etc. Such as Figure 3a and Figure 3bAs shown, the extracted first entity word and second entity word can be as marked in the text box.

[0055] Considering that the entity words in the sample Q&A text corpus of fewer sub-business domains will become low-frequency entity words, that is, they appear fewer times in the co-occurrence graph constructed subsequently, which causes the vector representation learning to shift towards high-frequency entity words, and the vector representation of low-frequency entity words is not accurate enough. Based on this, a third entity word is selected to increase the co-occurrence relationship of low-frequency entity words in the co-occurrence graph. The third entity word can be extracted from the description text based on the second preset entity word dictionary, and the second preset entity word dictionary can include entity words in multiple dimensions for describing sub-business domains in the medical business domain. This application does not limit this. Taking the medical field as an example, the second preset entity word dictionary can include entity words corresponding to departments and diseases respectively. In this way, the third entity word can be regarded as heterogeneous with the first entity word and the second entity word. Among them, the disease can correspond to the subdivision of the department. As an example, as Figure 3a shown, the third entity word can be extracted from Figure 3a shown 3022, such as pediatrics, ophthalmology, fundus disease.

[0056] Among them, the first preset entity word dictionary and the second preset entity word dictionary can be set accordingly based on the entity words associated with the medical business domain. This application does not limit this.

[0057] S505, use the first entity word, the second entity word and the third entity word as nodes, set edges between the first entity word and the second entity word in the same sample Q&A text corpus, set edges between the first entity word and the third entity word in the same sample Q&A text corpus, and set edge weights for each edge to construct a co-occurrence graph.

[0058] As Figure 4a shown, the first entity word, the second entity word and the third entity word extracted from Figure 3a and Figure 3b can be used as nodes, set edges between the first entity word and the second entity word in the same sample Q&A text corpus, set edges between the first entity word and the third entity word in the same sample Q&A text corpus, and set edge weights for each edge to construct a co-occurrence graph. For example, if the first entity word "diarrhea after breastfeeding" and the third entity word "fundus disease" belong to the same sample Q&A text corpus, an edge can be set between "diarrhea after breastfeeding" and "fundus disease"; if the first entity word "breast milk diarrhea" and the second entity word "AAA" belong to the same sample Q&A text corpus, an edge can be set between "breast milk diarrhea" and "AAA". The above-mentioned AAA and BBB can refer to the corresponding medical entity words. This is just a representation here and is not limited. Set edges in this way in turn, and edge weights can be set for each edge, so that Figure 4a the co-occurrence graph shown can be formed. It should be noted thatFigure 4a The co-occurrence graph is just an example. In actual applications, the sample Q&A text corpus can be large, so as to form a large-scale co-occurrence graph. Among them, the edge weights of each edge can be obtained based on at least one of the co-occurrence probabilities of the nodes connected by each edge, the node types, and the semantic similarity information between the nodes. For example, the edge weight of an edge can be determined based on the semantic similarity information between the two nodes (two entity words) connected by the edge, and this edge weight can be positively correlated with the semantic similarity information. The semantic similarity information here can be determined based on the vector distance of each of the two entity words, and the present disclosure does not limit this. Or the weight of an edge can be determined based on the sharing probability of the two nodes connected by the edge, and this edge weight can be positively correlated with the semantic similarity information. By including three types of entity words in the co-occurrence graph, the Q&A information covered by the co-occurrence graph is made more abundant, and the co-occurrence relationship of entity words in sub-business fields with less sample Q&A text corpus can be balanced. Furthermore, in the subsequent vector representation learning of entity words, the learning accuracy of high-frequency entity words and low-frequency entity words can be balanced.

[0059] Optionally, considering that it is desired to more deeply mine oral and low-frequency synonym pairs, and the frequencies of entity words may vary greatly. For example, in the medical field, the frequencies of "cold", "fever", etc. are relatively high, and there may be tens of thousands or even hundreds of thousands of edges in the co-occurrence graph. This not only seriously affects the training performance of the graph network, but may also mask those low-frequency medical synonyms in the generation of text sequences by graph walking, resulting in the model being biased towards the head medical vocabulary during embedding learning. Based on this, high-frequency entity words and their corresponding edges can be deleted. Specifically, the first entity word, the second entity word, and the third entity word can be used as nodes, and an edge can be set between the first entity word and the second entity word in the same sample Q&A text corpus, and an edge can be set between the first entity word and the third entity word in the same sample Q&A text corpus to obtain a first node graph; and the number of edges corresponding to each node in the first node graph can be obtained; thus, the nodes in the first node graph with the number of edges greater than the threshold and their corresponding edges can be deleted to obtain a second node graph. Further, edge weights can be set for the edges in the second node graph to obtain a co-occurrence graph, as Figure 4a shown. It should be noted that Figure 4a the weights of all edges are not shown, and only an example of w1 is given. As an example, the threshold can be 10,000, and the present application does not limit this. By deleting the nodes in the first node graph with the number of edges greater than the threshold and their corresponding edges, the proportion of low-frequency entity words and oral entity words in the co-occurrence graph is made more balanced, so that the accuracy of vector representation learning of oral synonyms and low-frequency entity word synonyms can be improved.

[0060] S205, according to the similarity between word vectors, obtain pairs of entity words with similar meanings in the entity word set.

[0061] In the embodiments of this specification, the distance between word vectors can be obtained as the similarity between word vectors. Thus, entity words corresponding to node pairs with a similarity greater than the similarity threshold can be used as entity word pairs with similar meanings, that is, synonym pairs.

[0062] By introducing entity words in three types of texts into the co-occurrence graph, that is, entity words respectively extracted from the problem text, the response text, and the description text of the response account, the co-occurrence times of low-frequency word entity words in the co-occurrence graph can be effectively balanced, the accuracy learning of the feature vector representation of low-frequency words can be improved, so that the vector representation of entity words by the target graph embedding model trained according to this co-occurrence graph is more accurate, and further the accuracy of synonym mining for low-frequency words can be improved; and due to the introduction of the description text, the colloquial entity words from the problem text in the co-occurrence graph can have an association relationship with the entity words indicating the sub-business field, making the vector representation learning of colloquial entity words more sufficient, so that the synonym mining based on the target graph embedding model can improve the accuracy and comprehensiveness of synonym mining for colloquial entity words; that is, using the target graph embedding model trained according to this co-occurrence graph to perform synonym mining can improve the accuracy and comprehensive coverage of synonym mining, and further improve the search accuracy of the Q&A text data in this medical business field, even in the case of facing colloquial search terms or low-frequency search terms in the medical business field, the search accuracy is relatively high.

[0063] Optionally, the method may further include: obtaining preset synonym information; the preset synonym information represents multiple preset synonym pairs, that is, it may include multiple preset synonym pairs. The preset synonym information may refer to existing synonym information. Thus, the preset synonym information can be updated according to the entity word pair to obtain target synonym information. The update here may be to replace at least one of the original preset synonym pairs, or add the entity word pair to the preset synonym information, so that the preset synonym information is supplemented, and the present application does not limit this. By updating the preset synonym information with the entity word pair, the target synonym information can dynamically adapt to the changes of entity words in the medical business field and improve the accuracy of the target synonym information in terms of timeliness.

[0064] In practical applications, question-and-answer text data can be searched based on target synonym information. Based on this, the method may further include: obtaining a keyword to be searched. For example, by obtaining the search term input by the user, this search term can be used as the keyword to be searched. Furthermore, synonyms of the keyword to be searched can be obtained from the target synonym information as at least one associated search term; thus, text data search processing can be performed according to the keyword to be searched and the associated search terms to obtain target text data, such as medical question-and-answer text data. Since the synonym pairs in the target synonym information cover rich spoken synonyms and low-frequency synonyms, the search can be more accurate, improving the accuracy of the target text data. By using the target synonym information including rich and accurate synonym pairs to search for question-and-answer text, the search accuracy of the question-and-answer text can be improved. Even if the keyword to be searched is relatively colloquial or a low-frequency word in the medical business field, relatively accurate question-and-answer text data can still be obtained.

[0065] Referring to Figure 4a , when the edges in the co-occurrence graph have edge weights, the low-frequency entity words can be further balanced to improve the accuracy of the vector representation learning of the low-frequency entity words. It should be noted that Figure 4a the edge weights in

[0066] are not all marked. Correspondingly, when the edges in the co-occurrence graph have edge weights, the method may further include a process for determining the edge weights. For example, it may include: obtaining a first target node and a second target node connected by a target edge, where the target edge is any edge in the co-occurrence graph; and determining the co-occurrence probability of the first target node and the second target node and their respective node types; determining a first weight of the target edge according to the co-occurrence probability and the node types; obtaining semantic similarity information between the first target node and the second target node; determining a second weight of the target edge according to the semantic similarity information; thus, the edge weight of the target edge can be obtained based on the first weight and / or the second weight.

[0067] W1 = co-occurrence probability * node consistency information

[0068] Among them, the co-occurrence probability = the number of co-occurrences of node pair P1 - P2 / (the number of co-occurrences of all node pairs containing node P1 + the number of co-occurrences of all node pairs containing node P2). This node pair can refer to two nodes with an edge between them. The node consistency information can be determined based on the node type. For example, if the node types are the same, the node consistency information can be 1; if the node types are different, the node consistency information can be 0.5. This application does not make any limitations in this regard. The node type can refer to the type of the corresponding entity word in the medical business field. Taking the medical field as an example, the node type can refer to multiple dimensions for describing a certain type of disease, which can be the dimensions in the above-mentioned first preset entity word dictionary and the second preset entity word dictionary. For example, the node type can be department, disease, symptom, drug, examination method, etc. This application does not make any limitations in this regard.

[0069] For W2, word2vec can be used to obtain the semantic representations of the first target node and the second target node, so that the semantic similarity distance between the semantic representations can be calculated as the semantic similarity information between the first target node and the second target node.

[0070] Correspondingly, the edge weight W can be W1 or W2, or the edge weight W = W1 * W2. In this way, the edge weight can reflect at least one of the co-occurrence probability of the node pair, the node type, and the semantic similarity of the nodes, so that the edge weight can effectively represent the association relationship between the nodes.

[0071] In practical applications, considering that the entity words in the less sample Q&A text corpus may be low-frequency entity words, which may lead to inaccurate vector representation learning. Based on this, edge weights are introduced for random walks, so that the text sequence of low-frequency entity words can effectively express the low-frequency entity words. As an example, the preset graph embedding model may include a node walk model and a preset word vector model. The training process of the preset graph embedding model may include: taking each node in the co-occurrence graph as a starting point, performing weighted random walks to adjacent nodes based on the node walk model and edge weights, for example, preferentially walking to neighbor nodes with high edge weights, to obtain the text sequence corresponding to each node; and training the preset word vector model according to the text sequence to obtain the target graph embedding model. For example, unsupervised iterative training can be performed until the iterative cutoff condition is met, and the preset graph embedding model when the iterative cutoff condition is met is used as the target graph embedding model. Among them, the preset graph embedding model can be a preset Node2Vec, a preset MetaPath2Vec, etc. The node walk model can be a node walk algorithm, and the preset word vector model can be word2vec, such as the skip-gram model, which is not limited in this application. Through the random walk based on edge weights, the obtained text sequence can better express the semantics of the nodes. Combining with the preset word vector model, the acquisition efficiency and accuracy of the feature vectors corresponding to each node are improved.

[0072] As an example, the preset graph embedding model is a preset Node2Vec. Node2Vec can include a random walk algorithm based on DFS and BFS and a skip-gram language model, and the skip-gram language model is a kind of preset word vector model. The Node2Vec is a network embedding algorithm. Taking a social network as an example, network embedding is to represent the nodes in the network with a low-dimensional vector, and these vectors should be able to reflect some characteristics of the original network. For example, if two nodes in the original network have a similar structure, then the vectors represented by these two nodes should also be similar. Specifically, the training process can include the following two processes:

[0073] 1. Biased random walk with edge weights

[0074] It is possible to start from each node in the co-occurrence graph and walk towards its neighbor nodes in a biased manner according to the edge weights, which is equivalent to random weighted sampling among the neighbor nodes. For example, preferentially walk to neighbor nodes with high edge weights to generate a text sequence of a preset length. Here, the preset length can refer to the preset number of nodes in the text sequence, which is not limited in this application. After the biased random walk, the text sequence corresponding to the node can be as Figure 6 shown.

[0075] Specifically, during the random walk of Node2Vec, the strategy for guiding the random walk can be BFS or DFS, and this application does not make any limitations in this regard. Among them, BFS is Breath First Search. Breath First Search is to select a node for traversal each time. When traversing, the unvisited adjacent nodes of the current node need to be added to a queue, and then the above traversal process is repeated for the adjacent nodes of the current node in turn until there are no nodes to be traversed in the queue. On this basis, in the embodiments of this specification, it is necessary to stop when the sequence obtained by the random walk reaches the preset length to obtain the text sequence corresponding to the node.

[0076] DFS is Depth First Search. Depth First Search is that every time a node is traversed, if this node has been traversed, then return, return to the upper layer (that is, the so-called backtracking); if the node meets the traversal conditions, add the current node to the traversed set, and then randomly select an adjacent node of this node, and repeat the above traversal process. Backtrack to the starting point until there are no other vertices to traverse. On this basis, in the embodiments of this specification, it is necessary to stop when the sequence obtained by the random walk reaches the preset length to obtain the text sequence corresponding to the node.

[0077] 2. Learning of the skip-gram language model

[0078] For each node in the biased random walk, map it to a representation vector, that is, perform vector mapping to form the representation vector of the given node. In the embodiments of this specification, the hidden layer of the skip-gram language model can be used to generate the feature vector of the node. For example, the text sequence corresponding to each node can be input into the skip-gram language model for vector representation processing to obtain the feature vector corresponding to each node. For example Figure 4b As shown, share the embedding vector of each entity word in the graph, and this embedding vector can be represented as R d , where d can be the vector dimension of the embedding vector, and this disclosure does not make any limitations on the vector dimension.

[0079] The skip-gram language model can maximize the probability that neighbor nodes appear in the random walk, and can be regarded as using the "center word" to predict the "context word". Take Figure 7For example, hierarchical Softmax can be used for binary classification, that is, to predict the probability of whether vj appears given vi, where vi is the vector representation of the i-th node in the co-occurrence graph, and vj is the vector representation of the j-th node in the co-occurrence graph. Based on this, the central vector in the feature vector can be input into hierarchical Softmax to obtain the prediction probability of the context vector. Thus, the prediction probability can be continuously iteratively optimized until the iterative termination condition is met. In one example, the iterative termination condition may refer to the prediction probability satisfying a probability threshold, which is not limited in this application. Correspondingly, Node2Vec when the iterative termination condition is met can be used as the target graph embedding model.

[0080] For Figure 7 example, when the corpus in the Q&A text data in the medical business field is short, such as in the medical field. The context window can be set to 1 (that is, one word on the left and right of the node) to improve the training efficiency. For example, the text sequence u v4 corresponding to the entity word 4 (W k = 4) for random walk: 4, 3, 1, 5, 7, where 1 is the central word, and 3 and 5 are context words. The corresponding feature vectors can be v4, v3, v1, v5, v7; the vector dimensions of v4, v3, v1, v5, v7 can be d. Skip-gram predicts its context vectors v3 and v5 through the central vector v1 (Φ(v1)), so as to maximize the probability of v3 and v5 under the condition of v1 until the iterative termination condition is met, then the target skip-gram language model can be obtained, and then the target graph embedding model can be obtained.

[0081] It should be noted that in the case where the edges in the co-occurrence graph do not have edge weights, random walks can also be performed in the co-occurrence graph starting from each node based on a preset node random walk algorithm to obtain the text sequences corresponding to each node. Furthermore, the skip-gram language model can also be trained according to the text sequences to obtain the target graph embedding model. The target graph embedding model can include the skip-gram language model or can include the preset node random walk algorithm and the skip-gram language model. The word vectors generated by the above target graph embedding model can be the output of the hidden layer of the skip-gram language model in the target graph embedding model, which is not limited in this disclosure.

[0082] Figure 8 FIG. shows a block diagram of a data processing device according to an embodiment of the present application. As Figure 8 shown, the device may include:

[0083] An acquisition module 801, configured to acquire a set of entity words to be processed;

[0084] A vector representation module 803, configured to perform vector representation processing on the entity words in the entity word set based on a target graph embedding model to obtain word vectors of each entity word; the graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, and the co-occurrence graph is constructed based on entity words respectively extracted from question texts, reply texts, and description texts of reply accounts included in multiple sample Q&A text corpora in the medical business domain, and the description text is used to indicate the sub-business domain corresponding to the reply account in the medical business domain; edges in the co-occurrence graph have edge weights, and the edge weight of each edge is obtained based on at least one of the co-occurrence probability of the nodes connected by each edge, the node type, and the semantic similarity information between the nodes, the nodes are the entity words, and the node type represents a preset dimension belonging to the corresponding entity word and used to describe the medical business domain;

[0085] An entity word pair acquisition module 805, configured to acquire entity word pairs with similar meanings in the entity word set according to the similarity between the word vectors.

[0086] In a possible implementation manner, the apparatus may further include:

[0087] A Q&A text corpus acquisition module, configured to acquire the multiple sample Q&A text corpora from the Q&A text data associated with the medical business domain, and each sample Q&A text corpus includes a question text, a reply text, and a description text of a reply account that feedbacks the reply text;

[0088] An entity word extraction module, configured to extract a first entity word from the question text, extract a second entity word from the reply text, and extract a third entity word from the description text;

[0089] A co-occurrence graph construction module, configured to use the first entity word, the second entity word, and the third entity word as nodes, set edges between the first entity word and the second entity word in the same sample Q&A text corpus, set edges between the first entity word and the third entity word in the same sample Q&A text corpus, and set edge weights for each edge to construct the co-occurrence graph.

[0090] In a possible implementation manner, the above co-occurrence graph construction module may include:

[0091] A first node graph acquisition unit, configured to use the first entity word, the second entity word, and the third entity word as nodes, set edges between the first entity word and the second entity word in the same sample Q&A text corpus, and set edges between the first entity word and the third entity word in the same sample Q&A text corpus to obtain a first node graph;

[0092] An edge number acquisition unit, configured to acquire the number of edges corresponding to each node in the first node graph;

[0093] A co-occurrence graph acquisition unit, configured to delete nodes with the number of edges greater than a threshold and corresponding edges in the first node graph to obtain a second node graph;

[0094] Set edge weights for the edges in the second node graph to obtain the co-occurrence graph.

[0095] In a possible implementation manner, the edges in the co-occurrence graph have edge weights; the apparatus may further include:

[0096] A target edge-connected node acquisition module, configured to acquire a first target node and a second target node connected by a target edge, where the target edge is any edge in the co-occurrence graph;

[0097] A co-occurrence probability and node type determination module, configured to determine the co-occurrence probability of the first target node and the second target node and their respective node types;

[0098] A first weight determination module, configured to determine a first weight of the target edge according to the co-occurrence probability and the node type;

[0099] A semantic similarity information acquisition module, configured to acquire semantic similarity information between the first target node and the second target node;

[0100] A second weight determination module, configured to determine a second weight of the target edge according to the semantic similarity information;

[0101] An edge weight acquisition module, configured to obtain the edge weight of the target edge based on the first weight and / or the second weight.

[0102] In a possible implementation manner, the preset graph embedding model includes a node random walk model and a preset word vector model, and the apparatus may further include:

[0103] A node random walk module, configured to use each node in the co-occurrence graph as a starting point, and perform weighted random walks to adjacent nodes based on the node random walk model and the edge weights to obtain a text sequence corresponding to each node;

[0104] A training module, configured to train the preset word vector model according to the text sequence to obtain the target graph embedding model.

[0105] In a possible implementation manner, the above-mentioned Q&A text corpus acquisition module may include:

[0106] An initial text corpus acquisition unit, configured to acquire a plurality of initial sample Q&A text corpora from the Q&A text data associated with the medical business domain;

[0107] A determination unit, configured to determine the content ratio information between the question text and the answer text in each initial sample Q&A text corpus, and / or the number of views of each initial sample Q&A text corpus;

[0108] A text corpus filtering unit, configured to filter out the sample Q&A text corpora with the content ratio information greater than a ratio threshold and / or the number of views less than a number threshold from the multiple initial sample Q&A text corpora, to obtain the multiple sample Q&A text corpora.

[0109] In a possible implementation manner, the apparatus may further include:

[0110] A preset synonym information acquisition module, configured to acquire preset synonym information; the preset synonym information represents multiple preset synonym pairs;

[0111] A synonym update module, configured to update the preset synonym information according to the entity word pair to obtain target synonym information.

[0112] In a possible implementation manner, the apparatus may further include:

[0113] A to-be-searched keyword acquisition module, configured to acquire a to-be-searched keyword;

[0114] An associated search term acquisition module, configured to acquire synonyms of the to-be-searched keyword from the target synonym information as at least one associated search term;

[0115] A target text data acquisition module, configured to perform text data search processing according to the to-be-searched keyword and the associated search terms to obtain target text data.

[0116] Regarding the apparatus in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0117] On the other hand, the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing methods provided in the above various optional implementation manners.

[0118] Figure 9 The block diagram of an electronic device for data processing according to an embodiment of the present application is shown. The electronic device may be a server, and its internal structure diagram may be as Figure 9As shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a data processing method.

[0119] Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0120] In an exemplary embodiment, there is also provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the data processing method in the embodiment of this application.

[0121] In an exemplary embodiment, there is also provided a storage medium. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the data processing method in the embodiment of this application.

[0122] In an exemplary embodiment, there is also provided a computer program product containing instructions. When it runs on a computer, the computer executes the data processing method in the embodiment of this application.

[0123] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0124] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only to be regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0125] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a set of entity words to be processed; Performing vector representation processing on the entity words in the set of entity words based on a target graph embedding model to obtain word vectors of each entity word; the target graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, and the co-occurrence graph is constructed based on entity words respectively extracted from question texts, reply texts, and description texts of reply accounts included in multiple sample Q&A text corpora in the medical business domain, and the description text is used to indicate the sub-business domain corresponding to the reply account in the medical business domain; the edges in the co-occurrence graph have edge weights, and the edge weight of each edge is obtained based on at least one of the co-occurrence probability of the nodes connected by each edge, the node type, and the semantic similarity information between the nodes, the nodes are the entity words, and the node type represents a preset dimension for describing the medical business domain to which the corresponding entity word belongs; Obtaining entity word pairs with similar meanings in the set of entity words according to the similarity between the word vectors; Wherein, the construction of the co-occurrence graph includes: obtaining the multiple sample Q&A text corpora from the Q&A text data associated with the medical business domain, and each sample Q&A text corpus includes a question text, a reply text, and a description text of a reply account that feedbacks the reply text; extracting first entity words from the question text, second entity words from the reply text, and third entity words from the description text; using the first entity words, the second entity words, and the third entity words as nodes, setting edges between the first entity words and the second entity words in the same sample Q&A text corpus, setting edges between the first entity words and the third entity words in the same sample Q&A text corpus, and setting edge weights for each edge to construct the co-occurrence graph.

2. The method according to claim 1, characterized in that, The step of using the first entity words, the second entity words, and the third entity words as nodes, setting edges between the first entity words and the second entity words in the same sample Q&A text corpus, setting edges between the first entity words and the third entity words in the same sample Q&A text corpus, and setting edge weights for each edge to construct the co-occurrence graph includes: Using the first entity words, the second entity words, and the third entity words as nodes, setting edges between the first entity words and the second entity words in the same sample Q&A text corpus, and setting edges between the first entity words and the third entity words in the same sample Q&A text corpus to obtain a first node graph; Obtaining the number of edges corresponding to each node in the first node graph; Deleting the nodes with the number of edges greater than a threshold and the corresponding edges in the first node graph to obtain a second node graph; Setting edge weights for the edges in the second node graph to obtain the co-occurrence graph.

3. The method according to claim 1 or 2, characterized in that, The setting of edge weights includes: Obtaining a first target node and a second target node connected by a target edge, where the target edge is any edge in the co-occurrence graph; Determining the co-occurrence probability of the first target node and the second target node and their respective node types; Determining a first weight of the target edge according to the co-occurrence probability and the node type; Obtain the semantic similarity information between the first target node and the second target node; Determine the second weight of the target edge according to the semantic similarity information; Obtain the edge weight of the target edge based on the first weight and / or the second weight.

4. The method according to claim 3, characterized in that, The preset graph embedding model includes a node random walk model and a preset word vector model, and the method further includes: Taking each node in the co-occurrence graph as a starting point, performing weighted random walks to adjacent nodes based on the node random walk model and the edge weight, and obtaining a text sequence corresponding to each node; Training the preset word vector model according to the text sequence to obtain the target graph embedding model.

5. The method according to claim 1, characterized in that, The obtaining the multiple sample Q&A text corpora from the Q&A text data associated with the medical business field includes: Obtain multiple initial sample Q&A text corpora from the Q&A text data associated with the medical business field; Determine the content ratio information between the question text and the answer text in each initial sample Q&A text corpus, and / or the number of views of each initial sample Q&A text corpus; Filter out the sample Q&A text corpora with the content ratio information greater than the ratio threshold and / or the number of views less than the number threshold from the multiple initial sample Q&A text corpora to obtain the multiple sample Q&A text corpora.

6. The method according to claim 1, characterized in that, The method further includes: Obtain preset synonym information; the preset synonym information represents multiple preset synonym pairs; Update the preset synonym information according to the entity word pair to obtain the target synonym information.

7. The method according to claim 6, characterized in that, The method further includes: Obtain the keyword to be searched; Obtain the synonyms of the keyword to be searched from the target synonym information as at least one associated search term; Perform text data search processing according to the keyword to be searched and the associated search terms to obtain the target text data.

8. A data processing device, characterized in that, Includes: An obtaining module, configured to obtain a set of entity words to be processed; A vector representation module, configured to perform vector representation processing on the entity words in the set of entity words based on the target graph embedding model to obtain word vectors of the entity words; the target graph embedding model is obtained by training a preset graph embedding model according to a co-occurrence graph, the co-occurrence graph is constructed based on entity words respectively extracted from question texts, answer texts, and description texts of answer accounts included in multiple sample Q&A text corpora in the medical business field, the description text is used to indicate the sub-business field corresponding to the answer account in the medical business field; the edges in the co-occurrence graph have edge weights, and the edge weight of each edge is obtained based on at least one of the co-occurrence probability of the nodes connected by each edge, the node type, and the semantic similarity information between the nodes, the nodes are the entity words, and the node type represents a preset dimension for describing the medical business field to which the corresponding entity word belongs; An entity word pair obtaining module, configured to obtain entity word pairs with similar meanings in the set of entity words according to the similarity between the word vectors; Wherein, the apparatus further includes: A question-and-answer text corpus acquisition module for obtaining the multiple sample question-and-answer text corpora from the question-and-answer text data associated with the medical business domain, where each sample question-and-answer text corpus includes a question text, an answer text, and a description text of the answer account that provides feedback on the answer text; An entity word extraction module for extracting first entity words from the question text, second entity words from the answer text, and third entity words from the description text; A co-occurrence graph construction module for using the first entity words, the second entity words, and the third entity words as nodes, setting edges between the first entity words and the second entity words in the same sample question-and-answer text corpus, setting edges between the first entity words and the third entity words in the same sample question-and-answer text corpus, and setting edge weights for each edge to construct the co-occurrence graph.

9. The device according to claim 8, characterized in that, The co-occurrence graph construction module includes: A first node graph acquisition unit for using the first entity words, the second entity words, and the third entity words as nodes, setting edges between the first entity words and the second entity words in the same sample question-and-answer text corpus, and setting edges between the first entity words and the third entity words in the same sample question-and-answer text corpus to obtain a first node graph; An edge number acquisition unit for obtaining the number of edges corresponding to each node in the first node graph; A co-occurrence graph acquisition unit for deleting the nodes in the first node graph with the number of edges greater than a threshold and the corresponding edges to obtain a second node graph; setting edge weights for the edges in the second node graph to obtain the co-occurrence graph.

10. The device according to claim 8 or 9, characterized in that, The device further includes: A target edge connected node acquisition module for obtaining a first target node and a second target node connected by a target edge, where the target edge is any edge in the co-occurrence graph; A co-occurrence probability and node type determination module for determining the co-occurrence probability of the first target node and the second target node and their respective node types; A first weight determination module for determining a first weight of the target edge according to the co-occurrence probability and the node type; A semantic similarity information acquisition module for obtaining semantic similarity information between the first target node and the second target node; A second weight determination module for determining a second weight of the target edge according to the semantic similarity information; an edge weight acquisition module for obtaining the edge weight of the target edge based on the first weight and / or the second weight.

11. The device according to claim 10, characterized in that, The preset graph embedding model includes a node random walk model and a preset word vector model, and the device further includes: A node random walk module for using each node in the co-occurrence graph as a starting point and performing weighted random walks to adjacent nodes based on the node random walk model and the edge weights to obtain a text sequence corresponding to each node; A training module for training the preset word vector model according to the text sequence to obtain the target graph embedding model.

12. The device according to claim 8, characterized in that, The question-and-answer text corpus acquisition module includes: An initial text corpus acquisition unit for obtaining multiple initial sample question-and-answer text corpora from the question-and-answer text data associated with the medical business domain; A determination unit, configured to determine the content ratio information between the question text and the answer text in each initial sample Q&A text corpus, and / or the number of views of each initial sample Q&A text corpus; A text corpus filtering unit, configured to filter out the sample Q&A text corpora with the content ratio information greater than a ratio threshold and / or the number of views less than a number threshold from the multiple initial sample Q&A text corpora, to obtain the multiple sample Q&A text corpora.

13. The device according to claim 8, characterized in that, The apparatus further includes: A preset synonym information acquisition module, configured to acquire preset synonym information; the preset synonym information represents multiple preset synonym pairs; A synonym update module, configured to update the preset synonym information according to the entity word pairs to obtain target synonym information.

14. The device according to claim 13, characterized in that, The apparatus further includes: A to-be-searched keyword acquisition module, configured to acquire a to-be-searched keyword; An associated search term acquisition module, configured to acquire synonyms of the to-be-searched keyword from the target synonym information as at least one associated search term; A target text data acquisition module, configured to perform text data search processing according to the to-be-searched keyword and the associated search terms to obtain target text data.

15. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the method according to any one of claims 1 to 7.

16. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

17. A computer program product, characterized in that, Including computer instructions, when the computer instructions are executed by the processor, the computer is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Automatic question and answer method and apparatus, and storage medium

    CN108304437A

  • Synonym mining method and device for question and answer retrieval system

    CN110442760A