Problem solution matching method and device, electronic equipment and storage medium
By constructing network diagrams and similarity calculations, high similarity and matching problem solutions are recommended, which solves the problem of insufficient combination of scientific research problems and solutions in the existing technology, and promotes scientific research innovation and interdisciplinary cooperation.
Patent Information
- Application Number
- CN202510420078.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
It is difficult to find deep semantic correlation scientific research problems and solution combinations in existing scientific and technological papers, and fail to explore the corresponding solutions to scientific research problems outside the literature library, lacking recommendations for novel scientific research problems and interdisciplinary research inspiration.
By constructing a network diagram, we calculate the similarity and matching degree of the pending scientific research problems and the existing problem types, we recommend high similarity and matching problem solutions, and use pre-trained general information extraction models and similarity calculation models such as SimBert, and combine them with LightGCN algorithm for correlation mining.
It has achieved accurate solutions for scientific research problems that do not exist in the network diagram, and promoted scientific research innovation and interdisciplinary research cooperation.
Smart Images

Figure CN120336478A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, this application relates to a method, apparatus, electronic device, and storage medium for matching problem solutions. Background Art
[0002] A scientific and technological paper is essentially a technical document about the solution to a scientific research problem. The scientific research problem is the task or goal that needs to be solved in scientific research, and the solution is a solution to the corresponding scientific research problem that includes elements such as technologies, models, and theories. The present invention extracts the texts representing scientific research problems and solutions in scientific and technological papers, and uses a recommendation algorithm to discover new combinations of existing scientific research problems and existing solutions, and scientific research problem-solution combinations that include new scientific research problems or new solutions. The discovered scientific research problem-solution combinations are called novel scientific research problem solutions.
[0003] Existing research on the extraction of scientific research problem-solutions in scientific and technological papers usually identifies key knowledge units such as scientific research problems and solutions from the perspective of word granularity or sentence granularity. However, the existing solutions are difficult to discover scientific research problem-solution combinations with deep semantic associations, and do not consider scientific research problems or solutions that do not exist in the existing literature. Mining solutions for scientific research problems outside the literature database, and even scientific research problems not studied in the scientific research field, may be more valuable in research and application, and can better inspire the inspiration of scientific research personnel and promote research progress. Therefore, it is necessary to study better recommendation algorithms to achieve the mining of solutions corresponding to more and more novel scientific research problems. Summary of the Invention
[0004] The purpose of this application aims to solve at least one of the above technical defects. The technical solutions provided by the embodiments of this application are as follows: In a first aspect, an embodiment of this application provides a method for matching problem solutions, including: Obtain a scientific research problem to be processed and the corresponding first problem type. If it is determined that there is no node representing the first problem type in the preset first network graph, determine the similarity between the first problem type and each second problem type. The first network graph includes a plurality of first nodes and a plurality of second nodes. Each first node represents a second problem type, and each second node represents a problem solution; for a pair of first nodes and second nodes that are connected in the first network graph, the connection indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes; Use the first nodes with a similarity not less than the first preset threshold as the first candidate nodes. For each first candidate node, use the second nodes that are connected to the first candidate node in the first network graph as the reference nodes of the first candidate node; For the reference node of each first candidate node, based on the similarity between the second problem type corresponding to the reference node and the first problem type of the first candidate node, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed; Use the problem solutions represented by the reference nodes with the first matching degree not less than the second preset threshold as the problem solutions of the scientific research problem to be processed.
[0005] In a second aspect, an embodiment of the present application provides a problem solution matching device, including: A first similarity determination module, configured to obtain a scientific research problem to be processed and a corresponding first problem type. If it is determined that there is no node representing the first problem type in a preset first network diagram, determine the similarity between the first problem type and each second problem type. The first network diagram includes multiple first nodes and multiple second nodes, each first node represents a second problem type, and each second node represents a problem solution; for a pair of first nodes and second nodes with a connection in the first network diagram, the connection indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes; A candidate node determination module, configured to use the first nodes with the similarity not less than the first preset threshold as the first candidate nodes. For each first candidate node, use the second nodes with a connection to the first candidate node in the first network diagram as the reference nodes of the first candidate nodes; A matching degree determination module, configured to, for the reference node of each first candidate node, based on the similarity between the second problem type corresponding to the reference node and the first problem type of the first candidate node, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed; A solution matching module, configured to use the problem solutions represented by the reference nodes with the first matching degree not less than the second preset threshold as the problem solutions of the scientific research problem to be processed.
[0006] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory; The processor executes the computer program to implement the method provided in the first aspect embodiment or any optional embodiment of the first aspect.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the first aspect embodiment or any optional embodiment of the first aspect is implemented.
[0008] The beneficial effects brought by the technical solution provided by the embodiment of the present application are: The solution provided by the embodiment of the present application, when obtaining a to-be-processed scientific research problem that does not exist in the first network diagram, calculates the similarity between the first problem type of the to-be-processed scientific research problem and each second problem type already existing in the first network diagram, and combines the corresponding relationships between each problem solution already existing in the first network diagram and each second problem type to determine the matching degree of each problem solution with the to-be-processed scientific research problem. Finally, one or more problem solutions with high matching degrees are used as the solutions to the to-be-processed scientific research problem. Through the above solution, it is possible to accurately recommend problem-solving solutions for scientific research problems that do not exist in the first network diagram, assist scientific research personnel in scientific research innovation, and also establish problem-solution associations across different disciplines to promote interdisciplinary research cooperation. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description of the embodiments of the present application.
[0010] Figure 1 It is a schematic flowchart of a problem solution matching method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of extracting scientific research problems and problem solutions in an example of an embodiment of the present application; Figure 3 It is a schematic flowchart of matching a to-be-processed scientific research problem in an example of an embodiment of the present application; Figure 4 It is a schematic flowchart of further mining the relationship between scientific research problems and problem solutions in an example of an embodiment of the present application; Figure 5 It is a schematic flowchart of classifying scientific research problems in an example of an embodiment of the present application; Figure 6 It is a structural block diagram of a problem solution matching device provided by an embodiment of the present application; Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] The following describes the embodiments of the present application with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0012] Those skilled in the art can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or their combinations supported by the technical field of the present application, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0013] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0014] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below by describing several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on or combine with each other. For the same terms, similar features and similar implementation steps in different embodiments, they will not be described repeatedly.
[0015] Figure 1 A flowchart of a problem solution matching method is provided for an embodiment of the present application. The execution subject of this method can be a terminal (such as a computer, a mobile phone, etc.), as Figure 1 shown, this method may include: Step S101, obtain the scientific research problem to be processed and the corresponding first problem type. If it is determined that there is no node representing the first problem type in the preset first network graph, then determine the similarity between the first problem type and each second problem type. The first network graph includes multiple first nodes and multiple second nodes. Each first node represents a second problem type, and each second node represents a problem solution; for a pair of first nodes and second nodes that are connected in the first network graph, the connection indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes.
[0016] In an embodiment of the present application, a scientific research problem may be a core question or goal to be solved in scientific research. The problem solution is a description of the process of how to solve the scientific research problem. The problem type may be the research direction, research field, or specific research goal of the scientific research problem, etc., which is not limited in the embodiments of the present application. The first network diagram may be a node relationship diagram, which includes multiple nodes and connections between the nodes. Each connection is used to connect one first node and one second node. The existing connected first node and second node indicate that the problem type represented by the first node can be solved by the problem solution represented by the second node. The similarity may represent the connection between the first problem type and the second problem type. Generally speaking, the higher the similarity, the closer the problem types represented by the first problem type and the second problem type are.
[0017] Specifically, the scientific research problem to be processed in the embodiments of the present application may be input by a user (such as a scientific researcher). After receiving the scientific research problem to be processed, semantic analysis may be performed on the scientific research problem to be processed first to determine the first problem type of the scientific research problem to be processed. The method of semantic analysis may be word segmentation or context feature extraction, etc., which is not limited in the embodiments of the present application. After obtaining the first problem type, compare the first problem type with the problem types (i.e., each second problem type) represented by each first node included in the pre-established first network diagram to determine whether there is a problem type in each second problem type that is the same as the first problem type. If there is a first problem type with the same type, the problem solutions represented by each second node that has a connection with the first node corresponding to the first problem type may be used as the solution to the scientific research problem to be processed. If there is no first problem type with the same type, it is necessary to dig out the corresponding problem solution from the first network diagram through further matching calculation steps. Since the corresponding relationship between each stored scientific research problem and each problem solution has been shown in the first network diagram, it is possible to consider starting from the association strength between each stored scientific research problem and the scientific research problem to be processed and indirectly exploring the relationship between the scientific research problem to be processed and each stored problem solution. For exploring the association strength between each stored scientific research problem and the scientific research problem to be processed, it can be measured by calculating the similarity between each stored scientific research problem and the scientific research problem to be processed respectively. Therefore, the similarity between each stored scientific research problem and the scientific research problem to be processed can be calculated first.
[0018] Step S102, use the first nodes with similarity not less than the first preset threshold as the first candidate nodes. For each first candidate node, use the second nodes that have a connection with the first candidate node in the first network diagram as the reference nodes of the first candidate node.
[0019] Specifically, after obtaining the similarity between the scientific research problems corresponding to each first node and the scientific research problem to be processed in step S101, each first node can be screened based on the similarity, that is, each first node with a similarity not less than the first preset threshold is used as a first candidate node. The first preset threshold can be set in advance. Not less than this threshold indicates a relatively high similarity between the two problem types. After determining each first candidate node, each second node with a connection to the first candidate node in the first network diagram can be used as a reference node, and the problem solutions represented by these reference nodes will be used as alternative problem solutions for the scientific research problem to be processed.
[0020] It should be noted that in the embodiments of the present application, the first preset threshold may not be set, but instead, the number of first candidate nodes can be set, that is, a certain number of first nodes with relatively high similarity are selected as first candidate nodes.
[0021] Step S103: For each reference node of each first candidate node, based on the similarity between the second problem type and the first problem type of the first candidate node corresponding to the reference node, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed.
[0022] In the embodiments of the present application, the first matching degree is used to represent whether the problem solution represented by the reference node can well solve the scientific research problem to be processed. Generally speaking, the larger the first matching degree, the better the problem solution represented by the reference node can solve the scientific research problem to be processed.
[0023] Specifically, since ultimately a suitable problem solution needs to be matched for the scientific research problem to be processed, it is also necessary to determine the matching degree between the scientific research problem to be processed and the problem solutions corresponding to each reference node. Specifically, since the second problem types represented by each first candidate node screened through the previous steps are already relatively close to the first problem type, and there is a corresponding relationship (i.e., a connection) between each second problem type and each reference node in the first network diagram, the relationship between the problem solution corresponding to each reference node and the second problem type represented by each first candidate node can be determined first according to the connection relationship, and then the determined relationship is combined with the similarity between the second problem type and the first problem type of the first candidate node corresponding to the reference node determined in step S102 to indirectly determine the first matching degree between the problem solution represented by the reference node and the scientific research problem to be processed.
[0024] Step S104: Use the problem solutions represented by the reference nodes with a first matching degree not less than the second preset threshold as the problem solutions for the scientific research problem to be processed.
[0025] In an embodiment of the present application, the second preset threshold may be preset, and not being less than this threshold indicates that the problem solution corresponding to the reference node can solve the scientific research problem to be processed.
[0026] Specifically, after determining the first matching degree between the problem solutions represented by each reference node and the scientific research problem to be processed, the problem solutions with the first matching degree not less than the second preset threshold can be taken as the solutions to the scientific research problem to be processed and output to the user.
[0027] It should be noted that in the embodiment of the present application, the second preset threshold may not be set, but instead, the number of reference nodes can be set, that is, the problem solutions represented by a certain number of reference nodes with a relatively high first matching degree are selected as the solutions to the scientific research problem to be processed.
[0028] The solution provided by the embodiment of the present application, when obtaining the scientific research problem to be processed that does not exist in the first network diagram, calculates the similarity between the first problem type of the scientific research problem to be processed and each second problem type existing in the first network diagram, and combines the corresponding relationships between each problem solution existing in the first network diagram and each second problem type to determine the matching degree between each problem solution and the scientific research problem to be processed. Finally, one or more problem solutions with a high matching degree are used as the solutions to the scientific research problem to be processed. Through the above solution, it is possible to accurately recommend problem solutions for scientific research problems that do not exist in the first network diagram, assist scientific research personnel in scientific research innovation, and also establish problem-solution associations in different disciplines to promote interdisciplinary research cooperation.
[0029] Based on the above various embodiments, as an alternative embodiment, the first network diagram is generated in the following manner: Obtain a plurality of scientific research documents from a preset database, input each scientific research document into a pre-trained general information extraction model, and obtain the scientific research problems solved by each scientific research document output by the pre-trained general information extraction model and the corresponding problem solutions; Perform word segmentation on each scientific research problem to obtain the first word segmentation result of each scientific research problem; perform word segmentation on each problem solution to obtain the second word segmentation result of each problem solution; Classify each scientific research problem based on each first word segmentation result to obtain a plurality of first sets, and classify each problem solution based on each second word segmentation result to obtain a plurality of second sets; wherein, each first set includes a plurality of scientific research problems of the same problem type, and each second set includes a plurality of problem solutions of the same type; Construct a second network graph based on the correspondence relationships between the first sets, the second sets, and the scientific research questions and the corresponding problem solutions; wherein, the second network graph includes a plurality of first nodes and a plurality of second nodes, and for a pair of first nodes and second nodes connected by a line in the second network graph, the line indicates that there is a correspondence relationship between the corresponding pair of first nodes and second nodes; For each first node in the second network graph, obtain at least one second candidate node that has not established a correspondence relationship with the first node from the second nodes included in the second network graph, determine the second matching degree between each second candidate node and the first node, and add a line between the first node and the preset number of second candidate nodes with the highest second matching degrees sorted from largest to smallest in the second network graph to obtain a first network graph.
[0030] In the embodiments of the present application, the preset database may be a database for storing various scientific research documents, such as a literature website, etc. The scientific research documents may be papers, patents, books, etc., and the embodiments of the present application do not limit this here. The general information extraction model is an AI (Artificial Intelligence) framework that uniformly processes multiple information extraction tasks, and is used in the embodiments of the present application to extract the scientific research questions and corresponding problem solutions included in the scientific research documents.
[0031] Specifically, the first nodes, the second nodes, and the lines included in the first network graph need to be constructed in advance, and the content represented by the first nodes and the second nodes needs to be first extracted from the scientific research documents. Therefore, as Figure 2 shown, a plurality of scientific research documents can be freely obtained from the preset database storing scientific research documents, and the obtained scientific research documents can come from different research fields, research objectives, or research directions, etc. After obtaining each scientific research document, all the obtained scientific research documents are input into the pre-trained general information extraction model, and the general information extraction model can extract the scientific research questions studied in each scientific research document and the corresponding problem solutions, obtaining a plurality of scientific research questions and a plurality of corresponding problem solutions. Optionally, since most documents will present the research questions and corresponding problem solutions in the abstract part, the abstract part of the scientific research documents can also be input into the model to improve the output efficiency.
[0032] Due to the large number of scientific research literatures, it is inevitable that there will be duplicate scientific research problems or duplicate problem solutions. When constructing a network diagram, considering the efficiency in subsequent matching calculations, the structure of the network diagram is often simplified as much as possible. Therefore, duplicate scientific research problems or problem solutions can be classified, and scientific research problems or problem solutions of the same category can be grouped into the same node in the network diagram. The classification method can be to perform semantic recognition on the text content of scientific research problems or solutions respectively, and group those with similar semantics into the same category. The semantic recognition method can be word segmentation. Perform word segmentation on scientific research problems to obtain the first word segmentation result, and perform word segmentation on problem solutions to obtain the second word segmentation result. Then, group the first word segmentation results with similar word segmentation results into the same category, with each category being the first set, and group the second word segmentation results with similar word segmentation results into the same category, with each category being the second set.
[0033] When constructing the network diagram, each first set is regarded as a first node, and each second set is also regarded as a second node. Then, according to the corresponding relationship between scientific research problems and solutions, connections are added between the first node and the second node to obtain the second network diagram. It can be understood that if the problem type corresponding to a first node contains multiple scientific research problems, and the multiple scientific research problems respectively correspond to problem solutions of different types, then connection relationships are set between the first node and all second nodes corresponding to different types of solutions. For example, the first node A contains scientific research problems X and Y, and the problem solution M corresponding to the scientific research problem X belongs to the second node B, and the problem solution N corresponding to the scientific research problem Y belongs to the second node C. Then, connection relationships are set between the first node A and the second node B and between the first node A and the second node C.
[0034] The second network diagram only contains the corresponding relationship between scientific research problems and corresponding problem solutions in the original scientific research literature. However, it is found that in many cases, the problem solutions corresponding to a scientific research problem type can also be used to solve another type of scientific research problem. Therefore, in the embodiments of this application, the corresponding relationship in this part will be further explored. Specifically, for each first node, first select at least one second candidate node that has no established relationship with the first node from each second node, and determine the second matching degree between each second candidate node and the first node. The second matching degree can be used to represent whether the problem solution corresponding to the second candidate node can solve the scientific research problem corresponding to the first node. Then, sort the second matching degrees from large to small, and obtain the top preset number of second nodes as the second candidate nodes. These second candidate nodes are the corresponding relationships mined with the first node. Then, add connections between these second candidate nodes and the first node. After all the connections of the first nodes are added, the first network diagram is obtained.
[0035] It should be noted that the pre-trained general information extraction model in the embodiments of the present application is obtained by fine-tuning an existing general information extraction model. The reason for fine-tuning is that the existing general information extraction model has strong versatility, but the specificity for extracting literature content is insufficient. Therefore, certain adjustments are needed. Specifically, a part of the literature can be selected first, and the text content containing scientific research questions and problem-solving solutions in the literature is marked (such as Figure 2 the data marking process in), and the marked literature is used as the training set. Then, the existing general information extraction model is trained with these training sets, and the model parameters of the model are adjusted according to the loss function value during the training process. After the training is completed, the pre-trained general information extraction model in the embodiments of the present application is obtained.
[0036] Based on the above embodiments, as an alternative embodiment, determining the similarity between the first problem type and each second problem type specifically includes: Performing text feature extraction on the first problem type to obtain a first vector corresponding to the first problem type, and performing text feature extraction on the second problem type to obtain a second vector corresponding to the second problem type; Determining the cosine similarity between the first vector and the second vector, and using the cosine similarity as the similarity.
[0037] Specifically, in the embodiments of the present application, the calculation of the similarity can be obtained through the SimBert (Similar Bidirectional Encoder Representations from Transformers) model. The advantage of this model is that it can enhance the semantic representation ability by simultaneously learning semantic similarity calculation, text generation, and retrieval tasks. Specifically, the text describing the first problem type and the text describing the second problem type can be input into the SimBert model together. The model can first perform text feature extraction on the two texts respectively by combining the context content of the texts to obtain a first vector and a second vector representing semantics, and then calculate the cosine similarity between the two vectors to obtain the similarity between the two semantics. This similarity can be used to represent the similarity between the first problem type and the second problem type.
[0038] Based on the above embodiments, as an alternative embodiment, based on the similarity between the second problem type of the first candidate node corresponding to the reference node and the first problem type, and whether there is a connection between the reference node and each first candidate node, determining the first matching degree between the reference node and the scientific research problem to be processed specifically includes: For each first candidate node, obtaining the association strength parameter of the first candidate node with respect to the reference node, and weighting the similarity based on the association strength parameter to obtain the associated similarity of the first candidate node; Sum the association similarities of each first candidate node to obtain the first matching degree; Among them, obtaining the association strength parameter of the first candidate node specifically includes: For each first candidate node, if there is a connection between the first candidate node and the reference node, then take the first parameter as the association strength parameter; if there is no connection between the first candidate node and the reference node, then take the second parameter as the association strength parameter; the first parameter is greater than the second parameter.
[0039] In the embodiments of the present application, the association strength parameter can represent whether there is a connection between the first candidate node and the reference node. The second parameter (such as 0) can be used to represent that there is no connection between the first candidate node and the reference node, and the first parameter (such as 1) can be used to represent that there is a connection between the first candidate node and the reference node. The association similarity can represent the matching degree between the reference node and the scientific research problem to be processed obtained through a certain first candidate node. The first matching degree can represent the matching degree between the reference node and the scientific research problem to be processed obtained through all first candidate nodes.
[0040] Specifically, as described above, in the embodiments of the present application, the matching degree between the reference node and the scientific research problem to be processed can be indirectly determined through the relationship between the reference node and the first candidate and the similarity between the first candidate node and the scientific research problem to be processed. Specifically, for each reference node, the association strength parameter of each first candidate node for the reference node can be first determined according to whether there is a connection between the reference node and each first candidate node in the first network diagram. For example, if there is a connection between the first candidate node and the reference node, the association strength parameter of the first candidate node relative to the reference node can be 1. On the contrary, if there is no connection between the first candidate node and the reference node, the association strength parameter of the first candidate node relative to the reference node can be 0; Next, for each first candidate node, the similarity between the first candidate node and the scientific research problem to be processed (that is, the similarity between the first problem type and the second problem type represented by the first candidate node) is weighted by the association strength parameter of the first candidate node, that is, the association similarity obtained by the reference node through the first candidate node is obtained, and then the association similarities are added up to form the first matching degree between the reference node and the scientific research problem to be processed. The specific formula expression can be as follows: =
[0041] Among them, represents the first matching degree, represents the similarity between the first candidate node and the scientific research problem to be processed, represents the association strength parameter.
[0042] The following combinesFigure 3 , the calculation process of the above first matching degree is specifically described. For example, Figure 3 as shown, the first network diagram includes four first nodes P1, P2, P3, P4 and five second nodes M1, M2, M3, M4, M5. The corresponding relationship between each first node and the second node is represented by the connection lines shown in the diagram. When a new problem P x (i.e., the scientific research problem to be processed) is obtained, and three problem solutions need to be matched for this new problem P x , then first, two first candidate nodes P2 and P4 are determined from the four first nodes P1, P2, P3, P4 (it can be that only two are selected according to actual needs, or P2 and P4 are screened out by setting a first preset threshold, that is, the first preset threshold is greater than S x3 and not greater than S x2 ). Then, it is found that in the original problem-solution network (i.e., the second network diagram), the second nodes corresponding to the first candidate node P2 (not marked in the diagram) are M2 and M4 (not marked in the diagram), and the second nodes corresponding to the first candidate node P4 (not marked in the diagram) are M1 and M3 (not marked in the diagram). Moreover, through the method provided by the embodiments of the present application, the corresponding relationships between the first node P2 and the second node M5 and between the first node P4 and the second node M2 are further mined (i.e., the mining results of new combinations of existing problems and solutions in the diagram). Since all five second nodes in the diagram have corresponding relationships with the first candidate nodes, these five second nodes are all used as reference nodes. According to the above relationships, the first matching degree of each reference node can be calculated. For example, for the reference node M1, since there is only a connection line between this reference node and the first candidate node P4, the first matching degree of M1 is S x4 (the similarity between the first candidate node P4 and the new problem Px) 1 (there is a connection line relationship between M1 and P4, so the correlation strength parameter is 1) + S x2 (the similarity between the first candidate node P2 and the new problem P x ) 0 (there is no connection line relationship between M1 and P2, so the correlation strength parameter is 0) = S x4 ; the calculation methods of the remaining reference nodes are similar to the above method and will not be elaborated here. Finally, according to the magnitude relationship of the first matching degrees, the problem solutions represented by the three reference nodes M1, M2, and M3 with the largest first matching degrees are taken as the problem solutions of the new problem P x .
[0043] Based on the above various embodiments, as an optional embodiment, determining the second matching degree between each second candidate node and the first node specifically includes: Determine the degree of association between the problem solution corresponding to each second candidate node and the problem type corresponding to the first node based on the corresponding relationships represented by the connections included in the first network diagram; Obtain the years when the solutions corresponding to the second nodes appear in the preset database, and determine the novelty degree of the problem solutions corresponding to each second candidate node according to the years of the second nodes; Determine the proposed frequency of the problem solution corresponding to each second candidate node based on the number of connections included in the first network diagram; Perform weighted summation on each score respectively based on the first preset weight of the degree of association, the second preset weight of the novelty degree, and the third preset weight of the proposed frequency to obtain the second matching degree.
[0044] In the embodiments of the present application, the degree of association can represent the association strength between the problem solution represented by the second candidate node and the problem type represented by the first node (for example, whether the technical fields studied are the same, whether the research goals are the same or similar, etc.). The novelty degree is used to represent the gap between the year when the problem solution corresponding to the second candidate node is proposed and the current year. The proposed frequency can represent the reliability of the problem solution. Generally speaking, the higher the proposed frequency of the solution, the higher the corresponding reliability.
[0045] Specifically, in the embodiments of the present application, the matching degree between the second candidate node and the first node will be measured from three aspects. The first aspect can be from the perspective of the association strength. For example, whether the technical problem fields solved by the problem solutions represented by the second candidate node are the same or similar to the technical problem fields of the problem types represented by the first node, etc. The more similar, the higher the degree of association. The second aspect can be from the perspective of the novelty degree. Generally speaking, the closer the year when the problem solution is newly proposed to the current year, the higher the novelty degree of the problem solution. The third aspect can be from the perspective of the proposed frequency. The higher the proposed frequency, the more times or scenarios the problem solution is often applied, and the feasibility and practicality have been verified by a large number of experiments. Then, the above three aspects are weighted and summed in turn through the preset weights, and the second matching degree between each second candidate node and the first node is obtained through comprehensive consideration. Specifically, it can be reflected by the following formula:
[0046] Among them, can represent the second matching degree, can represent the degree of association, can represent the novelty degree, can represent the proposed frequency, 、 、 Then the corresponding preset weights can be customized according to actual needs.
[0047] Based on the above various embodiments, as an alternative embodiment, the degree of association between the problem solution corresponding to each second candidate node and the first node is determined based on the corresponding relationships represented by the respective connections included in the first network diagram, specifically including: Extract the features of the types of problem solutions corresponding to each second node where there is a connection to the first node in the first network diagram to obtain the third vector of the first node; For each second candidate node, extract the features of the second problem types of each first node where there is a connection to the second candidate node in the first network diagram to obtain the fourth vector of the second candidate node, and perform an inner product operation on the third vector and the fourth vector to determine the degree of association between the first node and the second candidate node; Sort the degrees of association between each second candidate node and the first node from largest to smallest to determine the ranking of each second candidate node; For each second candidate node, use the ranking of the second candidate node as the input of the first preset function to obtain the degree of association of the second candidate node; where the first preset function is a decreasing function, and when the ranking is greater than the preset ranking, the function value of the first preset function is less than zero.
[0048] In the embodiments of the present application, the preset ranking is numerically the same as the preset number of second candidate nodes that need to add connections as described above. For example, when 10 second candidate nodes need to be selected to set connections with the first node, the preset ranking here is 10.
[0049] Specifically, in the embodiments of the present application, the degree of association between the problem solution corresponding to the second candidate node and the problem type of the first node can be calculated based on the LightGCN (Lightweight Graph Convolutional Network) algorithm. LightGCN has the characteristic of being able to mine the high-order connectivity in the graph and can better capture the collaborative signals (in the embodiments of the present application, it is the information hidden in the first node - second node connection, and this information can reveal the similarity between the first nodes or the second nodes), achieving a good recommendation effect. Therefore, this algorithm is selected as the new combination recommendation algorithm for problems and solutions. In the calculation process, first determine the vectors corresponding to the first node and each second candidate node (i.e., the third vector and the fourth vector), and these vectors are used to represent the existing connection relationships of each node. Then, perform an inner product operation on the third vector and each fourth vector pairwise. The inner product operation can be used to measure the consistency of the directions of two vectors. In the embodiments of the present application, it is used to measure the consistency between the problem research direction of the problem type and the problem-solving direction of the problem solution. Generally speaking, the larger the inner product, the higher the consistency, and the greater the corresponding degree of association.
[0050] Since only a part of the second candidate nodes need to be selected to add connections to the first node, after obtaining the association degrees between each second candidate node and the first node, it is necessary to sort each second candidate node based on the association degree first to obtain the ranking of each second candidate node, and then input the ranking into the first preset function. The function value output by the first preset function is the association degree of each second candidate node. The design of the first preset function is as follows:
[0051] Among them, represents the preset ranking, represents the ranking of each second candidate node. By observing the first preset function, it can be understood that this function is a decreasing function (that is, decreases as increases), and when is less than , the value of will be less than zero, which can intuitively reflect that the association degree between the second candidate node after the preset ranking and the first node is relatively low.
[0052] Based on the above various embodiments, as an alternative embodiment, obtain the years when the solutions corresponding to each second node appear in the preset database, and determine the novelty of the problem solution corresponding to each second candidate node for the first node according to the years of each second node, specifically including: For each second candidate node, obtain the latest year when the type of problem solution corresponding to the second candidate node appears in the preset database, and use the latest year as the input of the second preset function to obtain the novelty of the second candidate node; among them, the second preset function is an increasing function.
[0053] Specifically, in the embodiments of the present application, the calculation of novelty is mainly based on the latest proposed time of the problem solution. Specifically, since a type of problem solution may include multiple specific problem solutions, and the years when each problem solution is proposed are usually different, therefore, in the embodiments of the present application, for the type of problem solution represented by each second candidate node, only select the latest year among the years when each specific problem solution included in this type is proposed as the latest year, and use this latest year as the input of the second preset function. The function value output by the second preset function is the novelty of each second candidate node. The design of the second preset function is as follows:
[0054] Among them, represents the latest year of the second candidate node, Represents the earliest year in which all the problem-solving solutions included in the problem-solving solution types corresponding to all the second nodes in the first network diagram were proposed. Represents the latest year in which all the problem-solving solutions included in the problem-solving solution types corresponding to all the second nodes in the first network diagram were proposed. By observing the second preset function, it can be understood that this function is an increasing function (i.e., as increases and increases), at this time the larger it is (i.e., the later the proposed time), the greater the corresponding novelty degree.
[0055] It should be noted that in the embodiments of the present application, only from the perspective of the first node, a method for finding new relationships for each first node from the second nodes that have not established relationships is described in detail. And the method for finding new relationships for each second node from the first nodes that have not established relationships from the perspective of the second node also uses similar technical means. Although not elaborated herein in the embodiments of the present application, this part of the content should also be regarded as the protection scope of the present invention.
[0056] Based on the above various embodiments, as an optional embodiment, the proposed frequency of each second candidate node with respect to the first node is determined based on the number of connections included in the first network diagram, specifically including: Obtain the number of connections of each second candidate node in the first network diagram; For each second candidate node, use the number of connections as the input of the third preset function to obtain the proposed frequency of the second candidate node; wherein, the third preset function is an increasing function.
[0057] Specifically, in the embodiments of the present application, the calculation of the proposed frequency is mainly based on the number of problem types solved by each problem-solving solution type in the first network diagram. Specifically, for some problem types, although several problem types are different, they can all be solved by the same problem-solving solution type. Based on this, the embodiments of the present application can use the number of existing connections of each second candidate node as the number of times the second candidate node is proposed, and use this number of times as the input of the third preset function. The function value output by the third preset function is the proposed frequency of each second candidate node. The design of the third preset function is as follows:
[0058] wherein, represents the number of existing connections of the second candidate node, Denote the number of connections with the largest number of connections among the second nodes in the first network diagram. In the embodiments of the present application, the number of connections of each second candidate node is compared with the largest number of connections of the second nodes in the first network diagram, and the actual proposed frequency is calculated according to the comparison difference. By observing the third preset function, it can be understood that this function is an increasing function (i.e., increases as increases), and at this time the larger it is (i.e., the more connections the second candidate node has), the larger the corresponding proposed frequency is.
[0059] Next, in combination with Figure 4 , the construction process of the first network diagram provided by the present application will be introduced. As Figure 4 shown, first, a plurality of first nodes and a plurality of second nodes are generated according to a plurality of first sets and second sets (the light color in the figure represents the first nodes, and the dark color represents the second nodes). Then, according to each scientific research literature, the corresponding relationship between the problem types corresponding to each first node and the problem solution types corresponding to each second node is determined, and each node is organized and connected lines are set to obtain a second network diagram (i.e., the problem-solution network in the figure). Next, for each first node, a second node without a connected line is recommended for matching. First, each second candidate node is determined from each second node of each first node, and then the association degree of each corresponding second candidate node is calculated for each first node through a recommendation algorithm (the LightGCN algorithm corresponding to the embodiments of the present application) , and the novelty degree of each corresponding second candidate node is calculated for each first node according to the method described above in the embodiments of the present application and the proposed frequency , and then the sorting mechanism is improved by means of preset weights. According to the final sorting result, a new connection relationship is established for each first node (i.e., the new combination of existing problems and solutions in the figure).
[0060] Based on the above various embodiments, as an optional embodiment, the first word segmentation result includes word segments belonging to keywords; Based on each first word segmentation result, each scientific research problem is classified to obtain a plurality of first sets, specifically including: Based on each first word segmentation result, each scientific research problem is classified to obtain a plurality of first sets, specifically including: For two first word segmentation results with at least one same keyword word segment, each word segment belonging to the keyword in the two first word segmentation results and a preset prompt word are input into the large language model together. The large language model is instructed by the preset prompt word to match the similarity of the two first word segmentation results, and the matching result of the two first word segmentation results output by the large language model is obtained; wherein, the matching result is matching or not matching; If the matching result is a match, the scientific research questions corresponding to the two first word segmentation results are classified into the same first set.
[0061] In the embodiments of the present application, the large language model can be an existing general language model, such as ChatGPT (Chat Generative Pretrained Transformer), etc., or an improved model obtained by slightly improving the existing language model (such as the ERNIE (Enhanced Representation through Knowledge Integration) pre-trained model), etc. The embodiments of the present application do not make any limitations here. The preset prompt words can be keywords containing "matching" or similar synonyms, such as "Match A and B" or "Compare A and B", etc., for instructing the large model language to perform a matching operation.
[0062] Specifically, in order to classify different scientific research questions, such as Figure 5 As shown, the embodiments of the present application need to first perform a word segmentation operation on each scientific research question. The specific word segmentation operation can be carried out through a preset lexical analysis tool (such as the THULAC (THU Lexical Analyzer for Chinese) tool, etc.). That is, the text content describing the scientific research question is input into the lexical analysis tool, and the lexical analysis tool will decompose the text content into multiple word segments and analyze the part of speech of each word segment (such as nouns, verbs, adjectives, adverbs, etc.). Moreover, for the words representing the key parts in the text content, these words will be marked together (marked as keywords), and then the above content will be output as the first word segmentation result.
[0063] After obtaining the keywords in each scientific research question, these keywords can be used as the central words corresponding to the scientific research questions. To make the matching results more accurate, generally only two central words corresponding to scientific research questions are input into the large language model at one time. However, when there are many scientific research questions, inputting only two central words of scientific research questions each time will result in too low comparison efficiency. According to practical experience, if all the central words contained in the text content corresponding to two scientific research questions are different, then these two scientific research questions are almost impossible to be classified into the same type. Based on this, to improve the comparison efficiency, after determining the central words of each scientific research question, the central words of each scientific research question can be initially compared pairwise. If it is found that there is at least one identical central word in two scientific research questions, then it can be considered that there may be similarities between the scientific research questions corresponding to these two first-word segmentation results. At this time, these central words are input into the large language model and further analyzed and matched through the large language model. The large language model will directly output the result of whether the central words of the two scientific research questions match or do not match. If the output result is a matching result, then these two scientific research questions can be classified into the same type (i.e., classified into the same first set). It can be understood that if there is a scientific research question that cannot find any other scientific research question that can be successfully matched, then a first set can be set separately for this scientific research question, and this first set only includes this scientific research question.
[0064] Optionally, if there is no labeled keyword segmentation in the first-word segmentation result output by the lexical analysis tool, then the words with the part of speech of nouns can be used as the central words; if there are no nouns in the first-word segmentation result, then all the remaining words are used as the central words together.
[0065] It should be noted that the above methods for word segmentation and classification of each scientific research question can also be applied to problem solutions, and the corresponding results are the second-word segmentation results and the second set, which will not be elaborated here.
[0066] Figure 6 The structural block diagram of a problem solution matching device provided by an embodiment of the present application is as Figure 6 shown. The problem solution matching device 600 may include: a first similarity determination module 601, a candidate node determination module 602, a matching degree determination module 603, and a solution matching module 604, where The first similarity determination module 601 is used to obtain the scientific research problem to be processed and the corresponding first problem type. If it is determined that there is no node representing the first problem type in the preset first network diagram, the similarity between the first problem type and each second problem type is determined. The first network diagram includes a plurality of first nodes and a plurality of second nodes. Each first node represents a second problem type, and each second node represents a problem solution; for a pair of first nodes and second nodes that are connected in the first network diagram, the connection indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes. The candidate node determination module 602 is used to use the first nodes with a similarity not less than the first preset threshold as the first candidate nodes. For each first candidate node, the second nodes that are connected to the first candidate node in the first network diagram are used as the reference nodes of the first candidate node. The matching degree determination module 603 is used to, for the reference node of each first candidate node, based on the similarity between the second problem type of the first candidate node corresponding to the reference node and the first problem type, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed. The solution matching module 604 is used to use the problem solutions represented by the reference nodes with a first matching degree not less than the second preset threshold as the problem solutions of the scientific research problem to be processed. The solution provided by this application, when obtaining a scientific research problem that does not exist in the first network diagram, calculates the similarity between the first problem type of the scientific research problem to be processed and each existing second problem type in the first network diagram, and combines the corresponding relationship between each existing problem solution in the first network diagram and each second problem type to determine the matching degree between each problem solution and the scientific research problem to be processed. Finally, one or more problem solutions with a high matching degree are used as the solutions to the scientific research problem to be processed. Through the above solution, it is possible to accurately recommend problem-solving solutions for scientific research problems that do not exist in the first network diagram, assist scientific research personnel in scientific research innovation, and also establish problem-solution associations in different disciplines to promote cross-disciplinary research cooperation.
[0067] Based on the above various embodiments, as an optional embodiment, the device further includes a network diagram generation module, which is specifically used for: Obtain a plurality of scientific research documents from a preset database, input each scientific research document into a pre-trained general information extraction model, and obtain the scientific research problems solved by each scientific research document output by the pre-trained general information extraction model and the corresponding problem solutions of the scientific research problems; Perform word segmentation processing on each scientific research problem to obtain the first word segmentation result of each scientific research problem; perform word segmentation processing on each problem solution to obtain the second word segmentation result of each problem solution; Classify each scientific research question based on each first word segmentation result to obtain multiple first sets, and classify each problem solution based on each second word segmentation result to obtain multiple second sets; wherein, each first set includes multiple scientific research questions of the same problem type, and each second set includes multiple problem solutions of the same type; Construct a second network graph based on each first set, each second set, and the corresponding relationship between each scientific research question and each problem solution; wherein, the second network graph includes multiple first nodes and multiple second nodes, and for a pair of first nodes and second nodes connected by a line in the second network graph, the line indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes; For each first node in the second network graph, obtain at least one second candidate node that has not established a corresponding relationship with the first node from each second node included in the second network graph, determine the second matching degree between each second candidate node and the first node, and add a line between the first node and the preset number of second candidate nodes with the highest second matching degree sorted from large to small in the second network graph to obtain a first network graph.
[0068] Based on the above various embodiments, as an alternative embodiment, the first similarity determination module is specifically configured to: Extract text features from the first problem type to obtain a first vector corresponding to the first problem type, and extract text features from the second problem type to obtain a second vector corresponding to the second problem type; Determine the cosine similarity between the first vector and the second vector, and use the cosine similarity as the similarity.
[0069] Based on the above various embodiments, as an alternative embodiment, the matching degree determination module is specifically configured to: For each first candidate node, obtain the association strength parameter of the first candidate node with respect to the reference node, and weight the similarity based on the association strength parameter to obtain the association similarity of the first candidate node; Sum the association similarities of each first candidate node to obtain the first matching degree; Wherein, obtaining the association strength parameter of the first candidate node specifically includes: For each first candidate node, if there is a line between the first candidate node and the reference node, use the first parameter as the association strength parameter; if there is no line between the first candidate node and the reference node, use the second parameter as the association strength parameter; the first parameter is greater than the second parameter.
[0070] Based on the above various embodiments, as an alternative embodiment, the network graph generation module is further configured to: Determine the degree of association between the problem solution corresponding to each second candidate node and the problem type corresponding to the first node based on the corresponding relationships represented by the connections included in the first network diagram; Obtain the years in which the solutions corresponding to the second nodes appear in the preset database, and determine the novelty of the problem solutions corresponding to each second candidate node according to the years of the second nodes; Determine the proposed frequency of the problem solution corresponding to each second candidate node based on the number of connections included in the first network diagram; Perform weighted summation on each score respectively based on the first preset weight of the degree of association, the second preset weight of the novelty, and the third preset weight of the proposed frequency to obtain the second matching degree.
[0071] Based on the above various embodiments, as an alternative embodiment, the network diagram generation module can also be used for: Extract the features of the types of problem-solving solutions corresponding to the second nodes where the first node has connections in the first network diagram to obtain the third vector of the first node; For each second candidate node, extract the features of the second problem types of the first nodes where the second candidate node has connections in the first network diagram to obtain the fourth vector of the second candidate node, and perform an inner product operation on the third vector and the fourth vector to determine the degree of association between the first node and the second candidate node; Sort the degrees of association between each second candidate node and the first node from large to small to determine the ranking of each second candidate node; For each second candidate node, use the ranking of the second candidate node as the input of the first preset function to obtain the degree of association of the second candidate node; wherein, the first preset function is a decreasing function, and when the ranking is greater than the preset ranking, the function value of the first preset function is less than zero.
[0072] Based on the above various embodiments, as an alternative embodiment, the network diagram generation module can also be used for: For each second candidate node, obtain the latest year in which the type of problem solution corresponding to the second candidate node appears in the preset database, and use the latest year as the input of the second preset function to obtain the novelty of the second candidate node; wherein, the second preset function is an increasing function.
[0073] Based on the above various embodiments, as an alternative embodiment, the network diagram generation module can also be used for: Obtain the number of connections of each second candidate node in the first network diagram; For each second candidate node, use the number of connections as the input of the third preset function to obtain the proposed frequency of the second candidate node; wherein, the third preset function is an increasing function.
[0074] Based on the above various embodiments, as an alternative embodiment, the first word segmentation results include word segments belonging to keywords; The network diagram generation module can also be used for: Classifying each scientific research question based on each first word segmentation result to obtain a plurality of first sets, specifically including: Classifying each scientific research question based on each first word segmentation result to obtain a plurality of first sets, specifically including: For two first word segmentation results with at least one same keyword word segment, input each word segment belonging to the keyword and a preset prompt word in the two first word segmentation results into the large language model, and use the preset prompt word to instruct the large language model to match the similarity of the two first word segmentation results, and obtain the matching result of the two first word segmentation results output by the large language model; wherein, the matching result is matching or not matching; If the matching result is matching, classify the scientific research questions corresponding to the two first word segmentation results into the same first set.
[0075] The following refers to Figure 7 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server that executes the method shown in Figure 1 ) suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0076] The electronic device includes: a memory and a processor. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. Among them, the processor here can be referred to as the processing device 701 described below, and the memory may include at least one of the read-only memory (ROM) 702, random access memory (RAM) 703, and storage device 708 described below, as specifically shown below: As Figure 7As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0077] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0078] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present application are executed.
[0079] It should be noted that the above-mentioned computer-readable storage medium in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In this application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0080] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (for example, the Internet), and end-to-end networks (for example, ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0081] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately and not be assembled into the electronic device.
[0082] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to: Obtain the scientific research problem to be processed and the corresponding first problem type. If it is determined that there is no node representing the first problem type in the preset first network diagram, determine the similarity between the first problem type and each second problem type. The first network diagram includes multiple first nodes and multiple second nodes. Each first node represents a second problem type, and each second node represents a problem solution; for a pair of first nodes and second nodes that are connected in the first network diagram, the connection indicates that there is a corresponding relationship between the corresponding pair of first nodes and second nodes; use the first nodes with similarity not less than the first preset threshold as the first candidate nodes. For each first candidate node, use the second node that is connected to the first candidate node in the first network diagram as the reference node of the first candidate node; for the reference node of each first candidate node, based on the similarity between the second problem type of the first candidate node corresponding to the reference node and the first problem type, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed; use the problem solutions represented by the reference nodes with the first matching degree not less than the second preset threshold as the problem solutions of the scientific research problem to be processed.
[0083] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0085] The modules or units involved in the embodiments described in the present application can be implemented in software or in hardware. Among them, the name of a module or unit does not, in some cases, constitute a limitation on the unit itself. For example, the first constraint acquisition module can also be described as "the module for acquiring the first constraint".
[0086] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0087] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash Memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0088] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this text, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0089] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A problem solution matching method, characterized in that, include: Obtaining a scientific research problem to be processed and a corresponding first problem type, and if it is determined that there is no node representing the first problem type in a preset first network diagram, determining the similarity between the first problem type and each second problem type, wherein the first network diagram includes a plurality of first nodes and a plurality of second nodes, each first node represents a second problem type, and each second node represents a problem solution; For a pair of first nodes and a second node in the first network graph that have a connection line, the connection line indicates that there is a corresponding relationship between the corresponding pair of first nodes and the second node; Taking the first node whose similarity is not less than the first preset threshold as the first candidate node, and for each first candidate node, taking the second node in the first network graph that is connected to the first candidate node as the reference node of the first candidate node; For each reference node of the first candidate node, based on the similarity between the second question type of the first candidate node corresponding to the reference node and the first question type, and whether there is a connection between the reference node and each first candidate node, determine the first matching degree between the reference node and the scientific research problem to be processed; The problem-solving method represented by each reference node whose first matching degree is not less than a second preset threshold is used as the problem-solving method for the scientific research problem to be processed.
2. The method according to claim 1, wherein The first network diagram is generated by: Acquire multiple scientific research documents from a preset database, input each scientific research document into a pre-trained general information extraction model, and obtain the scientific research problems solved by each scientific research document output by the pre-trained general information extraction model and the problem solutions corresponding to the scientific research problems; Perform word segmentation processing on each scientific research question to obtain the first word segmentation result of each scientific research question; Perform word segmentation processing on each problem solution to obtain a second word segmentation result of each problem solution; Classifying each scientific research problem based on each first word segmentation result to obtain a plurality of first sets, and classifying each problem solution based on each second word segmentation result to obtain a plurality of second sets; wherein each first set includes a plurality of scientific research problems of the same problem type, and each second set includes a plurality of problem solutions of the same type; Based on the corresponding relationship between each first set, each second set, and each scientific research problem and each problem solution, a second network diagram is constructed; wherein the second network diagram includes a plurality of first nodes and a plurality of second nodes, and for a pair of first nodes and a second node with a connection in the second network diagram, a corresponding relationship exists between the pair of first nodes and the second node corresponding to the connection mark; For each first node in the second network diagram, at least one second candidate node that has no corresponding relationship with the first node is obtained from the second nodes included in the second network diagram, the second matching degree of each second candidate node with the first node is determined, and the connection lines are added between a preset number of second candidate nodes that are ranked in descending order by the second matching degree and the first node in the second network diagram to obtain the first network diagram.
3. The method according to claim 1, characterized in that, Determining the similarity between the first problem type and each second problem type includes: Performing text feature extraction on the first problem type to obtain a first vector corresponding to the first problem type, and performing text feature extraction on the second problem type to obtain a second vector corresponding to the second problem type; Determining the cosine similarity between the first vector and the second vector, and using the cosine similarity as the similarity.
4. The method according to claim 1, characterized in that Based on the similarity between the second problem type of the first candidate node corresponding to the reference node and the first problem type, and whether there is a connection between the reference node and each first candidate node, determining the first matching degree between the reference node and the scientific research problem to be processed includes: For each first candidate node, obtaining the association strength parameter of the first candidate node with respect to the reference node, and weighting the similarity based on the association strength parameter to obtain the association similarity of the first candidate node; Summing up the association similarities of each first candidate node to obtain the first matching degree; Among them, obtaining the association strength parameter of the first candidate node includes: For each first candidate node, if there is a connection between the first candidate node and the reference node, using the first parameter as the association strength parameter; if there is no connection between the first candidate node and the reference node, using the second parameter as the association strength parameter; the first parameter is greater than the second parameter.
5. The method according to claim 2, wherein Determining the second matching degree between each second candidate node and the first node includes: Based on the corresponding relationships represented by the connections included in the first network diagram, determining the degree of association between the problem solution corresponding to each second candidate node and the problem type corresponding to the first node; Obtaining the years in which the solutions corresponding to each second node appear in the preset database, and determining the novelty degree of the problem solutions corresponding to each second candidate node according to the years of each second node; Based on the number of connections included in the first network diagram, determining the proposed frequency of the problem solutions corresponding to each second candidate node; Based on the first preset weight of the degree of association, the second preset weight of the novelty degree, and the third preset weight of the proposed frequency, respectively performing weighted summation on each score to obtain the second matching degree.
6. The method according to claim 5, wherein Based on the corresponding relationships represented by the connections included in the first network diagram, determining the degree of association between the problem solution corresponding to each second candidate node and the first node includes: Performing feature extraction on the types of problem-solving solutions corresponding to each second node where the first node has a connection in the first network diagram to obtain a third vector of the first node; For each second candidate node, performing feature extraction on the second problem types of each first node where the second candidate node has a connection in the first network diagram to obtain a fourth vector of the second candidate node, and performing an inner product operation on the third vector and the fourth vector to determine the degree of association between the first node and the second candidate node; Sort the association degrees of each second candidate node with the first node from large to small to determine the ranking of each second candidate node; For each second candidate node, use the ranking of the second candidate node as the input of the first preset function to obtain the association degree of the second candidate node; wherein, the first preset function is a decreasing function, and when the ranking is greater than the preset ranking, the function value of the first preset function is less than zero.
7. The method according to claim 5, wherein The method for obtaining the years when the solutions corresponding to each second node appear in the preset database and determining the novelty degree of the problem solution corresponding to each second candidate node for the first node based on the years of each second node includes: For each second candidate node, obtain the latest year when the problem solution type corresponding to the second candidate node appears in the preset database, and use the latest year as the input of the second preset function to obtain the novelty degree of the second candidate node; wherein, the second preset function is an increasing function.
8. The method according to claim 5, wherein The method for determining the proposed frequency of each second candidate node with respect to the first node based on the number of connections included in the first network diagram includes: Obtain the number of connections of each second candidate node in the first network diagram; For each second candidate node, use the number of connections as the input of the third preset function to obtain the proposed frequency of the second candidate node; wherein, the third preset function is an increasing function.
9. The method according to claim 2, wherein The first word segmentation result includes word segments belonging to keywords; The method for classifying each scientific research problem based on each first word segmentation result to obtain multiple first sets includes: For two first word segmentation results with at least one identical keyword word segment, input the keyword word segments belonging to the two first word segmentation results and a preset prompt word into a large language model, and use the preset prompt word to instruct the large language model to match the similarity of the two first word segmentation results, and obtain the matching result of the two first word segmentation results output by the large language model; wherein, the matching result is match or no match; If the matching result is match, classify the scientific research problems corresponding to the two first word segmentation results into the same first set.
10. A problem solution matching device, characterized in that, It includes: A first similarity determination module, configured to obtain a scientific research problem to be processed and a corresponding first problem type, and if it is determined that there is no node representing the first problem type in a preset first network diagram, determine the similarity between the first problem type and each second problem type, the first network diagram includes multiple first nodes and multiple second nodes, each first node represents a second problem type, and each second node represents a problem solution; For a pair of first node and second node with a connection in the first network diagram, the connection indicates that there is a corresponding relationship between the corresponding pair of first node and second node; A candidate node determination module, configured to use the first nodes with a similarity not less than a first preset threshold as first candidate nodes, and for each first candidate node, use the second nodes connected to the first candidate node in the first network diagram as the reference nodes of the first candidate node; A matching degree determination module, configured to, for each reference node of the first candidate nodes, determine a first matching degree between the reference node and the scientific research problem to be processed based on the similarity between the second problem type of the first candidate nodes corresponding to the reference node and the first problem type, and whether there is a connection between the reference node and each first candidate node; A solution matching module, configured to use the problem solutions represented by the reference nodes with the first matching degree not less than a second preset threshold as the problem solutions of the scientific research problem to be processed.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-9 is implemented.