Information processing method and apparatus, and electronic device and computer readable storage medium
By constructing and analyzing heterogeneous networks between science, technology and industries, extracting and visualizing the skeleton structure, the problem of inaccurate data analysis results in the existing technology is solved, and accurate analysis and accurate prediction of industrial development trends are achieved.
Patent Information
- Application Number
- PCT/CN2024/133909
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-08
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-12
AI Technical Summary
In the information processing and data mining between science, technology and industries, it is difficult to accurately analyze and accurately predict the future development trend of the industry.
By obtaining object information of multiple objects associated with the target field, determining the reference relationship between multiple objects, building a heterogeneous network, and extracting the skeleton structure of the heterogeneous network, visualizing the skeleton structure to determine the analysis results of the target field.
The triple-related detection between science, technology and industries has been realized, data mining is carried out from multiple dimensions, and the development trajectory of the industry is deeply explored, the accuracy of data analysis results is improved, and the prediction accuracy of the future development trend of the industry is improved.
Smart Images

Figure CN2024133909_12062025_PF_FP_ABST
Abstract
Description
Information processing method, device, electronic device, and computer-readable storage medium Technical Field
[0001] The present application relates to the field of data mining technology, and more specifically, to an information processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] One of the key characteristics of contemporary scientific and technological development is the deepening integration of science, technology, and industrial development. Emerging and future industries rely heavily on scientific and technological advances, while industrial development, in turn, drives the advancement of scientific research and technological innovation. This creates a complex interactive pattern among science, technology, and industry. In-depth analysis of existing data on these science, technology, and industry will help explore the trajectory and future direction of industry development.
[0003] In the existing technology, correlation detection is usually carried out based on information data between science, technology and industry to achieve information processing and data mining. However, there is a problem of inaccurate data analysis results, which cannot accurately predict the future development trend of the industry. Summary of the Invention
[0004] The embodiments of the present application provide an information processing method, apparatus, electronic device, and computer-readable storage medium that can solve the problem of inaccurate data analysis results. The technical solution is as follows:
[0005] According to one aspect of an embodiment of the present application, there is provided an information processing method, the method comprising:
[0006] Obtaining object information of multiple objects associated with a target field; wherein the types of the multiple objects include papers, patents, and products;
[0007] Determining reference relationships between multiple objects based on object information; wherein the reference relationships include reference relationships between objects of different types and self-reference relationships between objects of the same type;
[0008] Construct a heterogeneous network based on the reference relationship and extract the skeleton structure of the heterogeneous network; the skeleton structure is used to represent the association relationship and interaction between multiple objects;
[0009] The skeleton structure is visualized so that the analysis results of the target area can be determined based on the visualized skeleton structure.
[0010] According to another aspect of an embodiment of the present application, there is provided an information processing device, the device comprising:
[0011] An acquisition module, configured to acquire object information of multiple objects associated with a target field; wherein the types of the multiple objects include papers, patents, and products;
[0012] A determination module, configured to determine reference relationships between multiple objects based on object information; wherein the reference relationships include reference relationships between objects of different types and reference relationships between objects of the same type;
[0013] The extraction module is used to construct a heterogeneous network based on the reference relationship and extract the skeleton structure of the heterogeneous network; wherein the skeleton structure is used to represent the association relationship and interaction between multiple objects;
[0014] The display module is used to visualize the skeleton structure so as to determine the analysis results of the target field based on the visualized skeleton structure.
[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, which includes: a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the method shown in the first aspect of the embodiment of the present application.
[0016] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method shown in the first aspect of the embodiments of the present application are implemented.
[0017] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps of the method shown in the first aspect of the embodiment of the present application are implemented.
[0018] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0019] The embodiment of the present application determines the reference relationship between multiple objects through the object information of multiple objects associated with the target field, then constructs a heterogeneous network based on the above reference relationship, extracts the skeleton structure of the heterogeneous network, and visualizes the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure. The embodiment of the present application realizes the ternary association detection between science (i.e., papers), technology (i.e., patents), and industry (i.e., products). Since the reference relationship can include the reference relationship between different types of objects and the self-reference relationship between the same type of objects, data mining can be performed from multiple dimensions. At the same time, the extracted skeleton structure can be used to characterize the association relationship and interaction between multiple objects, which is conducive to in-depth mining of the development trajectory of the industry; and, based on the visualized skeleton structure, it is possible to achieve accurate analysis of the target field, improve the accuracy of data analysis results, and improve the accuracy of prediction of future development trends of the industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0021] FIG1 is a schematic diagram of an application scenario of an information processing method provided by an embodiment of the present application;
[0022] FIG2 is a flow chart of an information processing method provided in an embodiment of the present application;
[0023] FIG3 is a schematic diagram of a visual skeleton structure in an information processing method provided in an embodiment of the present application;
[0024] FIG4 is a schematic diagram of a flowchart of another visualization skeleton structure in an information processing method provided in an embodiment of the present application;
[0025] FIG5 is a schematic diagram of another flow chart of a visualization skeleton structure in an information processing method provided in an embodiment of the present application;
[0026] FIG6 is a flow chart of an exemplary information processing method provided in an embodiment of the present application;
[0027] FIG7 is a schematic diagram of the structure of an information processing device provided in an embodiment of the present application;
[0028] FIG8 is a schematic structural diagram of an information processing electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0030] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B".
[0031] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0032] Optionally, embodiments of the present application may be applied to fields such as scientific and technological intelligence, data mining, and bibliometric analysis. For example, the multiple objects associated with the target field in the present application may be multiple entities in a complex network.
[0033] Complex networks are networks that exhibit self-organization, self-similarity, attractors, small-world properties, and scale-free nature. In the real world, numerous complex networks exist, consisting of entities connected to one another, such as social networks and computer networks. Research on complex networks primarily focuses on structural information such as path length, node density, and node connectivity within the network; as well as the community structure (local densely connected subgraphs or clusters of nodes) formed from the perspective of connectivity and node density. Community discovery primarily explores community structures and can be understood as clustering based on similar features or feature sets, helping us understand the inherent properties and functions of complex networks.
[0034] Complex networks can be divided into homogeneous networks and heterogeneous networks. Heterogeneous networks include not only multiple node types, but also complex relationships between different node types. The community structure in a heterogeneous network is generally defined as a community composed of nodes of the same type, with many-to-many relationships formed between communities.
[0035] Science is the theoretical study of the laws and nature of natural processes, matter, and life. Technology is the skills and methods people use to create technological artifacts through the use of tools, materials, and symbols in production practices. Industry, on the other hand, leverages science and technology to carry out large-scale production through specific organizational structures. Since the Industrial Revolution, the three have become increasingly interconnected, revealing the interrelationship between science, technology, and industry, which holds significant theoretical and practical value. However, existing technologies are still unable to accurately analyze and detect the interactive patterns among science, technology, and industry.
[0036] The information processing method, device, electronic device and computer-readable storage medium provided in this application are intended to solve the above technical problems in the prior art.
[0037] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0038] As shown in Figure 1, the information processing method of the present application can be applied to the scenario shown in Figure 1. Specifically, server 102 obtains object information of multiple objects associated with a target domain from a preset database, and then determines the reference relationships between the multiple objects based on the object information. Next, a heterogeneous network is constructed based on the reference relationships, the skeleton structure of the heterogeneous network is extracted, and the skeleton structure is visualized. Server 102 sends the visualized skeleton structure to client 101, and client 101 determines the analysis results of the target domain based on the visualized skeleton structure.
[0039] In the scenario shown in FIG1 , the above information processing method may be performed in a server, and in other scenarios, may also be performed in a terminal.
[0040] Those skilled in the art will understand that the “terminal” used here may be a mobile phone, a tablet computer, a PDA (Personal Digital Assistant), an MID (Mobile Internet Device), etc.; the “server” may be implemented as an independent server or a server cluster consisting of multiple servers.
[0041] An embodiment of the present application provides an information processing method, as shown in FIG2 , which can be applied to a server or terminal performing information processing. The method includes:
[0042] S201: Acquire object information of multiple objects associated with a target domain.
[0043] Among them, the types of multiple objects include papers, patents and products.
[0044] Taking the biomedical field as an example, we can retrieve papers from PubMed (a database that provides biomedical paper searches and abstracts) using the following search terms: ((treatment)AND(clinical)OR((effects)AND(efficacy)))AND(pharmacokinetics)AND(cancer)OR((metabolism)AND(cell)AND(action))AND((pharmacology)OR(acute)OR(chronic))OR(((Anti-inflammatory drugs)AND(Cardiovascular drugs))OR(Psychotropic drugs)AND(Vaccines)OR((Anticancer drugs)AND(Opioid drugs)AND(Analgesic drugs)))OR(((Immunomodulatory drugs)AND(Immunosuppressant drugs))OR((Antimalarial drugs)AND(Antineoplastic drugs)AND(Antidiabetic drugs) drugs), the publication time was limited to before December 31, 2019, and a total of 214,907 papers were retrieved. The object information of the papers can include: title, abstract, publication date and references, etc.
[0045] Patent documents were retrieved from the Derwent (Derwent Patent Database) database with the search formula: (TS=(treatment)AND TS=(clinical)AND(TS=(effects)AND TS=(efficacy)))OR TS=(pharmacokinetics)OR TS=(cancer)OR(TS=(metabolism)AND TS=(cell)AND TS=(action))AND(TS=(pharmacology)OR TS=(acute)OR TS=(chronic))OR((TS=(Anti-inflammatory drugs)OR TS=(Cardiovascular drugs))OR TS=(Psychotropic drugs)AND TS=(Vaccines)OR(TS=(Anticancer drugs)AND TS=(Opioid drugs)OR TS=(Analgesic drugs)))OR((TS=(Immunomodulatory drugs)AND TS=(Immunosuppressant drugs))OR(TS=(Antimalarial drugs)AND TS=(Antineoplastic drugs)OR TS=(Antidiabetic drugs))), with the publication time limited to before December 31, 2019, a total of 96,259 patents were retrieved, and the object information of the patents may include: title, abstract, application date and references, etc.
[0046] At the same time, XML (eXtensible Markup Language) data can be downloaded from the DrugBank database (a comprehensive database of drug and drug target information). A total of 13,339 drug products were extracted from the XML data. The object information of the drug products may include: synonyms, description information, reference information, etc.
[0047] S202: Determine reference relationships between multiple objects based on object information.
[0048] Citation relationships include citation relationships between objects of different types and self-citation relationships between objects of the same type. Citation relationships between objects of different types can include: patents citing papers, papers citing patents, papers citing products, products citing papers, products citing patents, and patents citing products. Self-citation relationships between objects of the same type can include: papers citing papers, patents citing patents, and so on.
[0049] Specifically, the object information of the paper can be processed in an automated manner, and the paper information in the references can be extracted to construct the citation relationship between the papers.
[0050] Next, a semi-automated construction method can be used to establish a citation relationship between the paper and the patent; the patent number can be automatically extracted from the object information of the paper (i.e., reference information), and a citation relationship between the paper and the patent can be constructed; at the same time, scientific non-patent literature information can be extracted from the object information of the patent (i.e., reference information), and then a citation relationship between the patent and the paper can be constructed.
[0051] At the same time, an automated construction method is used to process the object information of the patent (i.e., reference information), extract the patent information therein, and then construct the citation relationship between patents.
[0052] The citation relationship between the target paper and the target product can be constructed in the following way: first, the domain entity extraction method is used to extract the product entity from the title and abstract of the paper, and the citation relationship between the paper and the product is established based on the entity extraction results; second, the XML data downloaded from the DrugBank database is parsed, and the paper information in the reference is extracted in combination with the reference information of the drug product, and then the citation relationship between the product and the paper is constructed.
[0053] The citation relationship between patents and products can be constructed in the following ways: first, the domain entity extraction method is used to extract product entities from the title and abstract of the patent, and the citation relationship between the patent and the product is established based on the entity extraction results; second, the XML data downloaded from the DrugBank database is parsed, and the patent information in the reference is extracted in combination with the reference information of the drug product, and then the citation relationship between the product and the patent is constructed.
[0054] S203: construct a heterogeneous network based on the reference relationship and extract the skeleton structure of the heterogeneous network.
[0055] Among them, the skeleton structure is used to represent the association relationship and interaction between multiple objects.
[0056] Specifically, we can construct a heterogeneous citation network between science, technology, and industry based on the direction of knowledge flow in citation relationships (i.e., from cited objects to citing objects). The detailed steps for network construction and skeleton structure extraction are described in detail below.
[0057] S204 , visually displaying the skeleton structure so as to determine the analysis result of the target domain based on the visualized skeleton structure.
[0058] Among them, the analysis results can be used to characterize the development trajectory and trends of the target field.
[0059] Specifically, Pajek (a large-scale complex network analysis tool) software can be used to visualize the skeleton structure of the heterogeneous network.
[0060] In the embodiments of the present application, the visual skeleton structure is specifically described using Figures 3-5 as an example:
[0061] As shown in Figure 3, this demonstrates an industry model driven by both science and technology. Path 1's early stages primarily focused on basic research in cefaclor (paper IDs 2172 and 2173) and clarithromycin (paper IDs 513 and 2981). Later, the focus shifted to the development of related technologies. For example, patents with IDs 2763, 3109, 3300, and 3808 focus on the combined use of colchicine and macrolide antibiotics; patents with IDs 4610, 5748, 5396, and 721 focus on technologies related to colchicine administration methods and oral solution compositions. Path 2's early and mid-stages primarily focused on the development of methylphenidate sustained-release powders and aqueous suspensions (technology path 3251→614) and research on abuse-resistant amphetamine compounds (paper IDs 3060, 3059, and 385). Later, the focus was on the study of liridamole inhibiting human liver microsomal cytochrome p450 (paper IDs 10344, 10345, 10346) for the treatment of ADHD (Attention deficit and hyperactivity disorder) in children.
[0062] As shown in Figure 4, this demonstrates a science-driven industry model. Pathway 3 initiated macrocyclic quinoline drug technology for the treatment or prevention of hepatitis C virus infection (patent ID 3150). Based on this technology, basic research was conducted on protease inhibitors and combination inhibitors (research routes 9003 → 8188). Pathway 4 initially emphasized scientific research on omega-3 fatty acids (a type of long-chain, polyunsaturated fatty acid) (paper IDs 258, 257, and 3605). In the mid-term, technological development played a pivotal role. Patents IDs 3435, 3575, 3701, 3991, and 3980 focused on the development and improvement of methods for treating hypertriglyceridemia. Later research focused on dihexenylethyl (drug ID 7910), omega-3 acid ethyl (drug ID 8439), omega-3 fatty acids (drug ID 9238), and chamomile acid (drug ID 11776).
[0063] As shown in Figure 5, a technology-driven industry model is presented. Path 5 is based on the development of multi-dose drug delivery (patent ID 1205) and modified sustained-release formulations (patent ID 1575). Subsequently, the mid-term path focused on the development of methylphenidate sustained-release chewable tablets (technical route 3567→5678), and later on, dexmethylphenidate (drug ID 5822) was developed. Path 6 has a relatively complex technical route, a long technological innovation cycle, and multiple technical branches. The technology mainly focuses on compositions and methods for treating central nervous system-related diseases (technical route 1205→665), and finally, amantadine (drug ID 898) was developed.
[0064] The embodiment of the present application determines the reference relationship between multiple objects through the object information of multiple objects associated with the target field, then constructs a heterogeneous network based on the above reference relationship, extracts the skeleton structure of the heterogeneous network, and visualizes the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure. The embodiment of the present application realizes the ternary association detection between science (i.e., papers), technology (i.e., patents), and industry (i.e., products). Since the reference relationship can include the reference relationship between different types of objects and the self-reference relationship between the same type of objects, data mining can be performed from multiple dimensions. At the same time, the extracted skeleton structure can be used to characterize the association relationship and interaction between multiple objects, which is conducive to in-depth mining of the development trajectory of the industry; and, based on the visualized skeleton structure, it is possible to achieve accurate analysis of the target field, improve the accuracy of data analysis results, and improve the accuracy of prediction of future development trends of the industry.
[0065] In an embodiment of the present application, a possible implementation method is provided, wherein the above-mentioned construction of a heterogeneous network based on a reference relationship includes:
[0066] S301, determining the citing object, cited object, and citing direction corresponding to each citation relationship;
[0067] S302, using citing objects and cited objects as nodes, using citation directions as links between nodes, and constructing an initial heterogeneous network based on the nodes and links;
[0068] S303: Optimize the initial heterogeneous network to obtain a heterogeneous network.
[0069] In an embodiment of the present application, a possible implementation is provided, wherein the above-mentioned optimization of the initial heterogeneous network to obtain a heterogeneous network includes:
[0070] S401, extracting the closed-loop structure in the initial heterogeneous network.
[0071] Among them, the closed-loop structures in heterogeneous networks are divided into four categories:
[0072] (1) Contains both drug and paper nodes; (2) Contains both drug and patent nodes; (3) Contains both drug, paper, and patent nodes; (4) Contains only paper nodes.
[0073] S402: Filter out target links based on object information corresponding to nodes in the closed-loop structure.
[0074] Among them, the target link represents a link whose reference relationship does not conform to the reference logic.
[0075] S403: Delete the target link from the closed loop to obtain an optimized initial heterogeneous network, and use the optimized initial heterogeneous network as the heterogeneous network.
[0076] In the embodiments of this application, for the first three types of closed-loop structures, the temporal logical order of the reference relationships between nodes may be incorrect. Therefore, the embodiments of this application process the loops based on the time attributes corresponding to the nodes in the closed-loop structure. Because drug development must be based on scientific research or patented inventions, the specific processing logic for the loops is: the time attribute of the cited node is greater than or equal to the time attribute of the citing node, and links that do not conform to this logic are deleted.
[0077] For the fourth type of loop, we first checked for errors in the citations between papers within the closed loop structure. If so, we removed any incorrect citations. Secondly, for loops where all citations were correct, if papers within a closed loop had co-authors, their research could be roughly considered to be of the same type. Therefore, the entire closed loop structure can be considered as a whole, with new nodes constructed to replace the loop. In other words, each loop is considered a node. When other papers, patents, or products cite papers within the loop, they are considered to cite that node; and when papers within the loop cite other papers, patents, or products, they are considered to cite that node. After the aforementioned optimization preprocessing, the final heterogeneous network is obtained.
[0078] The present application provides a possible implementation method, wherein the above-mentioned extraction of the skeleton structure of the heterogeneous network includes:
[0079] S501: Obtain at least two subgraphs of a heterogeneous network.
[0080] Among them, a subgraph is an independent graph structure that is not connected to each other in a heterogeneous network.
[0081] S601, extracting the skeleton structure of the heterogeneous network based on the subgraph.
[0082] In an embodiment of the present application, a possible implementation method is provided, wherein the above-mentioned subgraph-based extraction of the skeleton structure of the heterogeneous network includes:
[0083] S601: Determine the number of nodes in each subgraph.
[0084] S602: When the number of nodes in the subgraph is greater than a preset threshold, the subgraph is used as a target subgraph.
[0085] The target subgraph may be a maximum weakly connected subgraph.
[0086] In the embodiment of the present application, the largest weakly connected subgraph can have a total of 41,200 edges and 16,147 nodes. Among them, patents citing patents have the most citation relationships, papers citing patents have the least citation relationships, papers have the most nodes, and pharmaceutical products have the least nodes.
[0087] S603, extracting the skeleton structure of the heterogeneous network from the target subgraph.
[0088] In another possible implementation, the above-mentioned extraction of the skeleton structure of the heterogeneous network from the target subgraph includes:
[0089] S701: Calculate the weight of each node and / or the weight of each link in the target subgraph.
[0090] The weight is used to represent the importance of the corresponding node or link in the target subgraph.
[0091] Specifically, node weight calculation methods include but are not limited to intermediacy, betweenness centrality, degree centrality, and closeness centrality. Link weight calculation methods include but are not limited to Search Path Count (SPC), Search Path Link Count (SPLC), and Search Path Node Pair (SPNP).
[0092] In an embodiment of the present application, the search path link counting method can be used to calculate the weights of links in the target subgraph, that is, the largest weakly connected subgraph.
[0093] S702 : Based on the weight of each node and / or the weight of each link, select a target link from the links in the target subgraph as a key route.
[0094] The key route can be one or more. Methods for determining the key route include, but are not limited to, selecting links with higher weights as key routes, selecting links between nodes with higher importance as key routes, selecting high-weight reference relationships of different reference types as key routes, and manually designating certain links as key routes.
[0095] S703: Extract the skeleton structure of the heterogeneous network according to the key routes.
[0096] Among them, the skeleton structure extraction method includes but is not limited to the main path analysis method, the path search method, etc.
[0097] An embodiment of the present application provides a possible implementation method for selecting a target link as a key route from links in a target subgraph based on the weight of each node and / or the weight of each link, including:
[0098] S801, when the weight of a node or link is greater than a preset weight threshold, the corresponding node or link is regarded as a key node or key link;
[0099] S802, taking the links between key nodes and key links as target routes;
[0100] S803: Filter out key routes from the target routes based on the type of the reference relationship corresponding to the target routes.
[0101] In the embodiment of the present application, the following 10 edges can be used as key routes: (1) the two links with the largest weight of patent citing papers, and the two links with the largest weight of papers citing patents; (2) the link with the largest weight among the other 6 citation relationships.
[0102] The embodiment of the present application selects high-weighted citation relationships of different citation types as key routes, which can fully explore the relationship between science, technology and industry, and ensure that the extracted skeleton structure can detect the science-technology-industry relationship.
[0103] To better understand the above information processing method, an example of the information processing method of the present application is described in detail below with reference to FIG6 . The method includes the following steps:
[0104] S901, obtaining object information of multiple entity objects such as papers, patents, and products associated with the target field.
[0105] S902: Determine reference relationships between multiple objects based on object information.
[0106] Citation relationships include citation relationships between objects of different types and self-citation relationships between objects of the same type. Citation relationships between objects of different types can include: patents citing papers, papers citing patents, papers citing products, products citing papers, products citing patents, and patents citing products. Self-citation relationships between objects of the same type can include: papers citing papers, patents citing patents, and so on.
[0107] S903, determining the citing object, cited object, and citing direction corresponding to each citing relationship; taking the citing object and cited object as nodes, and the citing direction as links between the nodes, and constructing an initial heterogeneous network based on the nodes and links.
[0108] S904, extracting the closed-loop structure in the initial heterogeneous network; filtering out target links based on object information corresponding to nodes in the closed-loop structure; deleting the target links from the closed-loop to obtain an optimized initial heterogeneous network, and using the optimized initial heterogeneous network as the heterogeneous network.
[0109] Among them, the target link represents a link whose reference relationship does not conform to the reference logic.
[0110] S905 , obtaining at least two subgraphs of the heterogeneous network; when the number of nodes in a subgraph is greater than a preset threshold, taking the subgraph as a target subgraph.
[0111] Among them, a subgraph is an independent graph structure that is not connected to each other in a heterogeneous network.
[0112] S906 , calculating the weight of each node and / or the weight of each link in the target subgraph; based on the weight of each node and / or the weight of each link, selecting a target link from the links in the target subgraph as a key route.
[0113] The weight is used to represent the importance of the corresponding node or link in the target subgraph.
[0114] S907: Extract the skeleton structure of the heterogeneous network based on the key routes.
[0115] The skeleton structure is used to represent the relationship and interaction between multiple objects. Methods for extracting the skeleton structure include but are not limited to main path analysis and path search.
[0116] S908 , visually displaying the skeleton structure so as to determine the analysis result of the target domain based on the visualized skeleton structure.
[0117] The embodiment of the present application determines the reference relationship between multiple objects through the object information of multiple objects associated with the target field, then constructs a heterogeneous network based on the above reference relationship, extracts the skeleton structure of the heterogeneous network, and visualizes the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure. The embodiment of the present application realizes the ternary association detection between science (i.e., papers), technology (i.e., patents), and industry (i.e., products). Since the reference relationship can include the reference relationship between different types of objects and the self-reference relationship between the same type of objects, data mining can be performed from multiple dimensions. At the same time, the extracted skeleton structure can be used to characterize the association relationship and interaction between multiple objects, which is conducive to in-depth mining of the development trajectory of the industry; and, based on the visualized skeleton structure, it is possible to achieve accurate analysis of the target field, improve the accuracy of data analysis results, and improve the accuracy of prediction of future development trends of the industry.
[0118] The embodiment of the present application provides an information processing device, as shown in FIG7 , the information processing device 70 may include: an acquisition module 701 , a determination module 702 , an extraction module 703 and a display module 604 ;
[0119] The acquisition module 701 is used to acquire object information of multiple objects associated with the target field; wherein the types of the multiple objects include papers, patents and products;
[0120] Determining module 702, for determining reference relationships between multiple objects based on object information; wherein the reference relationships include reference relationships between objects of different types and reference relationships between objects of the same type;
[0121] Extraction module 703, used to construct a heterogeneous network based on the reference relationship and extract the skeleton structure of the heterogeneous network; wherein the skeleton structure is used to represent the association relationship and interaction between multiple objects;
[0122] The display module 704 is used to visualize the skeleton structure so as to determine the analysis result of the target domain based on the visualized skeleton structure.
[0123] In an embodiment of the present application, a possible implementation is provided. When constructing a heterogeneous network based on a reference relationship, the extraction module 703 is configured to:
[0124] Determine the citing object, cited object, and citation direction for each citation relationship;
[0125] The citing objects and cited objects are regarded as nodes, and the citation directions are regarded as links between nodes. The initial heterogeneous network is constructed based on the nodes and links.
[0126] The initial heterogeneous network is optimized to obtain a heterogeneous network.
[0127] In an embodiment of the present application, a possible implementation is provided. When the extraction module 703 optimizes the initial heterogeneous network to obtain the heterogeneous network, it is configured to:
[0128] Extracting closed-loop structures from the initial heterogeneous network;
[0129] Filter out target links based on the object information corresponding to the nodes in the closed-loop structure; wherein the target links represent links whose reference relationships do not conform to the reference logic;
[0130] The target link is deleted from the closed loop to obtain an optimized initial heterogeneous network, and the optimized initial heterogeneous network is used as the heterogeneous network.
[0131] In an embodiment of the present application, a possible implementation is provided. When extracting the skeleton structure of a heterogeneous network, the extraction module 703 is configured to:
[0132] Obtain at least two subgraphs of the heterogeneous network; wherein the subgraphs are independent graph structures in the heterogeneous network that are not connected to each other;
[0133] Extracting skeleton structures of heterogeneous networks based on subgraphs.
[0134] In an embodiment of the present application, a possible implementation is provided. When extracting the skeleton structure of a heterogeneous network based on a subgraph, the extraction module 703 is configured to:
[0135] Determine the number of nodes in each subgraph;
[0136] When the number of nodes in a subgraph is greater than a preset threshold, the subgraph is used as the target subgraph;
[0137] Extracting the skeleton structure of heterogeneous networks from target subgraphs.
[0138] In an embodiment of the present application, a possible implementation is provided. When extracting the skeleton structure of the heterogeneous network from the target subgraph, the extraction module 703 is configured to:
[0139] Calculate the weight of each node and / or the weight of each link in the target subgraph; wherein the weight is used to represent the importance of the corresponding node or link in the target subgraph;
[0140] Based on the weight of each node and / or the weight of each link, select a target link from the links in the target subgraph as a key route;
[0141] Extract the skeleton structure of heterogeneous networks based on key routes.
[0142] In an embodiment of the present application, a possible implementation is provided. When the extraction module 703 selects a target link as a key route from links in the target subgraph based on the weight of each node and / or the weight of each link, the extraction module 703 is configured to:
[0143] When the weight of a node or link is greater than a preset weight threshold, the corresponding node or link is regarded as a key node or key link;
[0144] Use the links between key nodes and key links as target routes;
[0145] Based on the type of reference relationship corresponding to the target route, key routes are filtered out from each target route.
[0146] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, and will not be repeated here.
[0147] The embodiment of the present application determines the reference relationship between multiple objects through the object information of multiple objects associated with the target field, then constructs a heterogeneous network based on the above reference relationship, extracts the skeleton structure of the heterogeneous network, and visualizes the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure. The embodiment of the present application realizes the ternary association detection between science (i.e., papers), technology (i.e., patents), and industry (i.e., products). Since the reference relationship can include the reference relationship between different types of objects and the self-reference relationship between the same type of objects, data mining can be performed from multiple dimensions. At the same time, the extracted skeleton structure can be used to characterize the association relationship and interaction between multiple objects, which is conducive to in-depth mining of the development trajectory of the industry; and, based on the visualized skeleton structure, it is possible to achieve accurate analysis of the target field, improve the accuracy of data analysis results, and improve the accuracy of prediction of future development trends of the industry.
[0148] In an embodiment of the present application, an electronic device is provided, including a memory, a processor and a computer program stored on the memory, the processor executes the above-mentioned computer program to implement the steps of the information processing method, which can be achieved compared with the relevant technology: the embodiment of the present application determines the reference relationship between multiple objects through the object information of multiple objects associated with the target field, and then constructs a heterogeneous network based on the above-mentioned reference relationship, and extracts the skeleton structure of the heterogeneous network, and visualizes the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure. The embodiment of the present application realizes the ternary association detection between science (i.e., papers), technology (i.e., patents), and industry (i.e., products). Since the reference relationship can include the reference relationship between different types of objects and the self-reference relationship between objects of the same type, data mining can be performed from multiple dimensions. At the same time, the extracted skeleton structure can be used to characterize the association relationship and interaction between multiple objects, which is conducive to in-depth mining of the development trajectory of the industry; and, based on the visualized skeleton structure, it is possible to achieve accurate analysis of the target field, improve the accuracy of data analysis results, and improve the accuracy of prediction of future development trends of the industry.
[0149] In an optional embodiment, an electronic device is provided, as shown in FIG8 , wherein the electronic device 80 shown in FIG8 includes: a processor 801 and a memory 803. The processor 801 and the memory 803 are connected, such as via a bus 802. Optionally, the electronic device 80 may further include a transceiver 804, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 804 is not limited to one, and the structure of the electronic device 80 does not constitute a limitation on the embodiments of the present application.
[0150] The processor 801 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor 801 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0151] Bus 802 may include a path for transmitting information between the aforementioned components. Bus 802 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Bus 802 may be divided into an address bus, a data bus, a control bus, and so on. For ease of illustration, FIG8 shows only one thick line, but this does not indicate that there is only one bus or only one type of bus.
[0152] The memory 803 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.
[0153] The memory 803 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 801. The processor 801 is used to execute the computer program stored in the memory 803 to implement the steps shown in the above method embodiment.
[0154] The electronic devices include, but are not limited to, mobile terminals such as mobile phones, notebook computers, PADs, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0155] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0156] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, so that when the computer device executes the computer instructions, the following conditions are achieved:
[0157] Obtaining object information of multiple objects associated with a target field; wherein the types of the multiple objects include papers, patents, and products;
[0158] Determining reference relationships between multiple objects based on object information; wherein the reference relationships include reference relationships between objects of different types and self-reference relationships between objects of the same type;
[0159] Construct a heterogeneous network based on the reference relationship and extract the skeleton structure of the heterogeneous network; the skeleton structure is used to represent the association relationship and interaction between multiple objects;
[0160] The skeleton structure is visualized so that the analysis results of the target area can be determined based on the visualized skeleton structure.
[0161] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0162] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0163] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. An information processing method, wherein: include: Obtaining object information of multiple objects associated with a target field; wherein the types of the multiple objects include papers, patents, and products; Determining reference relationships between the multiple objects according to the object information; wherein the reference relationships include reference relationships between objects of different types and self-reference relationships between objects of the same type; Constructing a heterogeneous network based on the reference relationship, and extracting a skeleton structure of the heterogeneous network; wherein the skeleton structure is used to characterize the association relationship and interaction between the multiple objects; The skeleton structure is visualized so as to determine the analysis result of the target domain based on the visualized skeleton structure.
2. The method according to claim 1, wherein: The constructing a heterogeneous network based on the reference relationship includes: Determining the citing object, cited object and citing direction corresponding to each citing relationship; Taking the citing object and the cited object as nodes, taking the citing direction as the link between the nodes, and constructing an initial heterogeneous network according to the nodes and the links; The initial heterogeneous network is optimized to obtain the heterogeneous network.
3. The method according to claim 2, wherein: The step of optimizing the initial heterogeneous network to obtain the heterogeneous network includes: Extracting a closed-loop structure in the initial heterogeneous network; Filtering out target links based on object information corresponding to nodes in the closed-loop structure; wherein the target links represent links whose reference relationships do not conform to reference logic; The target link is deleted from the closed loop to obtain an optimized initial heterogeneous network, and the optimized initial heterogeneous network is used as the heterogeneous network.
4. The method according to claim 1, wherein: The extracting the skeleton structure of the heterogeneous network includes: Acquire at least two subgraphs of the heterogeneous network; wherein the subgraphs are independent graph structures in the heterogeneous network that are not connected to each other; A skeleton structure of the heterogeneous network is extracted based on the subgraph.
5. The method according to claim 4, wherein: The extracting the skeleton structure of the heterogeneous network based on the subgraph includes: Determine the number of nodes in each subgraph; When the number of nodes in a subgraph is greater than a preset threshold, the subgraph is used as a target subgraph; A skeleton structure of the heterogeneous network is extracted from the target subgraph.
6. The method according to claim 5, wherein: The step of extracting the skeleton structure of the heterogeneous network from the target subgraph includes: Calculate the weight of each node and / or the weight of each link in the target subgraph; wherein the weight is used to represent the importance of the corresponding node or link in the target subgraph; Based on the weights of the nodes and / or the weights of the links, selecting a target link from the links in the target subgraph as a key route; The skeleton structure of the heterogeneous network is extracted according to the key routes.
7. The method according to claim 6, wherein: The selecting a target link as a key route from the links in the target subgraph based on the weights of the nodes and / or the weights of the links comprises: When the weight of the node or link is greater than a preset weight threshold, the corresponding node or link is regarded as a key node or key link; Using the links between the key nodes and the key links as target routes; Based on the type of the reference relationship corresponding to the target route, a key route is screened out from each target route.
8. An information processing device, wherein: include: An acquisition module, used to acquire object information of multiple objects associated with a target field; wherein the types of the multiple objects include papers, patents and products; A determination module, configured to determine reference relationships between the plurality of objects according to the object information; wherein the reference relationships include reference relationships between objects of different types and reference relationships between objects of the same type; An extraction module, used to construct a heterogeneous network based on the reference relationship and extract a skeleton structure of the heterogeneous network; wherein the skeleton structure is used to characterize the association relationship and interaction between the multiple objects; The display module is used to visualize the skeleton structure so as to determine the analysis result of the target field based on the visualized skeleton structure.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Information processing method and device, electronic equipment and storage medium
CN114691814A
Information processing method and device, electronic equipment and computer readable storage medium
CN117669572A
Mining strong relevance between heterogeneous entities from their co-ocurrences
US20150332158A1