Method and system for analyzing emerging technology based on hypergraph network
By constructing and updating a hypergraph network and combining it with an emerging technology indicator system, the problem of capturing multi-dimensional correlations and temporal relationships in emerging technology identification is solved, achieving higher-precision emerging technology identification.
Patent Information
- Application Number
- CN202510717421.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
When identifying emerging technologies, existing technologies have difficulty accurately capturing multi-dimensional correlations and temporal relationships, resulting in low recognition accuracy and easily missing low-frequency but emerging technical terms.
A hypergraph network-based method is used to construct the first hypergraph network of the patent dataset. The second hypergraph network is generated by obtaining the embedding vectors of key attributes and inputting them into the hypergraph network model for update. The emerging technology indicator system is used to evaluate the technology evolution path.
It improves the accuracy of identifying emerging technologies, can more comprehensively capture the temporal characteristics and correlation characteristics of technological evolution, and breaks through the limitations of traditional methods' single structure and weak relationship modeling capabilities.
Smart Images

Figure CN120671723A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis, and in particular to an emerging technology analysis method and system based on a hypergraph network. Background Art
[0002] Emerging technologies refer to patented technologies that are significantly innovative and develop relatively rapidly. Their development and application have a great impact on society. However, since it is difficult to accurately identify emerging technologies in a large number of patent texts, when the predicted emerging technologies are put into production and applied, they cannot meet design requirements to the maximum extent.
[0003] In practical applications, qualitative and quantitative methods are usually used to help experts identify and predict the development trends and potential impacts of emerging technologies. This method has achieved automatic identification of emerging technologies to a certain extent. However, since most studies only focus on the judgment of novelty, they often ignore the correlation between indicators of various dimensions of emerging technologies, resulting in insufficient precision in the identification results of emerging technologies, and thus low accuracy in determining emerging technologies. Summary of the Invention
[0004] This application provides an emerging technology analysis method and system based on a hypergraph network, which can improve the accuracy of emerging technology identification.
[0005] In a first aspect, the present application provides an emerging technology analysis method based on a hypergraph network, the method comprising: obtaining multiple patent data sets to be tested, and for any of the patent data sets, determining multiple key attributes of the patent data sets, wherein the key attributes include at least technical themes; constructing a first hypergraph network of the patent data set, wherein the first hypergraph network comprises multiple hyperedges, and the hyperedges comprise multiple key attributes, and for any of the key attributes in the first hypergraph network, obtaining an embedding vector of the key attribute; obtaining multiple embedding vectors of the first hypergraph network, and inputting the embedding vectors into a hypergraph network model, wherein the hypergraph network model updates the first hypergraph network to obtain a second hypergraph network of the patent data set; obtaining the first hypergraph network and the second hypergraph network of each of the patent data sets, and determining an emerging technology path based on the first hypergraph network and the second hypergraph network, and generating emerging technology evaluation information representing industry development trends based on the emerging technology path.
[0006] In one possible implementation, the patent dataset includes multiple patents to be tested; for any of the patent datasets, determining multiple key attributes of the patent dataset includes: for any of the patents to be tested, obtaining patent record information of the patent to be tested, wherein the patent record information includes title, primary classification number, secondary classification number, publication date, applicant, agent and agency structure; obtaining one or more technical subjects of the patent to be tested, determining each of the patent record information and the technical subject as the key attributes of the patent to be tested, and obtaining the key attributes of each of the patents to be tested in the patent dataset as the key attributes of the patent dataset.
[0007] In one possible embodiment, obtaining one or more technical themes of the patent to be tested includes: obtaining patent text information of multiple patents to be tested, performing preprocessing operations on the patent text information, the preprocessing operations including word segmentation and deduplication, wherein the patent text information includes title, abstract and claims; performing term extraction on the preprocessed patent text information to obtain multiple key technical terms, and determining multiple technical themes based on the key technical terms, wherein any of the technical themes contains multiple key technical terms; for any of the patents to be tested, obtaining multiple key technical terms of the patent to be tested, and determining one or more technical themes of the patent to be tested based on the key technical terms.
[0008] In one possible implementation, constructing the first hypergraph network of the patent data set includes: for any patent to be tested in the patent data set, obtaining multiple key attributes of the patent to be tested, and using the multiple key attributes of the patent to be tested as nodes; connecting each of the nodes according to a preset hyperedge type to form multiple hyperedges of the patent to be tested, and the multiple hyperedges of each of the patents to be tested constitute the first hypergraph network.
[0009] In one possible embodiment, for any of the key attributes in the first hypergraph network, obtaining the embedding vector of the key attribute includes: for any of the key attributes, obtaining multiple adjacent key attributes of the key attribute in the first hypergraph network, the adjacent key attributes being key attributes that are on the same hyperedge as the key attribute; determining the proximity relationship between the key attribute and each of the adjacent key attributes, and generating an adjacency list of the key attribute based on the multiple proximity relationships of the key attributes; performing a random walk in the adjacency list according to a preset jump probability and a preset residence probability, determining multiple node sequences of the key attribute according to the random walk results, and generating the embedding vector of the key attribute based on the multiple node sequences.
[0010] In one possible implementation, the embedding vector is input into a hypergraph network model, and the hypergraph network model updates the first hypergraph network to obtain a second hypergraph network of the patent data set, including: the hypergraph network model generates a feature matrix of the first hypergraph network based on the embedding vector, and generates multiple initial predicted hyperedges based on the feature matrix, and any of the initial predicted hyperedges includes multiple key attributes; obtains a prediction score for each of the initial predicted hyperedges, and when the prediction score meets a preset threshold, uses the initial predicted hyperedge as a target predicted hyperedge; obtains multiple original hyperedges of the first hypergraph network, and the original hyperedges and the target predicted hyperedges constitute the second hypergraph network.
[0011] In one possible embodiment, the hypergraph network model is a pre-trained model, and the training steps of the hypergraph network model include: obtaining a training set hypergraph network, the training set hypergraph network including feature vectors of multiple key attributes, performing attention convolution on the multiple feature vectors according to preset parameters to construct a feature matrix; determining multiple predicted hyperedges according to the feature matrix, determining the true label of each predicted hyperedge, and determining the predicted label of each predicted hyperedge according to the loss function, calculating the difference value between the predicted label and the true label, and updating the preset parameters of the hypergraph network model according to the difference value.
[0012] In one possible implementation, there is a temporal relationship between each of the patent data sets; obtaining the first hypergraph network and the second hypergraph network of each of the patent data sets, and determining the emerging technology based on the first hypergraph network and the second hypergraph network includes: for any of the patent data sets, determining one or more technical themes of the first hypergraph network of the patent data set, and obtaining the embedding vector of each of the technical themes; using the embedding vector of the technical theme as the technical theme feature vector of the current patent data set, and constructing a technology evolution path diagram based on the temporal relationship between the technical theme feature vector and each of the patent data sets, wherein the technology evolution path diagram includes multiple technology evolution paths; for any technology evolution path, determining the emerging technology indicators of the technology evolution path based on the first hypergraph network and the second hypergraph network, and determining whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path.
[0013] In one possible embodiment, the technology evolution path includes multiple technology theme feature vectors; constructing a technology evolution path diagram based on the temporal relationship between the technology theme feature vectors and each of the patent data sets includes: for any two technology theme feature vectors in different patent data sets, calculating the similarity between the two technology theme feature vectors; when the similarity is within a first threshold range, determining the two technology theme feature vectors as an inheritance relationship; when the similarity is within a second threshold range, determining the two technology theme feature vectors as a differentiation relationship; constructing multiple technology evolution paths based on the inheritance relationship and the differentiation relationship, and constructing the multiple technology evolution paths into a technology evolution path diagram based on the temporal relationship between each of the patent data sets.
[0014] In a possible implementation, the emerging technology indicators include technology innovation indicators, technology influence indicators and interdisciplinary indicators; determining the emerging technology indicators of the technology evolution path based on the first hypergraph network and the second hypergraph network, and determining whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path includes: for any of the technology theme feature vectors, obtaining the technology innovation parameters, technology influence parameters and interdisciplinary parameters of the technology theme feature vector; for any of the technology evolution paths, determining the technology innovation indicators, technology influence indicators and interdisciplinary indicators of the technology evolution path based on the technology innovation parameters, technology influence parameters and interdisciplinary parameters of each of the technology themes in the technology evolution path; weighting the technology innovation indicators, technology influence indicators and interdisciplinary indicators according to a preset first weight, a preset second weight and a preset third weight, respectively, and determining whether the technology evolution path is an emerging technology path based on the weighted results.
[0015] In a possible implementation, obtaining the technical innovation parameters, technical influence parameters and interdisciplinary parameters of the technical theme feature vector includes: determining the technical theme corresponding to the technical theme feature vector, obtaining the first classification number and the second classification number of the technical theme, and determining the technical innovation index of the current technical theme feature vector based on the first classification number and the second classification number; obtaining the first degree of the technical theme in the corresponding first hypergraph network and the second degree in the corresponding second hypergraph network, and determining the technical influence index of the technical theme feature vector based on the first degree and the second degree; obtaining the total number of the first classification numbers of the technical theme and the total number of the second classification numbers of the technical field to which the technical theme belongs, and determining the interdisciplinary index of the technical theme feature vector based on the total number of the first classification numbers and the total number of the second classification numbers.
[0016] According to a second aspect of the present application, there is provided an emerging technology identification device, comprising: a patent data acquisition unit for acquiring a plurality of patent data sets to be tested, and for any of the patent data sets, determining a plurality of key attributes of the patent data set, wherein the key attributes include at least a technical subject; a hypergraph network construction unit for constructing a first hypergraph network of the patent data set, wherein the first hypergraph network includes a plurality of hyperedges, and the hyperedges include a plurality of the key attributes, and for any of the key attributes in the first hypergraph network, obtaining an embedding vector of the key attribute; a hypergraph network update unit for acquiring a plurality of the embedding vectors of the first hypergraph network, inputting the embedding vectors into a hypergraph network model, and updating the first hypergraph network using the hypergraph network model to obtain a second hypergraph network of the patent data set; and an emerging technology determination unit for acquiring the first hypergraph network and the second hypergraph network of each of the patent data sets, and determining an emerging technology based on the first hypergraph network and the second hypergraph network.
[0017] The technical solution provided by this embodiment of the present application achieves accurate prediction of emerging technologies by performing hypergraph modeling and vector analysis on patent data. Specifically, a first hypergraph network of patent data at each stage is constructed, and each technical feature in the first hypergraph network is vector-embedded and associated. The second hypergraph network is randomly updated based on the embedded vectors, and the emerging technology path is analyzed based on the technical theme feature vectors in the first hypergraph network and the second hypergraph network to generate evaluation information for emerging technologies. At the same time, by constructing an emerging technology indicator system, the relevance and novelty evaluation of emerging technology identification are enriched, further improving the accuracy of emerging technology identification.
[0018] It can be seen that the technical solution provided by this embodiment of the present application does not need to rely on traditional single-dimensional indicators or simple graph network models to identify emerging technologies. Instead, it uses a pre-trained model to update the hypergraph network, so that the hypergraph network can represent high-order correlation relationships between multiple types of nodes. It breaks through the limitations of traditional technology identification methods such as single structure, weak relationship modeling capabilities, and lack of dynamic time series analysis. It performs emerging technology identification on the basis of updating the hypergraph network through hyperedge prediction based on the hypergraph network model, thereby improving the accuracy of emerging technology identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A schematic diagram of the steps of a hypergraph network-based emerging technology analysis method provided in an embodiment of the present application; Figure 2 A schematic diagram showing the connection of some hyperedges of a first hypergraph network provided in one embodiment of the present application; Figure 3 A schematic diagram of the steps of a method for training a hypergraph network model provided in one embodiment of the present application; Figure 4 A schematic diagram of the steps for determining emerging technologies provided for one embodiment of the present application; Figure 5 A schematic diagram of the structure of a technology evolution roadmap provided for one embodiment of the present application; Figure 6 A schematic diagram of the structure of an identification device of an emerging technology provided in one embodiment of the present application; Figure 7 A schematic structural diagram of a computer device provided in accordance with one embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0022] In addition, the descriptions of "first", "second", etc. in this application are for descriptive purposes only and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values can be based on additional conditions or values beyond the described values in practice.
[0023] With the continuous development of various technological fields, the identification of emerging technologies has emerged. Generally speaking, emerging technologies refer to innovative, relatively rapidly developing patented technologies. By identifying emerging technologies, we can understand the development trends and potential impacts of a particular technology field, provide new technical directions for corporate research, and thus promote social, economic, and scientific development.
[0024] In related technologies, qualitative or quantitative methods are usually used to identify emerging technologies. Qualitative methods rely entirely on independent expert evaluations, which are difficult to meet large-scale needs and may lead to identification bias due to subjectivity. Most commonly, quantitative methods are used to identify emerging technologies. These quantitative methods refer to the quantitative identification of emerging technologies through big data technology. Specifically, keyword statistics are performed on large amounts of patent data, and semantic relationships between keywords are mined by building word networks and graph embedding models, thereby enabling the identification of emerging technologies under large-scale data.
[0025] However, since emerging technology identification typically assesses novelty and creativity, and related technologies often focus only on static analysis of keyword quantity and time, and technology is rapidly evolving, traditional identification methods tend to overlook the multidimensional changes and temporal relationships of technology, resulting in a relatively simplistic assessment of innovation and low accuracy in emerging technology identification. Furthermore, as emerging technologies are in their early stages of development, the number of literature and the frequency of keyword terms are relatively low. In practical applications, related technologies tend to omit low-frequency but emerging technical terms, leading to biased identification of emerging technologies and low accuracy.
[0026] In view of this, one or more embodiments of the present application provide an emerging technology analysis method and system based on a hypergraph network, which can solve the above problems. By dynamically analyzing patent data through the multi-attribute connection and update of the hypergraph network, the temporal characteristics and correlation characteristics of technological evolution can be effectively captured, thereby improving the accuracy of emerging technology identification.
[0027] See also Figure 1 One embodiment of the present application provides a method for analyzing emerging technologies based on a hypergraph network, which may include the following steps: S1: Acquire multiple patent data sets to be tested, and for any of the patent data sets, determine multiple key attributes of the patent data set, where the key attributes at least include a technical subject.
[0028] The above-mentioned patent dataset can be understood as a collection of patent data consisting of multiple patents in a certain technical field, and the above-mentioned key attributes can be understood as the basic information of each patent, including the patent title, main classification number, sub-classification number, technical subject, publication date, applicant, agent, patent agency and other attribute fields of patent information. In the identification and prediction of emerging technologies, the technical subject, as an important attribute representing the content of the patent text, reflects the technical innovation and core content of the patent. Taking the technical subject as the key attribute can accurately capture the technical essence of the patent and provide the core basis for subsequent analysis. Among them, the technical subject represented by each patent classification number or the keyword extraction based on the patent text can be used to determine the technical subject of the patent text by analyzing the keywords or counting their frequency.
[0029] S3: Construct a first hypergraph network of the patent dataset, wherein the first hypergraph network includes multiple hyperedges, and the hyperedges include multiple key attributes. For any key attribute in the first hypergraph network, obtain an embedding vector of the key attribute.
[0030] The first hypergraph network is used to characterize the potential connections between key attributes in a patent. The first hypergraph network allows for consideration of the multi-dimensional technical relationships between key attributes, ensuring comprehensiveness and accuracy in identifying emerging technologies. Figure 2 , Figure 2 This is a schematic diagram of the hyperedge connection of part of the first hypergraph network. Specifically, for any patent, the key attributes of the patent are connected as nodes to form a hyperedge of the first hypergraph network. The above hyperedge is used to characterize a certain technical connection between multiple key attributes. Furthermore, for any node of the first hypergraph network, the embedding vector of the node is constructed according to the hyperedge where the current node is located, that is, the embedding vector of each key attribute is constructed. The key attributes are represented by the embedding vector, and while setting the hyperedge to retain the correlation between the key attributes, it is convenient for subsequent model analysis and link prediction of the key attributes. For example, Figure 5 As shown in the figure, a hyperedge is constructed using the patent title, primary classification number, secondary classification number, and technical subject as key attributes, and another hyperedge is constructed using the patent title, primary classification number, publication date, and technical subject as key attributes. Furthermore, for the key attribute of technical subject, a word embedding method can be used to construct an embedding vector of the technical subject in the current hyperedge to obtain a vector representation of the technical subject, which is then used to predict and determine emerging trends in the technical subject. It should be noted that since multiple technical subjects may appear in a patent text, in the hyperedge containing the key attribute of technical subject, multiple technical subjects are treated as connectable nodes.
[0031] S5: Obtain multiple embedding vectors of the first hypergraph network, input the embedding vectors into a hypergraph network model, and the hypergraph network model updates the first hypergraph network to obtain a second hypergraph network of the patent dataset.
[0032] The second hypergraph network is an updated version of the first hypergraph network, containing more hyperedges than the first. The first hypergraph network is constructed based on the initial key attributes of the patent data, reflecting known explicit correlations. However, within each key attribute, there are often hidden or indirect higher-order correlations, especially between technical themes, which are not directly accessible from the original data. Therefore, by analyzing and processing the embedding vectors using a hypergraph network model, these higher-order correlations and latent relationships can be captured to form the second hypergraph network, thereby more comprehensively representing the complex structure of the patent data.
[0033] The aforementioned hypergraph network model is a pre-trained model used to identify potential new edges in the initial hypergraph network, reflecting potential technology evolution paths and emerging technology trends. During this update process, the hypergraph network model may discover connections between certain technology topics that were not sufficiently emphasized in the initial network, or that certain combinations of key attributes indicate new development directions. By integrating and analyzing the embedding vectors of key attributes across different dimensions, new hyperedges are further constructed to achieve link prediction and improve the accuracy of emerging technology identification.
[0034] S7: Obtain the first hypergraph network and the second hypergraph network of each of the patent data sets, determine an emerging technology path based on the first hypergraph network and the second hypergraph network, and generate emerging technology evaluation information representing industry development trends based on the emerging technology path.
[0035] The above-mentioned emerging technology paths are the development paths of emerging technologies represented by multiple technology themes with a temporal evolution relationship. Based on the above-mentioned emerging technology paths, emerging technology evaluation information representing industry development trends can be obtained. The above-mentioned emerging technology evaluation information can provide decision-making support for enterprises, investors, etc., helping them understand the direction and focus of industry technology development and accurately grasp the opportunities of emerging technologies.
[0036] In this embodiment, multiple technology theme evolution paths are determined through the first hypergraph network and the second hypergraph network, and whether the technology theme evolution path is an emerging technology path is measured by the emerging technology indicator system. Specifically, based on the first hypergraph network and the updated second hypergraph network, the evolutionary relationship of different technology themes can be determined by comparing and analyzing the similarity of technology themes from different patent data sets, thereby determining multiple technology theme evolution paths. On the basis of determining multiple technology theme evolution paths, emerging technology paths are determined by identifying which technology paths meet the characteristics of emerging technologies, such as the determination of characteristics such as innovation, influence, and interdisciplinary integration.
[0037] In this embodiment, by performing hypergraph modeling and vector analysis on patent data, accurate prediction of emerging technologies is achieved. Specifically, a first hypergraph network of patent data at each stage is constructed, and each technical feature in the first hypergraph network is vector-embedded and associated, the second hypergraph network is randomly updated according to the embedded vector, and the emerging technology path is analyzed based on the technical theme feature vectors in the first hypergraph network and the second hypergraph network to generate evaluation information of emerging technologies, which may include the development trend of emerging technologies, potential application fields, market prospects, etc. These evaluation information provide strong support for the R&D decisions, investment choices, strategic layouts of enterprises, and the formulation of technology policies by policymakers. At the same time, by performing feature evaluation on multiple technology evolution paths, the correlation evaluation and novelty evaluation of emerging technology identification are enriched, and the accuracy of emerging technology identification is further improved.
[0038] Based on the above ideas, the technical solution provided by this embodiment of the present application does not need to rely on traditional single-dimensional indicators or simple graph network models to identify emerging technologies. Instead, it uses a pre-trained model to update the hypergraph network, so that the hypergraph network can represent high-order correlation relationships between multiple types of nodes, breaking through the limitations of traditional technology identification methods such as single structure, weak relationship modeling capabilities and lack of dynamic timing analysis. On the basis of updating the hypergraph network through hyperedge prediction based on the hypergraph network model, emerging technology identification is performed, thereby improving the accuracy of emerging technology identification.
[0039] In one possible implementation, based on step S1 above, any patent dataset contains multiple patents to be tested, each of which has its own key attributes. The key attributes may be the patent title, primary classification number, secondary classification number, technical subject, publication date, applicant, agent, and patent agency recorded or represented by each patent to be tested. Specifically, for any of the patent datasets, determining the multiple key attributes of the patent dataset includes the following steps: S101: For any of the patents to be tested, obtain patent record information of the patent to be tested, wherein the patent record information includes title, primary classification number, secondary classification number, publication date, applicant, agent and agency structure.
[0040] S103: Obtain one or more technical themes of the patent to be tested, determine each of the patent record information and the technical themes as key attributes of the patent to be tested, and obtain the key attributes of each of the patents to be tested in the patent dataset as the key attributes of the patent dataset.
[0041] In this embodiment, the aforementioned patent record information can be obtained directly from the text of the patent document, and the aforementioned technical themes can be obtained indirectly through keyword extraction and analysis based on the patent text. By using the patent title, primary classification number, secondary classification number, technical theme, publication date, applicant, agent, and patent agency of any patent under test as key attributes for subsequent hyperedge connection and hypergraph construction, a multi-dimensional fusion expression of the patent under test can be achieved, thereby enriching the multi-dimensional associations between the various nodes of the patent under test and improving the accuracy of emerging technology identification.
[0042] In one embodiment, one or more technical themes of the patent to be tested can be obtained in the manner of term extraction and subject identification. Specifically, for any patent to be tested, the patent text information of the patent to be tested is first obtained. The above-mentioned patent text information includes the title, abstract and claims of the patent to be tested. These parts usually contain the core content and key information of the patent. Furthermore, the patent text information is subjected to pre-processing operations such as analysis and deduplication to achieve text splitting and removal of duplicate information. Furthermore, specific term extraction methods, such as natural language processing, word frequency statistics, text clustering and other term extraction technologies, can be used to extract terms from the pre-processed patent text information to obtain multiple key technical terms. The above-mentioned key technical terms are the core expressions of the patent text information. One or more technical themes can be determined by analyzing the key technical terms. Each technical theme contains multiple key technical terms for determining the current technical theme. In this embodiment, by extracting key technical terms from patent text information and determining technical themes, the core technical content of the patent to be tested can be grasped more accurately. Specifically, by performing word segmentation processing and term extraction on the collected patent text information of the patent to be tested, key technical terms reflecting the core content of the patent technology are identified and extracted from the pre-processed text, and one or more technical themes of the patent are determined by analyzing the correlation and co-occurrence relationship between the key technical terms. The obtained technical themes are used to represent the innovation points or research directions of the patent to be tested in a specific technical field. This precise method of determining technical themes can effectively avoid omissions or misjudgments caused by simple classification numbers or keyword statistics in traditional methods, thereby improving the accuracy of patent analysis.
[0043] In one possible implementation, constructing the first hypergraph network of the patent dataset in step S3 includes the following steps: S301: For any patent to be tested in the patent dataset, multiple key attributes of the patent to be tested are obtained, and the multiple key attributes of the patent to be tested are used as nodes.
[0044] S303: Connecting each of the nodes according to a preset hyperedge type to form multiple hyperedges of the patent to be tested, and the multiple hyperedges of each of the patents to be tested constitute the first hypergraph network.
[0045] In this implementation, key attributes are connected to form different hyperedges based on pre-set hyperedge types, each reflecting a different technical relationship. Specifically, these pre-set hyperedge types include technical theme association hyperedges, temporal evolution hyperedges, innovation subject technology hyperedges, technical classification evolution hyperedges, innovation network hyperedges, and technical feature hyperedges. Multiple key attributes of the patent under test are connected as nodes to form these multiple hyperedges.
[0046] Among them, the above-mentioned technical theme association hyperedge is used to reflect the multiple associations of the patent to be tested on the technical theme, the above-mentioned temporal evolution hyperedge is used to reflect the association between the time evolution trajectory of the patent to be tested and the technical theme, the above-mentioned innovation subject technology hyperedge is used to reflect the association between the innovation subject and the technical theme, the above-mentioned technology classification evolution hyperedge is used to reflect the association between the evolution trend of the technology field and the technical theme, the above-mentioned innovation network hyperedge is used to reflect the association between the innovation cooperation subject and the technical theme, and the above-mentioned technology feature hyperedge is used to reflect the association between the technical characteristics of the patent to be tested and the technical theme.
[0047] Among them, each hyperedge has a specific node connection relationship, and multiple key attributes are connected to each other according to the node connection relationship of the preset hyperedge type. Specifically, the hyperedge composed of title, main classification number, sub-classification number and technical theme is used as the technical theme association hyperedge, the hyperedge composed of title, publication date, technical theme and main classification number is used as the time evolution hyperedge, the hyperedge composed of applicant, main classification number, technical theme, title and agent is used as the innovation subject technology hyperedge, the hyperedge composed of main classification number, sub-classification number, publication date and technical theme is used as the technology classification evolution hyperedge, the hyperedge composed of applicant, agency, agent and main classification number is used as the innovation network hyperedge, and the hyperedge composed of title, technical theme, main classification number, applicant and sub-classification number is used as the technology feature hyperedge.
[0048] In this implementation, by connecting nodes according to pre-set hyperedge types, a comprehensive modeling of the various complex relationships within a patent dataset is possible. This first hypergraph network not only represents the relationships between different attributes within the patent under test, but also the associations between different patents within the same patent dataset. This more accurately describes the complex interactions between multiple key attributes within the patent data, providing a structured data foundation for subsequent patent analysis and hyperedge prediction.
[0049] In a possible implementation, based on the above step S3, for any key attribute in the first hypergraph network, obtaining the embedding vector of the key attribute includes the following steps: S311: For any of the key attributes, obtain multiple adjacent key attributes of the key attribute in the first hypergraph network, where the adjacent key attributes are key attributes that are on the same hyperedge as the key attribute.
[0050] S313: Determine the proximity relationship between the key attribute and each of the adjacent key attributes, and generate an adjacency list of the key attribute based on the multiple proximity relationships of the key attributes.
[0051] S315: Perform a random walk in the adjacency table according to a preset jump probability and a preset residence probability, determine a plurality of node sequences of the key attribute according to the random walk result, and generate an embedding vector of the key attribute according to the plurality of node sequences.
[0052] The adjacent key attributes of the aforementioned key attributes refer to other key attributes on the same hyperedge, which are used to reflect the layout structure and association relationship of the current key attribute in the first hypergraph network. The adjacency table of the key attribute is determined based on the multiple adjacent relationships of the key attributes, which is then used for random walks based on the adjacency table. It should be noted that since the key attributes of different patents to be tested may be the same, multiple adjacent key attributes of the same key attribute may come from the hyperedge of the same patent to be tested or from the hyperedges of different patents to be tested.
[0053] In this embodiment, a random walk is performed in the adjacency table according to a preset jump probability and a preset residence probability to obtain multiple node sequences obtained by the random walk. Specifically, when performing a random walk, a certain key attribute is used as a starting point to randomly jump to and reside in the next node, with a first probability of jumping and a second probability of staying. Preferably, the first probability and the second probability are both 50%. The above-mentioned jump can be understood as jumping from the current key attribute node to the adjacent key attribute in the adjacency table. The above-mentioned preset jump probability determines the possibility of jumping to a certain adjacent node. The above-mentioned residence can be understood as staying in another key attribute within the hyperedge to which the current node belongs. The above-mentioned preset residence probability determines the possibility of staying in a certain key attribute within the hyperedge to which the current node belongs.
[0054] In this embodiment, multiple node sequences are obtained based on the random walk results. These node sequences are sequences of multiple key attributes starting with a certain key attribute. Based on these node sequences, embedding vectors for the key attributes are generated and input into the hypergraph network model for hypergraph update. It should be noted that these embedding vectors represent the key attributes in a low-dimensional vector space, thereby capturing both semantic and structural information about the key attributes, helping to improve the hypergraph network model's ability to analyze key attributes and, consequently, the accuracy of identifying emerging technologies.
[0055] Specifically, key attributes are extracted from the patent dataset to construct a hypergraph H = (V, E), where H represents the hypergraph, V represents the node set consisting of all key attributes in the hypergraph H, and E represents the hyperedge set consisting of all hyperedges in the hypergraph H. For any patent to be tested in the patent dataset, there exists a subset of hyperedges e∈E and a subset of nodes v∈V, such that e={v1,v2,v3} indicates that key attributes v1, v2, and v3 appear together in the same patent to be tested. Furthermore, if two or more key attributes appear in the same hyperedge, for example, v i and v j , then the above key attributes are determined as adjacent relationships in the adjacency table, expressed as N(v i )=N(v i )∪N(v j ) and N(v j )=N(v j )∪N(v i ), where the adjacency list is u is a node in the hyperedge, v is another node in the hyperedge, Indicates that both nodes are contained within the hyperedge.
[0056] Furthermore, a preset number of jump and stay operations are performed in the adjacency table N(v), and the preset jump probability and the preset stay probability can be calculated according to the following method: Among them, P′(v′ t |v t ) represents the preset jump probability, P″(v″ t |v t ) represents the preset residence probability, v t represents the key attribute of the starting point of the node, v′ t Represents a key attribute of an adjacency in the adjacency table, v″ t Indicates that v t Another key attribute of the same hyperedge, v″ t ∈e, |e| is the number of nodes in hyperedge e. For example, 10 walks are performed on each key attribute in the first hypergraph network, with each walk being 20 steps long, thereby generating multiple node sequences for the current key attribute. Based on these multiple node sequences, an embedding vector h(v) for the key attribute is generated using a word embedding method.
[0057] In one possible implementation, based on step S5 above, the embedding vector is input into a hypergraph network model, and the hypergraph network model updates the first hypergraph network to obtain a second hypergraph network of the patent dataset, including the following steps: S501: The hypergraph network model generates a feature matrix of the first hypergraph network according to the embedding vector, and generates a plurality of initial predicted hyperedges according to the feature matrix, wherein any of the initial predicted hyperedges includes a plurality of key attributes.
[0058] S503: Obtain a prediction score of each of the initial predicted hyperedges, and when the prediction score meets a preset threshold, use the initial predicted hyperedge as a target predicted hyperedge.
[0059] S505: Acquire multiple original hyperedges of the first hypergraph network, where the original hyperedges and the target predicted hyperedges constitute the second hypergraph network.
[0060] In this embodiment, a feature matrix is generated from each embedding vector of the first hypergraph network, and hyperedge prediction is performed based on the feature matrix to form a second hypergraph network. The above-mentioned feature matrix is used to represent the feature information of each key attribute in the first hypergraph network, including semantic information representing the technical subject and structural information of the hypergraph structure representing other key attributes. According to the feature matrix, through two layers of attention convolution and fully connected layer mapping, the hypergraph network model can generate multiple initial predicted hyperedges and prediction scores of each predicted hyperedge. Each initial predicted hyperedge contains multiple key attributes. The generated initial predicted hyperedges are used to reflect the potential associations that may exist between the key attributes, and to screen out target predicted hyperedges whose prediction scores meet the preset threshold to ensure the reliability and relevance of the predicted hyperedges. The original hyperedges of the first hypergraph network are combined with the target predicted hyperedges to form a second hypergraph network that represents the original technical associations and predicts potential associations.
[0061] The prediction score of the above initial predicted hyperedge can be expressed as Among them, |nodes| represents the number of key attributes, h i Represents the embedding vector of the i-th key attribute, σ represents the probability of the hyperedge output by the activation function, and σ is used to map the output result to the (0, 1) interval. Optionally, when the preset threshold is 0.5, the initial predicted hyperedge greater than 0.5 can be used as the target predicted hyperedge. Optionally, the prediction result can also be converted into a predicted label Indicates that the prediction is a positive sample, and the current initial prediction hyperedge constitutes the target prediction hyperedge. Indicates that the prediction is a negative sample, and the current initial predicted hyperedge is not the target predicted hyperedge.
[0062] In this embodiment, the second hypergraph network adds verified target prediction hyperedges on the basis of the original hypergraph network, so that the second hypergraph network can more comprehensively reflect the complex connections between technical topics in the patent dataset. The hypergraph network model can discover potential technical associations, thereby more accurately identifying emerging technology paths and trends, and improving the accuracy of emerging technology identification.
[0063] See also Figure 3 In this embodiment, the hypergraph network model is a pre-trained model, and the training steps of the hypergraph network model include: S511: Obtain a training set hypergraph network, wherein the training set hypergraph network includes feature vectors of multiple key attributes, and perform attention convolution on the multiple feature vectors according to preset parameters to construct a feature matrix.
[0064] S513: Determine multiple predicted hyperedges based on the feature matrix, determine the true label of each predicted hyperedge, and determine the predicted label of each predicted hyperedge based on the loss function, calculate the difference value between the predicted label and the true label, and update the preset parameters of the hypergraph network model according to the difference value.
[0065] The above preset parameters may include input dimension, hidden dimension, output dimension, number of attention heads, learning rate, number of training rounds or other model parameters. Among them, the above training process is an iterative multiple training. The above input dimension represents the dimensional value of the embedding vector of each key attribute, preferably 32, the above hidden dimension represents the output dimension of the first layer of attention convolution layer, preferably 128, the above output dimension is the output dimension of the second layer of attention convolution layer, preferably 32, the above input dimension and output dimension must be consistent, the above number of attention heads represents the number of heads of the multi-head attention mechanism in the first layer of attention convolution layer, and the above learning rate is the learning rate of the optimizer used to update the parameters. It should be noted that the above optimizer is specifically used to dynamically adjust the learning rate of each parameter during the training process. After each training round, the gradient of the loss to the model parameters is calculated by backpropagation, and the parameters are updated by the optimizer, so as to improve the training efficiency and stability. The update rule of the optimizer is Among them, θ t is the current parameter, θ t+1 is the updated parameter, η is the learning rate, m t and v t are the first and second moments of the gradient, respectively, and is a small constant that prevents the denominator from being zero.
[0066] In this embodiment, the hypergraph network model is able to more accurately predict potential hyperedges by continuously updating and learning the key attributes and model parameters. Specifically, a training set hypergraph network containing key attribute feature vectors is prepared as input data for model training, and the feature vectors in the training set hypergraph network are processed through attention convolution to construct a feature matrix that can reflect the importance and relevance of the key attributes. Furthermore, the hypergraph network model generates multiple predicted hyperedges based on the feature matrix, calculates the prediction score of each predicted hyperedge as score = σ(mean(h)), and assigns a prediction label to the predicted hyperedge according to the prediction score. At the same time, the true label corresponding to each predicted hyperedge is obtained, and the difference value between the predicted label and the true label is calculated. The above difference value is calculated in the following way, The difference value is represented by . The preset parameters of the model are updated based on the obtained difference value, and the above process is repeated until the model performance reaches the expected standard.
[0067] See also Figure 4In one possible implementation, the patent datasets have a temporal relationship with each other; obtaining the first hypergraph network and the second hypergraph network of each patent dataset, and determining the emerging technology based on the first hypergraph network and the second hypergraph network includes the following steps: S71: For any of the patent datasets, determine one or more technical themes of the first hypergraph network of the patent dataset, and obtain an embedding vector of each of the technical themes.
[0068] S73: Using the embedding vector of the technical theme as the technical theme feature vector of the current patent data set, and constructing a technical evolution path diagram according to the temporal relationship between the technical theme feature vector and each of the patent data sets, wherein the technical evolution path diagram includes multiple technical evolution paths.
[0069] S75: For any technology evolution path, determine the emerging technology indicators of the technology evolution path according to the first hypergraph network and the second hypergraph network, and determine whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path.
[0070] In this embodiment, since each patent dataset has a time-series relationship, the technical theme feature vectors of different patent datasets also have a time-series relationship. Figure 5 , Figure 5 This is a schematic diagram of a technology evolution path diagram. The different technical theme feature vectors separated by the dotted boxes come from different patent data sets and are arranged in chronological order to form different time windows of multiple dotted boxes. The gray circles in the figure represent technical theme feature vectors, and the various technical theme feature vectors are connected to form multiple technology evolution paths. Based on the similarity between the technical theme feature vectors of different patent data sets, multiple technology evolution paths for the evolution of technical themes can be determined, and then the emerging technology paths can be determined based on the emerging technology indicators of each technology evolution path. By mining the temporal correlation of patent data sets, this technology can dynamically track the technology evolution process and monitor the development context of the technology field in real time.
[0071] In this embodiment, the technology evolution path diagram includes multiple technology evolution paths, each of which includes multiple technology theme feature vectors. The technology evolution path diagram is used to display the development process of technology over different time windows. The technology evolution path diagram can represent the inheritance and differentiation relationship between different technology themes. The technology evolution path diagram is constructed based on the temporal relationship between the technology theme feature vectors and each of the patent data sets in the following steps: S731: For any two technical theme feature vectors in different patent data sets, calculate the similarity between the two technical theme feature vectors.
[0072] S733: When the similarity is within a first threshold range, the two technical theme feature vectors are determined to be in an inheritance relationship; when the similarity is within a second threshold range, the two technical theme feature vectors are determined to be in a differentiation relationship.
[0073] S735: Constructing a plurality of technology evolution paths according to the inheritance relationship and the differentiation relationship, and constructing the plurality of technology evolution paths into a technology evolution path diagram according to the temporal relationship between each of the patent data sets.
[0074] In this embodiment, different patent data sets are considered to be in different time windows due to the existence of a temporal relationship. When determining the evolution relationship between any two technical theme feature vectors of different patent data sets based on their similarity, the inheritance relationship or differentiation relationship is specifically determined based on the preset first threshold range and the preset second threshold range. Figure 5 For example, a solid line represents an inheritance relationship, and a dotted line represents a differentiation relationship. Preferably, for any two technical theme feature vectors in different patent data sets, when the similarity value is less than 0.2, it is considered that there is no evolutionary relationship between the two; when the similarity value is between 0.2 and 0.5, it is considered that there is a differentiation relationship between the two; when the similarity value is greater than 0.5 and less than 1.0, it is considered that there is an inheritance relationship between the two.
[0075] In this embodiment, based on the relevant information of the hyperedges where the technical features are located in the first hypergraph network and the second hypergraph network, the complex associations between technical topics are captured, thereby determining the emerging technology indicators of each technology evolution path. The above-mentioned emerging technology indicators include technology innovation indicators, technology influence indicators, and interdisciplinary indicators. Specifically, determining the emerging technology indicators of the technology evolution path based on the first hypergraph network and the second hypergraph network, and determining whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path includes the following steps: S751: For any of the technical theme feature vectors, obtain the technical innovation parameter, technical influence parameter and interdisciplinary degree parameter of the technical theme feature vector.
[0076] S753: For any of the technology evolution paths, determine the technology innovation index, technology influence index and interdisciplinary index of the technology evolution path based on the technology innovation parameters, technology influence parameters and interdisciplinary parameters of each of the technology themes in the technology evolution path.
[0077] S755: Weight the technological innovation index, the technological influence index and the interdisciplinary index according to a preset first weight, a preset second weight and a preset third weight respectively, and determine whether the technological evolution path is an emerging technology path based on the weighted results.
[0078] In this embodiment, each emerging technology indicator is weighted and summed by setting preset weights, and the emerging technology path is identified based on the total emerging technology score obtained by weighted calculation. The technology evolution path with a higher emerging technology total score is used as the emerging technology path, for example, the technology evolution path with the top five emerging technology total scores is used as the emerging technology path. Specifically, the above-mentioned TotalScore = A1×InnovationIndex+A2×ImpactIndex+A3×InterdisciplinaryIndex, wherein Innovation Index represents the technology innovation index, Impact Index represents the technology influence index, Interdisciplinary Index represents the interdisciplinary index, and A1, A2 and A3 represent the preset first weight, the preset second weight and the preset third weight, respectively. Preferably, the above-mentioned preset first weight, preset second weight and preset third weight are set to 0.4, 0.4 and 0.2, respectively.
[0079] In this embodiment, the technical innovation parameter of the technical theme feature vector is obtained based on the number of classification numbers of the technical themes in the first hypergraph network and the second hypergraph network. Specifically, the technical theme corresponding to the technical theme feature vector is determined, the number of first classification numbers and the number of second classification numbers of the technical theme are obtained, and the technical innovation index of the current technical theme feature vector is determined based on the number of first classification numbers and the number of second classification numbers. The above technical innovation index is expressed as Among them, Novelty content is the technological innovation index, N new_ipc is the number of first classification numbers, N total_ipc It should be noted that the first classification number is the number of classification numbers of the technical subject in the previous time window of the technical subject feature vector, and the second classification number is the number of classification numbers of the technical subject in the current time window of the technical subject feature vector.
[0080] In this embodiment, the technical influence parameter of the technical theme feature vector is obtained based on the degrees of the first hypergraph network and the second hypergraph network. Specifically, the first degree of the technical theme in the corresponding first hypergraph network and the second degree in the corresponding second hypergraph network are obtained, and the technical influence index of the technical theme feature vector is determined based on the first degree and the second degree. The above technical influence index is expressed as Degree(i)=Degreeoriginal (i)+Degree new (i), where Degree original (i) is the first degree, Degree new (i) is the second degree. It should be noted that the first degree is the degree of the technical theme in the first hypergraph network, and the second degree is the degree of the technical theme as a key attribute in the second hypergraph network. The above degrees can be understood as the number of hyperedges connected by the technical theme.
[0081] In this embodiment, the interdisciplinary index of the technical theme feature vector is obtained according to the number of classification numbers in the technical field to which the first hypergraph network and the second hypergraph network belong. Specifically, the total number of the first classification numbers of the technical theme and the total number of the second classification numbers in the technical field to which the technical theme belongs are obtained, and the interdisciplinary index of the technical theme feature vector is determined according to the total number of the first classification numbers and the total number of the second classification numbers. The above-mentioned interdisciplinary index Among them, Interdisciplinarity category Indicates the interdisciplinary index, N corss_discipline Indicates the total number of first classification numbers, N total_class Indicates the total number of second classification numbers. It should be noted that the total number of first classification numbers mentioned above is the number of IPC classification numbers in different technical fields, and the total number of second classification numbers mentioned above is the number of IPC classification numbers in the technical fields involved in the current technical subject. The total number of second classification numbers mentioned above is determined based on the IPC classification numbers of the original patents corresponding to all key technical terms under the current technical subject.
[0082] In another embodiment, when determining the interdisciplinary index, the interdisciplinary connection index also needs to be considered. Among them, CrossLnk Interdisciplinarity is the interdisciplinary connectivity index, N cross_link N represents the number of connections of target prediction hyperedges in different technical fields connected to technical topics. total_link Indicates the number of connections of the target prediction hyperedge connected to the technical theme. Further, the interdisciplinary index and the interdisciplinary connection index in the above embodiment are weighted and summed with specific weights, and the result of the weighted sum is used as the interdisciplinary index in this embodiment.
[0083] In a specific application example, for example, when conducting new data analysis in the field of new energy vehicles, Chinese patent data from 2016 to 2024 can be obtained for emerging technology identification. The Chinese patent data from 2016 to 2018, the Chinese patent data from 2018 to 2020, the Chinese patent data from 2020 to 2022, and the Chinese patent data from 2022 to 2024 are respectively used as multiple patent data sets to be tested with a time series relationship. By constructing and updating a hypergraph network for the above multiple patent data sets, based on the emerging technology indicators and the second hypergraph network, multiple technology evolution paths are determined in the technology theme embedding vector of the first hypergraph network, and the technology evolution paths with the top five comprehensive evaluation scores are taken as emerging technology paths.
[0084] For example, the technical topics of the first emerging technology path include, in chronological order, charging system and structural integration technology, integration of electric vehicle sensors and drive components, new energy applications of motors and control systems, and integration of charging guns and power management systems; The technical topics of the second emerging technology path include, in chronological order, drive motor and intelligent component integration technology, electric vehicle components and housing protection design, electric vehicle structure and new energy management technology, and new energy vehicle components and box design; The technical themes of the third emerging technology path include, in chronological order, electric vehicle brackets and modular design, and modular design and management of new energy vehicles; The technical topics of the fourth emerging technology path include, in chronological order, the integration technology of electric vehicles and high-voltage control systems, the integrated management of electric vehicle motors and controllers, the integration of new energy vehicle devices and battery systems, and the optimized design of electric vehicle devices and new energy systems; The technical themes of the fifth emerging technology path include, in chronological order, the design of battery and electric vehicle base systems, the integration of new energy vehicle batteries and structures, and the integrated design of housings and new energy power components.
[0085] Furthermore, a comparative analysis of the identified emerging technology paths with recent trends in the innovation and development of new energy vehicles and related corporate initiatives shows that they are generally consistent. On a Chinese patent dataset, the Hypergraph Attention Convolutional Model achieves an accuracy of up to 0.98 and a precision of up to 0.97. This demonstrates that the technical solution provided by this application, through the dynamic analysis of patent data through the multi-attribute connection and update of the Hypergraph network, effectively captures the temporal and correlation characteristics of technological evolution, thereby achieving a high degree of accuracy in the identification of emerging technologies.
[0086] See also Figure 6 , the present application also provides an identification device of an emerging technology, the device comprising: The patent data acquisition unit 100 is configured to acquire a plurality of patent data sets to be tested, and for any of the patent data sets, determine a plurality of key attributes of the patent data set, wherein the key attributes include at least a technical subject; A hypergraph network construction unit 200 is configured to construct a first hypergraph network of the patent dataset, wherein the first hypergraph network includes a plurality of hyperedges, each of the hyperedges includes a plurality of the key attributes, and for any of the key attributes in the first hypergraph network, obtain an embedding vector of the key attribute; A hypergraph network updating unit 300 is configured to obtain a plurality of embedding vectors of the first hypergraph network, input the embedding vectors into a hypergraph network model, and update the first hypergraph network using the hypergraph network model to obtain a second hypergraph network of the patent dataset; The emerging technology determination unit 400 is used to obtain the first hypergraph network and the second hypergraph network of each of the patent data sets, and determine the emerging technology based on the first hypergraph network and the second hypergraph network.
[0087] In one embodiment, the patent data acquisition unit 100 is specifically used to obtain patent record information of any patent to be tested, wherein the patent record information includes title, main classification number, sub-classification number, publication date, applicant, agent and agency structure, and obtain one or more technical themes of the patent to be tested, determine each patent record information and technical theme as the key attributes of the patent to be tested, and obtain the key attributes of each patent to be tested in the patent dataset as the key attributes of the patent dataset.
[0088] In one embodiment, the hypergraph network construction unit 200 is specifically used to obtain multiple key attributes of the patent to be tested for any patent to be tested in the patent data set, use the multiple key attributes of the patent to be tested as nodes, and connect each node according to a preset hyperedge type to form multiple hyperedges of the patent to be tested. The multiple hyperedges of each patent to be tested constitute a first hypergraph network.
[0089] In one embodiment, the hypergraph network update unit 300 is specifically used to generate a feature matrix of the first hypergraph network based on the embedded vector using the hypergraph network model, generate multiple initial predicted hyperedges based on the feature matrix, any initial predicted hyperedge includes multiple key attributes, obtain the prediction score of each initial predicted hyperedge, and when the prediction score meets the preset threshold, use the initial predicted hyperedge as the target predicted hyperedge to obtain multiple original hyperedges of the first hypergraph network, and the original hyperedges and the target predicted hyperedges constitute the second hypergraph network.
[0090] In one embodiment, the emerging technology determination unit 400 is specifically used to calculate the similarity between any two technical theme feature vectors in different patent data sets, determine the inheritance and differentiation relationship of the two technical theme feature vectors based on the first threshold range and the second threshold range, construct multiple technology evolution paths based on the inheritance relationship and differentiation relationship, determine whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path, perform objective quantitative evaluation of the emerging technology path, and generate emerging technology evaluation information that characterizes industry development trends.
[0091] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0092] An identification device of an emerging technology in an embodiment of the present application is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, or other devices that can provide the above functions.
[0093] See also Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 7 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 7 A processor 10 is taken as an example.
[0094] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0095] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0096] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0097] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0098] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0099] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0100] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0101] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0102] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, and devices according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0104] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0106] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0107] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0108] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
[0109] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A method for analyzing emerging technologies based on hypergraph networks, characterized in that: The method comprises: Acquire multiple patent data sets to be tested, and for any of the patent data sets, determine multiple key attributes of the patent data set, wherein the key attributes at least include a technical subject; Constructing a first hypergraph network of the patent dataset, wherein the first hypergraph network includes a plurality of hyperedges, each of the hyperedges includes a plurality of the key attributes, and obtaining an embedding vector of any key attribute in the first hypergraph network; Obtaining a plurality of embedding vectors of the first hypergraph network, inputting the embedding vectors into a hypergraph network model, and causing the hypergraph network model to update the first hypergraph network to obtain a second hypergraph network of the patent dataset; Obtain the first hypergraph network and the second hypergraph network of each of the patent data sets, determine an emerging technology path based on the first hypergraph network and the second hypergraph network, and generate emerging technology evaluation information representing industry development trends based on the emerging technology path.
2. The method according to claim 1, characterized in that The patent dataset includes multiple patents to be tested; for any of the patent datasets, determining multiple key attributes of the patent dataset includes: For any of the patents to be tested, obtaining patent record information of the patent to be tested, wherein the patent record information includes title, primary classification number, secondary classification number, publication date, applicant, agent, and agency structure; One or more technical themes of the patent to be tested are obtained, each of the patent record information and the technical themes are determined as key attributes of the patent to be tested, and the key attributes of each of the patent to be tested in the patent dataset are obtained as key attributes of the patent dataset.
3. The method according to claim 2, characterized in that One or more technical themes of the patent to be tested include: Obtaining patent text information of a plurality of the patents to be tested, and performing a preprocessing operation on the patent text information, wherein the preprocessing operation includes word segmentation and deduplication, wherein the patent text information includes a title, an abstract, and claims; Term extraction is performed on the preprocessed patent text information to obtain multiple key technical terms, and multiple technical themes are determined based on the key technical terms, wherein any technical theme contains multiple key technical terms; for any of the patents to be tested, multiple key technical terms of the patent to be tested are obtained, and one or more technical themes of the patent to be tested are determined based on the key technical terms.
4. The method according to claim 1 or 2, characterized in that Constructing the first hypergraph network of the patent dataset includes: For any patent to be tested in the patent dataset, multiple key attributes of the patent to be tested are obtained, and the multiple key attributes of the patent to be tested are used as nodes; Each of the nodes is connected according to a preset hyperedge type to form multiple hyperedges of the patent to be tested, and the multiple hyperedges of each of the patents to be tested constitute the first hypergraph network.
5. The method according to claim 1, wherein For any of the key attributes in the first hypergraph network, obtaining an embedding vector of the key attribute includes: For any of the key attributes, obtain a plurality of adjacent key attributes of the key attribute in the first hypergraph network, where the adjacent key attributes are key attributes that are on the same hyperedge as the key attribute; Determining a neighbor relationship between the key attribute and each of the adjacent key attributes, and generating an adjacency list of the key attribute based on the multiple neighbor relationships of the key attributes; A random walk is performed in the adjacency table according to a preset jump probability and a preset residence probability, a plurality of node sequences of the key attribute are determined according to the random walk result, and an embedding vector of the key attribute is generated according to the plurality of node sequences.
6. The method according to claim 1, characterized in that Inputting the embedding vector into a hypergraph network model, wherein the hypergraph network model updates the first hypergraph network to obtain a second hypergraph network of the patent dataset, comprising: generating, by the hypergraph network model, a feature matrix of the first hypergraph network according to the embedding vector, and generating a plurality of initial predicted hyperedges according to the feature matrix, wherein any of the initial predicted hyperedges includes a plurality of key attributes; Obtaining a prediction score for each of the initial predicted hyperedges, and taking the initial predicted hyperedge as a target predicted hyperedge if the prediction score meets a preset threshold; A plurality of original hyperedges of the first hypergraph network are obtained, wherein the original hyperedges and the target predicted hyperedges constitute the second hypergraph network.
7. The method according to claim 1 or 6, characterized in that The hypergraph network model is a pre-trained model, and the training steps of the hypergraph network model include: Obtaining a training set hypergraph network, the training set hypergraph network including feature vectors of multiple key attributes, and performing attention convolution on the multiple feature vectors according to preset parameters to construct a feature matrix; Determine multiple predicted hyperedges based on the feature matrix, determine the true label of each predicted hyperedge, and determine the predicted label of each predicted hyperedge based on the loss function, calculate the difference value between the predicted label and the true label, and update the preset parameters of the hypergraph network model based on the difference value.
8. The method according to claim 1, characterized in that There is a time sequence relationship between the patent data sets; obtaining the first hypergraph network and the second hypergraph network of each patent data set, and determining the emerging technology based on the first hypergraph network and the second hypergraph network includes: For any of the patent datasets, determining one or more technical themes of the first hypergraph network of the patent dataset, and obtaining an embedding vector of each of the technical themes; Using the embedding vector of the technical theme as the technical theme feature vector of the current patent dataset, and constructing a technology evolution path diagram based on the temporal relationship between the technical theme feature vector and each of the patent datasets, wherein the technology evolution path diagram includes multiple technology evolution paths; For any technology evolution path, the emerging technology indicators of the technology evolution path are determined according to the first hypergraph network and the second hypergraph network, and whether the technology evolution path is an emerging technology path is determined based on the emerging technology indicators of each technology evolution path.
9. The method according to claim 8, characterized in that The technology evolution path includes a plurality of technology theme feature vectors; constructing a technology evolution path diagram based on the temporal relationship between the technology theme feature vectors and each of the patent data sets includes: For any two technical theme feature vectors in different patent data sets, calculate the similarity between the two technical theme feature vectors; When the similarity is within a first threshold range, the two technical subject feature vectors are determined to be in an inheritance relationship; when the similarity is within a second threshold range, the two technical subject feature vectors are determined to be in a differentiation relationship; multiple technical evolution paths are constructed based on the inheritance relationship and the differentiation relationship, and the multiple technical evolution paths are constructed into a technical evolution path diagram based on the temporal relationship between each of the patent data sets.
10. The method according to claim 8, characterized in that The emerging technology indicators include technology innovation indicators, technology influence indicators, and interdisciplinary indicators. Determining the emerging technology indicators of the technology evolution path based on the first hypergraph network and the second hypergraph network, and determining whether the technology evolution path is an emerging technology path based on the emerging technology indicators of each technology evolution path includes: For any of the technical theme feature vectors, obtaining a technical innovation parameter, a technical influence parameter, and a disciplinary intersection parameter of the technical theme feature vector; For any of the technology evolution paths, determine the technology innovation index, technology influence index, and interdisciplinary index of the technology evolution path based on the technology innovation parameter, technology influence parameter, and interdisciplinary parameter of each of the technology themes in the technology evolution path; The technological innovation index, the technological influence index and the interdisciplinary index are weighted according to a preset first weight, a preset second weight and a preset third weight respectively, and whether the technological evolution path is an emerging technology path is determined based on the weighted results.
11. The method according to claim 10, characterized in that The technical innovation parameters, technical influence parameters and interdisciplinary degree parameters of the technical theme feature vector are obtained as follows: Determine the technical subject corresponding to the technical subject feature vector, obtain the number of first classification numbers and the number of second classification numbers of the technical subject, and determine the technical innovation index of the current technical subject feature vector based on the number of first classification numbers and the number of second classification numbers; Obtaining a first degree of the technical topic in the corresponding first hypergraph network and a second degree in the corresponding second hypergraph network, and determining the technical influence index of the technical topic feature vector based on the first degree and the second degree; The total number of first classification numbers of the technical subject and the total number of second classification numbers of the technical field to which the technical subject belongs are obtained, and the interdisciplinary index of the technical subject feature vector is determined based on the total number of the first classification numbers and the total number of the second classification numbers.
12. An identification device of emerging technology, characterized in that: The device comprises: A patent data acquisition unit is configured to acquire a plurality of patent data sets to be tested, and for any of the patent data sets, determine a plurality of key attributes of the patent data set, wherein the key attributes include at least a technical subject; a hypergraph network construction unit, configured to construct a first hypergraph network of the patent dataset, wherein the first hypergraph network includes a plurality of hyperedges, each of the hyperedges includes a plurality of the key attributes, and for any of the key attributes in the first hypergraph network, obtain an embedding vector of the key attribute; a hypergraph network updating unit, configured to obtain a plurality of embedding vectors of the first hypergraph network, input the embedding vectors into a hypergraph network model, and update the first hypergraph network using the hypergraph network model to obtain a second hypergraph network of the patent dataset; An emerging technology determination unit is used to obtain the first hypergraph network and the second hypergraph network of each of the patent data sets, and determine emerging technologies based on the first hypergraph network and the second hypergraph network.
Citation Information
Cited By
A subversive technology identification method based on superlink prediction
CN122757783A