Classification method mapping data set construction method and device based on thesis-patent pair

By constructing a paper-patent pair dataset and automatically determining the mapping relationship between scientific and technological classification labels, the problem of relying on expert judgment in existing technologies is solved, and the construction of a dataset with higher objectivity and scalability is achieved to support artificial intelligence applications.

CN120804319APending Publication Date: 2025-10-17BEIJING UNIV OF TECH

Patent Information

Application Number
CN202511088501.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The classification mapping dataset constructed by existing technologies is highly dependent on manual judgment by domain experts, resulting in high data objectivity, scalability, and update and maintenance costs, and is unable to effectively associate dispersed and heterogeneous scientific and technological resources.

Method used

By constructing a paper-patent pair dataset, the mapping relationship is automatically determined based on scientific taxonomy and technical taxonomy, directly related paper and patent combinations are screened out, and a taxonomy mapping dataset is constructed to reduce dependence on expert knowledge.

Benefits of technology

The objectivity and scalability of the taxonomy mapping dataset have been significantly improved, providing high-quality basic data support for large language model pre-training and cross-domain knowledge discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804319A_ABST
    Figure CN120804319A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a paper-patent pair-based classification method mapping data set construction method and device, and relates to the field of data mining, and the method comprises the steps: determining a paper data set and a patent data set of a target field; determining a paper-patent pair data set based on the patent data set and the paper data set; based on a scientific classification method and a technical classification method, supplementing a scientific classification method label corresponding to the paper and a technical classification method label corresponding to the patent in each paper-patent pair; for each paper-patent pair, determining a mapping relationship between a scientific classification method label corresponding to the academic paper and a technical classification method label corresponding to the patent in the paper-patent pair; and obtaining a classification method mapping data set based on a mapping relation associated with all paper-patent pairs. According to the embodiment of the invention, the data objectivity and expandability of the classification method mapping data set are improved, and high-quality basic data support is provided for artificial intelligence application scenes such as large language model pre-training and cross-domain knowledge discovery.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data mining, and in particular, the present disclosure relates to a mapping dataset construction method and device based on a paper-patent pair. BACKGROUND

[0002] With the transformation of the technology industry, frontier technologies such as artificial intelligence show a deeper dependence on high-quality structured datasets. The superposition of this industrial demand and technological iteration makes the strategic position of technological resources more prominent - these technological resources, including scientific data, knowledge models, and technical facilities, are the key matrix for cultivating high-quality structured datasets. However, at present, although the scale of technological resources is large, they are in a loose and isolated state, lacking effective interconnection, coordination, and configuration management, resulting in the phenomenon of "technological resource islands", which fails to fully realize the value and role of technological resources. Therefore, it is urgent to study how to associate dispersed, heterogeneous, complex, diverse, and massive technological resources to activate the "chemical reaction" between technological resources to generate new value, and further form a value circulation chain of "resource integration - high-quality data production - AI evolution - industry empowerment", ultimately promoting the positive feedback mechanism of production factor upgrading and technological revolution.

[0003] In terms of the association and aggregation of scientific literature resources, the current method is to construct a mapping dataset of classification methods between scientific classification methods and technical classification methods through scientific non-patent references (sNPR) or scientific corpus information in patent literature, so as to realize the aggregation of scientific literature resources. However, the mapping dataset of classification methods constructed by these methods highly depends on manual discrimination by experts, which not only is subject to the bias caused by subjective cognitive differences of experts, but also has obvious deficiencies in data objectivity, scalability, and update and maintenance costs. SUMMARY

[0004] The embodiment of the present disclosure provides a mapping dataset construction method and device of classification methods based on a paper-patent pair, which effectively solves the problem that the mapping dataset of classification methods constructed by the existing method is subject to subjective cognitive differences of experts, resulting in the deficiencies of the mapping dataset of classification methods in data objectivity, scalability, and update and maintenance costs. The technical solution provided by the present disclosure is as follows: According to the first aspect of the embodiment of the present disclosure, a mapping dataset construction method of classification methods based on a paper-patent pair is provided, and the method comprises: determining a paper dataset and a patent dataset of a target field; the paper dataset comprises a plurality of academic papers, and the patent dataset comprises a plurality of patents; determining a paper-patent pair dataset based on the patent dataset and the paper dataset; the paper-patent pair dataset comprises a plurality of paper-patent pairs, each of which comprises an academic paper and a patent; supplementing a scientific classification label corresponding to the academic paper and a technology classification label corresponding to the patent in each paper-patent pair based on the scientific classification and the technology classification; determining a mapping relationship between the scientific classification label corresponding to the academic paper and the technology classification label corresponding to the patent in each paper-patent pair; and obtaining a classification mapping dataset based on the mapping relationships associated with all paper-patent pairs; the classification mapping dataset is determined based on the scientific classification and the technology classification.

[0005] As an optional implementation, each academic paper in the paper dataset comprises a first data field; The determining of the paper-patent pair dataset based on the patent dataset and the paper dataset comprises: For each patent in the patent dataset, determining the cited literature of the patent and a second data field included in the cited literature, and determining whether the cited literature is a scientific non-patent citation according to the second data field; If the cited literature is a scientific non-patent citation, matching the content of the second data field with the content of the first data field of each academic paper in the paper dataset to determine an academic paper in the paper dataset that matches the cited literature; wherein the matching degree between the cited literature and the matched academic paper is greater than or equal to a first threshold value; determining the author of the academic paper that matches the cited literature and the inventor of the patent; if the author of the academic paper and the inventor of the patent are the same or partially the same, constructing a paper-patent pair based on the paper and the patent; determining the paper-patent pair dataset based on all paper-patent pairs associated with the patents in the patent dataset.

[0006] As an optional implementation, the determining of the paper-patent pair dataset based on the patent dataset and the paper dataset comprises: determining the inventor of each patent in the patent dataset, and determining the author of each academic paper in the paper dataset; matching the inventor of the patent with the author of the paper, and determining an academic inventor based on the matching result; determining an academic paper published by the academic inventor in the paper dataset, and determining a patent published by the academic inventor in the patent dataset; determine a similarity between the academic papers and the patents published by the academic inventor, and determine a paper-patent pair dataset based on the similarity.

[0007] As an optional implementation, the determining the paper-patent pair dataset based on the patent dataset and the paper dataset further includes: For each paper-patent pair in the paper-patent pair dataset, determine a confidence of the paper-patent pair based on a similarity between the patent author and the paper inventor in the paper-patent pair, and a text similarity between the patent and the paper. The confidence is used to represent a correlation between the paper and the patent in the paper-patent pair.

[0008] As an optional implementation, the determining the mapping relationship between the scientific classification tags and the technical classification tags in the paper-patent pair further includes: filter out the paper-patent pairs in the paper-patent pair dataset with a confidence less than or equal to a second threshold; and / or, filter out the paper-patent pairs in the paper-patent pair dataset that lack scientific classification tags and / or technical classification tags; and / or, filter out the paper-patent pairs in the paper-patent pair dataset that lack novelty, where the difference between the publication time of the paper and the application time of the patent exceeds a grace period for not losing novelty.

[0009] As an optional implementation, the determining, for each paper-patent pair, the mapping relationship between the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in the paper-patent pair includes: determine a hierarchical structure of the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in each paper-patent pair; the hierarchical structure includes at least one level, and each child level has only one parent level; determine a mapping relationship between the scientific classification tags of a leaf category and the technical classification tags of the leaf category based on the hierarchical structure; based on the hierarchical structure, aggregate the mapping relationship between the scientific classification tags of the leaf category and the technical classification tags of the leaf category level by level upwards to determine a mapping relationship between the scientific classification tags of each level and the technical classification tags of the corresponding level; determine the mapping relationship between the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in the paper-patent pair based on the mapping relationship between the scientific classification tags of each level and the technical classification tags of the corresponding level.

[0010] As an optional implementation, the determining, for each paper-patent pair, the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent of the paper-patent pair further includes: determining a hierarchical structure of the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent of each paper-patent pair; the hierarchical structure includes at least one level, and a child level corresponds to a plurality of parent levels; determining, based on the hierarchical structure, the scientific classification label of each level corresponding to the academic paper and the technical classification label of each level corresponding to the patent of the paper-patent pair; fully connecting the scientific classification label of each level and the technical classification label of the corresponding level in sequence; obtaining the mapping relationship between the scientific classification label of each level and the technical classification label of the corresponding level; determining, based on the mapping relationship between the scientific classification label of each level and the technical classification label of the corresponding level, the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent of the paper-patent pair.

[0011] According to a second aspect of the embodiments of the present disclosure, a classification mapping dataset construction device based on paper-patent pairs is provided, and the device includes: a first processing module configured to determine a paper dataset and a patent dataset of a target field; the paper dataset includes a plurality of academic papers, and the patent dataset includes a plurality of patents; a second processing module configured to determine a paper-patent pair dataset based on the patent dataset and the paper dataset; the paper-patent pair dataset includes a plurality of paper-patent pairs, each paper-patent pair including an academic paper and a patent; a third processing module configured to supplement, based on a scientific classification and a technical classification, a scientific classification label corresponding to an academic paper and a technical classification label corresponding to a patent of each paper-patent pair; a fourth processing module configured to determine, for each paper-patent pair, a mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent of the paper-patent pair; based on the mapping relationships associated with all paper-patent pairs, a classification mapping dataset is obtained; the classification mapping dataset is determined based on the scientific classification and the technical classification.

[0012] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method of any one of the first aspect.

[0013] According to a fourth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the steps of the method according to any one of the first aspect.

[0014] The technical solutions provided by the embodiments of the present disclosure have the following beneficial effects: The embodiments of the present disclosure provide a classification method mapping dataset construction method and device based on a paper-patent pair. The embodiments of the present disclosure screen out paper and patent combinations with direct correlation in the target field to form a paper-patent pair dataset, and construct a classification method mapping data based on a large-scale paper-patent pair data. The embodiments of the present disclosure do not need to rely on experts in the target field to manually distinguish the mapping relationship, effectively overcome the limitations brought by the over-reliance on expert knowledge in the construction method of the traditional classification method mapping dataset, significantly improve the objectivity and scalability of the classification method mapping dataset, and provide high-quality basic data support for artificial intelligence application scenarios such as large language model pre-training and cross-domain knowledge discovery. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings needed to be used in the description of the embodiments of the present disclosure will be briefly introduced.

[0016] Figure 1 A flowchart of a classification method mapping dataset construction method based on a paper-patent pair is provided for the embodiments of the present disclosure. Figure 2 A hierarchical aggregation mapping diagram based on a scientific classification method and a technical classification method is provided for the embodiments of the present disclosure. Figure 3 A single-layer full-connection mapping diagram based on a scientific classification method and a technical classification method is provided for the embodiments of the present disclosure. Figure 4 A first layer-by-layer full-connection mapping diagram based on a scientific classification method and a technical classification method is provided for the embodiments of the present disclosure. Figure 5 A second layer-by-layer full-connection mapping diagram based on a scientific classification method and a technical classification method is provided for the embodiments of the present disclosure. Figure 6 A mapping atlas based on a scientific classification method and a technical classification method is provided for the embodiments of the present disclosure. Figure 7 A structural diagram of a classification method mapping dataset construction device based on a paper-patent pair is provided for the embodiments of the present disclosure. Figure 8 A structural diagram of an electronic device is provided for the embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions of the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.

[0018] Those skilled in the art can understand that the singular forms "a", "an" and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the terms "comprise" and "include" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can mean that the element and the other element are connected through an intermediate element. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The term "and / or" used herein means that at least one of the items defined by the term, for example, "A and / or B" or "A, B" means that A is implemented, or B is implemented, or A and B are implemented.

[0019] In order to make the purposes, technical solutions and advantages of the present disclosure clearer, the embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings.

[0020] With the transformation of the technology industry, frontier technologies such as artificial intelligence show a deeper dependence on high-quality structured data sets. The superposition of this industrial demand and technological iteration has highlighted the strategic position of technological resources - these include scientific data, knowledge models, and technological facilities, which are the key precursors for cultivating high-quality structured data sets. However, current technological resources, although large in scale, are in a state of fragmentation and isolation, lacking effective interconnection, coordination, and configuration management, resulting in the phenomenon of "technological resource islands" and failing to fully realize the value and role of technological resources. Therefore, it is urgent to study how to associate dispersed, heterogeneous, complex, diverse, and massive technological resources to activate the "chemical reaction" between technological resources to generate new value, and further form a value circulation chain of "resource integration - high-quality data production - AI evolution - industry empowerment", ultimately promoting the upgrading of production factors and the formation of a positive feedback mechanism for technological revolution.

[0021] In the aspect of the association and aggregation of scientific and technological literature resources, the current method mainly constructs the classification mapping dataset between the scientific subject classification and the technology classification by the scientific non-patent reference (sNPR) in the patent literature or the corpus information of the scientific and technological purpose, so as to realize the aggregation of the scientific and technological literature resources. However, the classification mapping dataset constructed by these methods highly depends on the manual discrimination of the field experts, which is not only subject to the deviation caused by the subjective cognitive difference of the experts, but also has obvious deficiencies in the aspects of data objectivity, scalability and updating and maintenance cost.

[0022] The classification mapping dataset construction method and device based on the paper-patent pair provided by the present disclosure aim to solve at least one of the above technical problems of the prior art.

[0023] The technical solutions of the embodiments of the present disclosure and the technical effects of the technical solutions of the present disclosure will be described below through the description of several exemplary embodiments. It should be pointed out that the following embodiments can be mutually referenced, borrowed or combined. For the same terms, similar features and similar implementation steps in different embodiments, they will not be described repeatedly.

[0024] Figure 1 A flowchart of a classification mapping dataset construction method based on a paper-patent pair provided by an embodiment of the present disclosure is shown in FIG. 1. As shown in FIG. 1, the method comprises the following steps. Figure 1 S101, determining a paper dataset and a patent dataset of a target field; the paper dataset comprises a plurality of academic papers, and the patent dataset comprises a plurality of patents.

[0025] Specifically, in the embodiment of the present disclosure, academic papers of all fields or a specific field are collected, denoted as a paper dataset R. The data fields of each academic paper in the paper dataset include but are not limited to: paper title, abstract, author name, author institution, journal name, publication year, volume number, issue number, page number and subject label to which the paper belongs. At the same time, patent data of all fields or a specific field are collected, denoted as a patent dataset T. The data fields of each patent in the patent dataset include but are not limited to: patent title, abstract, inventor name, patentee, publication date, cited literature and IPC classification label.

[0026] Specifically, in the embodiment of the present disclosure, the paper dataset can be academic papers on the OpenAlex website; wherein, OpenAlex is a global free academic literature database developed by the non-profit organization OurResearch; and the patent dataset can be patent data of the United States Patent and Trademark Office (USPTO) with publication years of 2000-2010.

[0027] ​Specifically, in the embodiments of the present disclosure, the target field can be at least one of a natural science field, a social science field, and an engineering technology field.

[0028] S102, based on the patent dataset and the paper dataset, determine a paper-patent pair dataset; the paper-patent pair dataset includes a plurality of paper-patent pairs, and each paper-patent pair includes an academic paper and a patent.

[0029] Specifically, in the embodiments of the present disclosure, there is a correlation between the academic paper and the patent describing the same research. Therefore, the paper-patent pair dataset can be constructed based on the academic paper and the patent describing the same research.

[0030] Specifically, in the embodiments of the present disclosure, the paper-patent pair dataset includes a plurality of paper-patent pairs, and each paper-patent pair includes an academic paper and a patent; and the academic paper and the patent have a correlation. For example: the author of the academic paper and the inventor of the patent are the same, and the text similarity between the academic paper and the patent is greater than or equal to a preset threshold; or, the academic paper is cited by the patent, and the author of the academic paper and the inventor of the patent are the same or partially the same.

[0031] S103, based on the scientific classification method and the technical classification method, supplement the scientific classification method label corresponding to the academic paper and the technical classification method label corresponding to the patent in each paper-patent pair.

[0032] Specifically, in the embodiments of the present disclosure, the scientific classification method includes but is not limited to Web Of Science subject classification method, OpenAlex Topic, OpenAlex Concept, and Chinese Library Classification (CLC); and the technical classification method includes but is not limited to International Patent Classification (IPC) and Cooperative Patent Classification (CPC).

[0033] Specifically, in the embodiments of the present disclosure, based on the scientific classification method and the technical classification method, the scientific classification method label is supplemented for the academic paper in each paper-patent pair, and the technical classification method label is supplemented for the patent in the paper-patent pair. For example: when the scientific classification method is OpenAlex Concept classification method, the OpenAlex Concept system label of the paper in the paper-patent pair is supplemented; and when the technical classification method is International Patent Classification (IPC), the IPC label of the patent in the paper-patent pair is supplemented.

[0034] S104, for each paper-patent pair, determine the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair; based on the mapping relationship associated with all paper-patent pairs, obtain a classification mapping dataset; the classification mapping dataset is determined based on the scientific classification and the technical classification.

[0035] Specifically, in the embodiments of the present disclosure, the hierarchical structure of the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair is determined. If the hierarchical structure is a "tree structure", that is, the parent hierarchical category corresponding to the child hierarchical category is unique, the mapping relationship between the scientific classification label of the leaf category and the technical classification label of the leaf category can be determined based on the hierarchical structure; then, the mapping relationship between the scientific classification label of each level and the technical classification label of the corresponding level is determined by aggregating upwards level by level, as the mapping relationship between the scientific classification label corresponding to the paper and the technical classification label corresponding to the patent in the paper-patent pair.

[0036] Specifically, in the embodiments of the present disclosure, if the hierarchical structure does not present a tree structure, that is, the parent hierarchical category corresponding to the child hierarchical category can have multiple, it is difficult or biased to aggregate upwards; therefore, based on the hierarchical structure, the scientific classification label of each level corresponding to the academic paper and the technical classification label of each level corresponding to the patent in the paper-patent pair can be determined; the scientific classification label of each level is sequentially fully connected to the technical classification label of the corresponding level; the mapping relationship between the scientific classification label of each level and the technical classification label of the corresponding level is obtained, thereby determining the mapping relationship between the scientific classification label corresponding to the paper and the technical classification label corresponding to the patent.

[0037] The embodiments of the present disclosure screen out the paper and patent combination with direct association in the target field, constitute a paper-patent pair dataset, and determine the mapping relationship between the scientific classification label and the technical classification label based on a large-scale paper-patent pair data, without relying on experts in the target field to manually determine the mapping relationship, effectively overcoming the limitations brought by the excessive dependence on expert knowledge in the traditional classification mapping dataset construction method, significantly improving the objectivity and scalability of the classification mapping dataset, and providing high-quality basic data support for artificial intelligence application scenarios such as large language model pre-training and cross-domain knowledge discovery.

[0038] On the basis of each of the above embodiments, as an optional embodiment, each academic paper in the paper dataset comprises a first data field; Based on the patent dataset and the paper dataset, a paper-patent pair dataset is determined, comprising: For each patent in the patent dataset, determine the citation of the patent and the second data field included in the citation, and determine whether the citation is a scientific non-patent citation according to the second data field; If the citation is a scientific non-patent citation, match the content of the second data field with the content of the first data field of each academic paper in the paper dataset to determine the academic paper in the paper dataset that matches the citation; wherein the matching degree of the citation and the matched academic paper is greater than or equal to a first threshold value; Determine the author of the matched academic paper and the inventor of the patent; if the author of the academic paper and the inventor of the patent are the same or partially the same, construct a paper-patent pair based on the paper and the patent; Based on all paper-patent pairs associated with all patents in the patent dataset, determine a paper-patent pair dataset.

[0039] Specifically, in the embodiments of the present disclosure, for each patent in the patent dataset, scientific non-patent citations in the patent citation are identified, and the methods used include but are not limited to: first, identify the citations in the patent text by machine learning, rule matching or natural language processing and other technical means, parse the citations into title, author name, journal name, document type identifier, publication year, volume number, issue number and page number, etc. Fields, and these fields are used as second data fields; second, cross-verify the parsed second data fields with the citations of the patent to correct and improve the parsing results; for example: when the academic paper is cited in the patent text, only the name and author of the academic paper may be listed, and other field information of the academic paper is omitted; therefore, the complete field information of the academic paper can be obtained from the reference literature page of the patent, and the field information of the academic paper is supplemented through cross-verification, so as to improve the second data field. Third, determine whether the citation is a scientific non-patent citation based on the second data field; including but not limited to: when the journal name, publication year, volume number, issue number, etc. Fields exist in the second data field, it can be determined that the citation is a scientific non-patent citation; when the patent number, disclosure number, patent type, etc. Fields exist in the second data field, it can be determined that the citation is not a scientific non-patent citation. For example: when the journal name is “Journal of the American Chemical Society” (Chinese name: “American Chemical Society”), it can be determined that the citation is an academic paper in the field of chemistry, and the field of chemistry belongs to the field of natural sciences, therefore, the citation is a scientific non-patent citation; if the document type identifier contains “Patent No.”, “USPTO”, “CN” and other identifiers, it can be determined that the citation is a patent, therefore, the citation is not a scientific non-patent citation.

[0040] Specifically, in the embodiments of the present disclosure, after determining that the reference document is a scientific non-patent reference, the second data field can be matched with the first data field of the academic papers in the paper data set, so as to determine the academic papers in the paper data set that match the reference document.

[0041] In some optional embodiments, the matching degree of the reference document and each academic paper in the paper data set can be determined according to the number, accuracy and other indicators of the fields in the second data field that are successfully matched with the fields in the first data field, and the academic paper with a matching degree greater than or equal to a first threshold value is determined as the academic paper that matches the reference document.

[0042] Specifically, in the embodiments of the present disclosure, the first data field of each academic paper in the paper data set includes but is not limited to: paper title, abstract, author name, author institution, journal name, publication year, volume number, issue number, page number and subject label to which the paper belongs. As can be seen, the first data field and the second data field overlap. Therefore, the matching degree of the reference document and each academic paper in the paper data set can be determined by matching the content of the first data field and the content of the second data field. If the matching degree is greater than or equal to the first threshold value, it means that the reference document has a high correlation with a certain academic paper in the paper data set, and the academic paper is determined as the academic paper that matches the reference document.

[0043] Specifically, in the embodiments of the present disclosure, if a certain patent explicitly references a certain paper, and the inventor of the patent partially or completely overlaps with the author of the scientific non-patent reference paper, a paper-patent pair is formed by the patent and the paper. All paper-patent pairs associated with the patents in the patent data set are aggregated to obtain a paper-patent pair data set. It should be noted that if a certain patent references multiple papers, a paper-patent pair is formed by the patent and each referenced paper.

[0044] In the embodiments of the present disclosure, by matching the second data field with the first data field and setting the first threshold value, non-academic documents (such as technical manuals, news, etc.) or low-correlation references are excluded, ensuring that the screened papers are core documents that are truly referenced by patents and belong to the academic research category, laying a foundation for subsequent construction of classification method mapping data set.

[0045] Based on the above embodiments, as an optional embodiment, based on the patent data set and the paper data set, the paper-patent pair data set is determined, including: determining the inventor of each patent in the patent data set, and determining the author of each academic paper in the paper data set; matching the inventor of the patent with the author of the paper, and determining the academic inventor based on the matching result; Identify the academic papers published by academic inventors in the paper dataset, and identify the patents published by academic inventors in the patent dataset; Determine the similarity between academic papers and patents published by academic inventors, and based on the similarity, identify a dataset of paper-patent pairs.

[0046] Specifically, in the embodiments of the present disclosure, the paper-patent pair dataset is determined by identifying academic inventors in the patent dataset and the paper dataset. For example: First, the paper author and patent inventor data are preprocessed, including name standardization, institution name standardization, research field keyword extraction, etc. Second, the names of authors and inventors are disambiguated. Specific operations include but are not limited to: rule-based disambiguation, combining information such as institution information, cooperation network, research field, keywords, etc. to assist in identification and disambiguation; author / inventor clustering disambiguation method based on co-authorship network. Third, the disambiguated paper author list is associated and matched with the patent inventor list to identify academic inventors who appear in both the paper and the patent.

[0047] Specifically, in the disclosed embodiment, for an identified academic inventor, all of their published papers and patent applications are summarized, and the similarity between the papers and patents in terms of text, such as titles and abstracts, is calculated. Based on this similarity, a paper-patent pair is constructed. For example, if the similarity between a paper published by an academic inventor and a patent application is greater than or equal to a preset threshold, it indicates that the technical field and direction studied by the paper and the patent are the same or highly similar. The paper and the patent are then associated to obtain a paper-patent pair. Then, the paper-patent pairs associated with all academic inventors are summarized to obtain a paper-patent pair dataset.

[0048] For example, for a paper-patent pair in the paper-patent pair dataset, the English title of the paper is "Development of an underwater mass-spectrometry system for in situ chemical analysis", and the Chinese title is "Development of an underwater in situ chemical analysis mass spectrometry system"; the patent publication number is US6727498B2, the English title is "Portable underwater mass spectrometer", and the Chinese title is "Portable underwater mass spectrometer"; the author of the paper and the inventor of the patent highly overlap, and the titles and abstracts of the paper and the patent are highly similar. The paper and the patent are associated to obtain a paper-patent pair.

[0049] In the embodiments of the present disclosure, the academic inventors are determined through matching of patent inventors and paper authors, which can accurately screen the subjects active in both academic research and technical invention, and define the core authors for the paper-patent pair dataset, thereby avoiding irrelevant author interference in the literature.

[0050] Based on the above embodiments, as an optional embodiment, the paper-patent pair dataset is determined based on the patent dataset and the paper dataset, and further includes: For each paper-patent pair in the paper-patent pair dataset, the confidence of the paper-patent pair is determined based on the similarity between the patent authors and the paper inventors in the paper-patent pair, and the text similarity between the patent and the paper. The confidence is used to represent the association between the paper and the patent in the paper-patent pair.

[0051] Specifically, in the embodiments of the present disclosure, for each paper-patent pair in the paper-patent pair dataset, the confidence, i.e., the reliability of the association between the two, needs to be further calculated. When calculating the confidence, the following two dimensions can be referred to: Subject similarity: comparing the matching degree between the authors (such as inventors or invention teams) of the patent and the inventors (such as authors or research teams) of the paper. For example, if they are completely the same or partially overlap, the subject similarity is high.

[0052] Text similarity: comparing the similarity of the text content of the patent and the paper. For example, if they discuss the same technical field, or the technical solution of the patent is based on the research results of the paper, the text similarity between the patent and the paper is high.

[0053] Specifically, in the embodiments of the present disclosure, the higher the subject similarity and the text similarity, the closer the association between the paper and the patent, and the higher the confidence. The confidence is a quantitative description of the association strength of the paper-patent pair. The higher the confidence, the more closely related the paper and the patent are in the creation subject or content, and the more reliable the association relationship is. The lower the confidence, the more likely it is a formal reference (such as accidental mention), and the actual association relationship is weak.

[0054] For example, in a paper-patent pair, the English label of the paper is "Protein Kinase Expression during Murine Mammary Development," and the Chinese title is "Protein Kinase Expression during Murine Mammary Development." The patent publication number is US7368113B2, and the English title is "Hormonally up-regulated, neu-tumor-associated kinase," and the Chinese title is "Hormonally up-regulated, neu-tumor-associated kinase." The confidence level of the paper-patent pair is determined based on the subject matter and text similarity between the paper and patent.

[0055] In the disclosed embodiment, the confidence level is calculated by using the “similarity between the patent author and the paper inventor” and the “text similarity between the patent and the paper” to convert the correlation between the paper and the patent into a measurable quantitative indicator, thus providing data support for the subsequent screening of paper-patent pairs.

[0056] Based on the above embodiments, as an optional embodiment, determining the mapping relationship between scientific classification labels and technical classification labels in a paper-patent pair also includes: Filtering out paper-patent pairs in the paper-patent pair dataset whose confidence is less than or equal to a second threshold; and / or, Filter out paper-patent pairs with missing scientific taxonomy labels and / or technical taxonomy labels in the paper-patent pair dataset; and / or; Filter out paper-patent pairs lacking novelty in the paper-patent pair dataset; the difference between the publication time of the paper and the patent application time in the paper-patent pairs lacking novelty exceeds the grace period for not losing novelty.

[0057] Specifically, in the embodiment of the present disclosure, after determining the paper-patent pair dataset, filtering the paper-patent pairs in the paper-patent pair dataset can be divided into the following situations: Filter by confidence level, removing paper-patent pairs with a confidence level less than or equal to a second threshold. The second threshold is a preset value. When the confidence level of a paper-patent pair is less than or equal to the second threshold, it indicates a weak correlation. Therefore, it is removed from the paper-patent pair dataset, leaving only paper-patent pairs with strong correlations. For example, paper-patent pairs with a confidence level less than or equal to 0.8 can be removed.

[0058] Filtering by classification label integrity, filtering out papers-patent pairs missing scientific taxonomy labels and / or technical taxonomy labels, wherein scientific system taxonomy labels and technical system taxonomy labels are important basis for subsequent construction of taxonomy mapping dataset. If a papers-patent pair is missing one of the labels, or both labels, the papers-patent pair cannot participate in the construction of the taxonomy mapping dataset.

[0059] Filtering by time relevance, filtering out papers-patent pairs lacking novelty, where "lack of novelty" has a clear time standard: when the interval between the publication time of the paper and the application time of the patent in the papers-patent pair exceeds the "non-loss of novelty grace period" (for example: 6 months or 12 months), the papers-patent pair is considered to lack direct timeliness correlation. For example: a patent is applied for 2 years after the paper is published, then the patent does not belong to the immediate technology transfer of the paper; therefore, such papers-patent pairs with too long time intervals can be excluded.

[0060] It should be noted that in the embodiments of the present application, "and / or" means that the above-mentioned three filtering methods of papers-patent pairs can be used alone or in combination.

[0061] In the embodiments of the present disclosure, by filtering the papers-patent pairs, the irrelevant or low-quality samples of the papers-patent pair dataset are greatly reduced, the relevance, integrity and timeliness of the core samples are guaranteed, and the accuracy of the constructed taxonomy mapping dataset is ensured.

[0062] On the basis of the above-mentioned embodiments, as an optional embodiment, for each papers-patent pair, the mapping relationship between the scientific taxonomy label corresponding to the academic paper and the technical taxonomy label corresponding to the patent in the papers-patent pair is determined, comprising: determining the hierarchical structure of the scientific taxonomy label corresponding to the academic paper and the technical taxonomy label corresponding to the patent in each papers-patent pair; the hierarchical structure includes at least one level, and each sub-level has only one parent level; based on the hierarchical structure, determining the mapping relationship between the scientific taxonomy label of the leaf category and the technical taxonomy label of the leaf category; based on the hierarchical structure, the mapping relationship between the scientific taxonomy label of the leaf category and the technical taxonomy label of the leaf category is aggregated level by level upwards to determine the mapping relationship between the scientific taxonomy label of each level and the technical taxonomy label of the corresponding level; based on the mapping relationship between the scientific taxonomy label of each level and the technical taxonomy label of the corresponding level, determining the mapping relationship between the scientific taxonomy label corresponding to the academic paper and the technical taxonomy label corresponding to the patent in the papers-patent pair.

[0063] Specifically, in the embodiments of the present disclosure, the hierarchical structure is similar to a "tree structure", and contains multiple levels of categories, each sub-level category has only one parent-level category, and there is no cross-ownership between the sub-level category and the parent-level category; wherein, the "leaf category" refers to the specific category at the bottom of the hierarchical structure, which cannot be further subdivided, and can be regarded as the category at the bottom of the hierarchical structure.

[0064] Specifically, in the embodiments of the present disclosure, for each publication-patent pair, based on the hierarchical structure of the publication-patent pair, the scientific classification label of the leaf category and the technical classification label of the leaf category are determined; based on the scientific classification label of the leaf category and the technical classification label of the leaf category, a mapping relationship between the scientific classification label of the leaf category and the technical classification label of the leaf category is established; then, the mapping relationship between the scientific classification label of each level and the technical classification label of the corresponding level is obtained by aggregating step by step upwards along the hierarchical structure. It can be understood that for the classification of the "tree structure", the relationship between the sub-level and the parent-level is determined and unique, so the mapping relationship between the scientific classification label and the technical classification label of the upper level can be directly determined by using the aggregation method, and the mapping relationship between the scientific classification label and the technical classification label of other levels does not need to be determined by using the step-by-step full connection mapping method.

[0065] As shown in Figure 2 , for the patent literature and scientific literature describing the same research, it can be regarded as a publication-patent pair (Publication-Patent Pair, PPP). The technical classification label corresponding to the patent literature and the scientific classification label corresponding to the scientific literature are determined; wherein, the hierarchical structure of the technical classification corresponding to the patent literature is three levels, which are: three-level technical category, two-level technical category and one-level technical category; the hierarchical structure of the scientific classification of the scientific literature is three levels, which are: three-level scientific category, two-level scientific category and one-level scientific category. The three-level technical category and the three-level scientific category are the respective bottom level, which cannot be further subdivided, i.e. the leaf category; based on the scientific classification leaf category of the publication in the publication-patent pair and the technical classification leaf category of the patent, a scientific-technical three-level classification system mapping is constructed; since the sub-level category has only one parent-level category, the mapping relationship between the three-level technical category and the three-level scientific category can be aggregated upwards to obtain the scientific-technical two-level classification system mapping and the scientific-technical one-level classification system mapping.

[0066] It should be noted that Figure 2 the classification system mapping at each level can be regarded as the mapping between the scientific classification and the technical classification at the corresponding level.

[0067] In the embodiments of the present disclosure, based on the mapping relationship between the leaf target tags, the mapping relationship between the target tags of each hierarchical level is obtained by further aggregating the mapping relationship upward. The "aggregated association" feature of the hierarchical mapping makes it unnecessary to reconstruct the entire mapping relationship when adjusting the classification locally, thereby effectively improving the scalability of the classification mapping dataset.

[0068] Based on the above embodiments, as an optional embodiment, for each paper-patent pair, the mapping relationship between the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in the paper-patent pair is determined, and the method further includes: determining the hierarchical structure of the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in each paper-patent pair; the hierarchical structure includes at least one hierarchical level, and a child hierarchical level corresponds to a plurality of parent hierarchical levels; based on the hierarchical structure, determining the scientific classification tags of each hierarchical level corresponding to the academic paper and the technical classification tags of each hierarchical level corresponding to the patent in the paper-patent pair; sequentially performing full connection mapping between the scientific classification tags of each hierarchical level and the technical classification tags of the corresponding hierarchical level; and obtaining the mapping relationship between the scientific classification tags of each hierarchical level and the technical classification tags of the corresponding hierarchical level; based on the mapping relationship between the scientific classification tags of each hierarchical level and the technical classification tags of the corresponding hierarchical level, determining the mapping relationship between the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in the paper-patent pair.

[0069] Specifically, in the embodiments of the present disclosure, part of the classification does not present a tree structure (such as a spindle structure, in which the number of categories of the intermediate hierarchical level is more than that of the root node and the leaf node), the parent category of its subcategory may have multiple, and the upward aggregation operation is more difficult, or may cause deviation, so the mapping can be realized by layer-by-layer full connection association.

[0070] Specifically, in the embodiments of the present disclosure, for each paper-patent pair, "full connection mapping" means that in each hierarchical level, all scientific classification tags under the hierarchical level are fully associated with the technical classification tags of the corresponding hierarchical level, that is, each scientific classification tag under the same hierarchical level is mapped to each technical classification tag under the corresponding hierarchical level. The mapping relationship obtained by full connection of the tags of each hierarchical level is summarized, and finally the mapping relationship between the scientific classification tags corresponding to the academic paper and the technical classification tags corresponding to the patent in the paper-patent pair is formed.

[0071] For example, Figure 3As shown in FIG. 1, the scientific classification corresponding to the scientific literature has only one level of tags, and the technical classification corresponding to the patent literature also has only one level of tags; therefore, the mapping relationship between the scientific classification tags and the technical classification tags is determined by a single-layer full connection. Specifically, the scientific categories of the scientific literature include genes, antibodies, and receptors; the technical categories of the patent literature include A61K (medicine) and G01N (analysis); by single-layer full connection of the scientific categories and the technical categories, the mapping relationship is obtained, including genes-A61K, antibodies-A61K, and receptors-A61K, and genes-G01N, antibodies-G01N, and receptors-G01N; it can be understood that single-layer full connection mapping is to map each scientific category with each technical category to obtain the corresponding mapping relationship.

[0072] As shown in FIG. 2, the scientific classification corresponding to the scientific literature has multiple levels of tags, and the technical classification corresponding to the patent literature also has multiple levels of tags; the mapping relationship between the scientific classification tags and the technical classification tags is determined by layer-by-layer full connection. Specifically, the scientific categories of the scientific literature include: a first-level scientific category of biology, a third-level scientific category of genes, and a five-level scientific category of RNA interference; the technical categories of the patent literature include a five-level technical category C12N 15 / 113 and a four-level technical category C07K 16 / 00. Figure 4

[0073] Among them, the five-level technical category C12N 15 / 113 can be split into a first-level technical category C, a third-level technical category C12N, and a five-level technical category C12N 15 / 113; the four-level technical category can be split into a first-level technical category C and a third-level technical category C07K.

[0074] The scientific categories and the five-level technical category C12N 15 / 113 are mapped by layer-by-layer full connection to obtain the following mapping relationship: The first-level mapping relationship is biology-C; the third-level mapping relationship is genes-C12N; and the five-level mapping relationship is RNA interference-C12N 15 / 113.

[0075] The scientific categories and the four-level technical category C07K 16 / 00 are mapped by layer-by-layer full connection to obtain the following mapping relationship: The first-level mapping relationship is biology-C; and the third-level mapping relationship is genes-C07K.

[0076] It should be noted that, since the four-level technical category C07K 16 / 00 cannot be split into a five-level technical category, there is no five-level mapping relationship between the five-level scientific category RNA interference and the four-level technical category C07K 16 / 00.

[0077] As shown in FIG. 3, the scientific classification corresponding to the scientific literature has only one level of tags, and the technical classification corresponding to the patent literature also has only one level of tags; therefore, the mapping relationship between the scientific classification tags and the technical classification tags is determined by a single-layer full connection. Specifically, the scientific categories of the scientific literature include genes, antibodies, and receptors; the technical categories of the patent literature include A61K (medicine) and G01N (analysis); by single-layer full connection of the scientific categories and the technical categories, the mapping relationship is obtained, including genes-A61K, antibodies-A61K, and receptors-A61K, and genes-G01N, antibodies-G01N, and receptors-G01N; it can be understood that single-layer full connection mapping is to map each scientific category with each technical category to obtain the corresponding mapping relationship. Figure 5 ​As shown, the scientific classification corresponding to the scientific literature has multiple levels, and the technical classification corresponding to the patent literature also has multiple levels; the mapping relationship between the scientific classification and the technical classification is determined through the full connection relationship at each level. Specifically, the scientific categories of the scientific literature include: the first-level scientific category: biology, the second-level scientific category: cell biology, the third-level scientific category: kinase, the fourth-level scientific category: mammary gland development, and the fifth-level scientific category: transgenic mouse; the technical categories of the patent literature include: the fifth-level technical category C07H 21 / 002 and the fourth-level technical category A61K 48 / 00.

[0078] Among them, the fifth-level technical category C07H 21 / 002 can be split into the first-level technical category C, the second-level technical category C07, the third-level technical category C07H, the fourth-level technical category C07H 21 / 00 and the fifth-level technical category C07H 21 / 02. The fourth-level technical category A61K 48 / 00 can be split into the first-level technical category A, the second-level technical category A61, the third-level technical category A61K and the fourth-level technical category A61K 48 / 00.

[0079] The scientific categories and the fifth-level technical category C07H 21 / 02 can be mapped in a full connection manner at each level to obtain the following mapping relationship: The first-level mapping relationship: biology-C; the second-level mapping relationship: cell biology-C07; the third-level mapping relationship: kinase-C07H; the fourth-level mapping relationship: mammary gland development-C07H 21 / 00; and the fifth-level mapping relationship: transgenic mouse-C07H 21 / 02.

[0080] The scientific categories and the fourth-level technical category A61K 48 / 00 can be mapped in a full connection manner at each level to obtain the following mapping relationship: The first-level mapping relationship: biology-C; the second-level mapping relationship: cell biology-A61; the third-level mapping relationship: kinase-A61K; and the fourth-level mapping relationship: mammary gland development-A61K 48 / 00.

[0081] It should be noted that, since the fourth-level technical category A61K 48 / 00 cannot be split into a fifth-level technical category, there is no fifth-level mapping relationship between the fifth-level scientific category transgenic mouse and the fourth-level technical category A61K 48 / 00.

[0082] It should also be noted that, Figures 3 to 5 Each level of mapping relationship in the above table can be regarded as the mapping between the scientific classification and the technical classification at the corresponding level.

[0083] In the disclosed embodiment, the layer-by-layer fully connected association does not need to forcibly determine the unique affiliation between the sub-category and the parent category. Instead, a mapping relationship between all scientific classification labels and technical classification labels at each level is directly established within the level. This fundamentally avoids the deviation caused by forced aggregation for the "multiple affiliations and difficult aggregation" characteristics of the non-tree structure.

[0084] In some optional embodiments, the scientific classification method adopts the OpenAlex concept classification method, and the technical classification method adopts the International Patent Classification IPC. The mapping diagram of the first-level categories of the OpenAlex concept classification method and the first-level categories of the International Patent Classification is as follows: Figure 6 shown; through Figure 6 As can be seen, scientific fields such as biology, chemistry, medicine, computer science, and materials science frequently interact deeply with technical fields such as Section C (Chemistry; Metallurgy), Section A (Necessities of Life), Section G (Physics), and Section H (Electrical Engineering). This phenomenon reflects a prominent trend in the development of contemporary science and technology: deep multidisciplinary integration and cross-disciplinary innovation. In contrast, humanities and social sciences such as art, history, economics, and sociology have less direct correlation with technical fields. This is primarily because research findings in these fields are often difficult to directly apply to technological innovation practices. Furthermore, technical fields such as Section D (Textiles; Papermaking), Section E (Fixed Structures), and Section F (Mechanical Engineering; Lighting; Heating; Weapons; Demolitions) have relatively little interaction with scientific fields. This is because innovation in these technical fields tends to rely more on engineering experience, practical skills, and the accumulation of traditional technologies rather than breakthroughs in basic scientific theory.

[0085] Figure 7 A schematic diagram of a structure of a device for constructing a classification mapping dataset based on paper-patent pairs provided in an embodiment of the present disclosure is shown in FIG. Figure 7 As shown, the device includes: a first processing module 7001, a second processing module 7002, a third processing module 7003, and a fourth processing module 7004. The first processing module 7001 is used to determine a paper dataset and a patent dataset in a target field; the paper dataset includes multiple academic papers, and the patent dataset includes multiple patents; The second processing module 7002 is used to determine a paper-patent pair dataset based on the patent dataset and the paper dataset; the paper-patent pair dataset includes multiple paper-patent pairs, and each paper-patent pair includes an academic paper and a patent; The third processing module 7003 is used to supplement the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in each paper-patent pair based on the scientific classification and technical classification; The fourth processing module 7004 is used to determine, for each paper-patent pair, the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair; based on the mapping relationship associated with all paper-patent pairs, obtain a classification mapping dataset; the classification mapping dataset is determined based on the scientific classification and the technical classification.

[0086] The apparatus for constructing a classification mapping dataset based on scientific classification and technological classification provided in the embodiments of the present disclosure can execute the method for constructing a classification mapping dataset based on paper-patent pairs provided in the embodiments of the present disclosure, and its implementation principles are similar. The actions performed by each module in the apparatus for constructing a scientific and technological classification mapping dataset based on paper-patent pairs provided in each embodiment of the present disclosure correspond to the steps in the method for constructing a classification mapping dataset based on paper-patent pairs provided in each embodiment of the present disclosure. For the detailed functional description of each module in the apparatus for constructing a scientific and technological classification mapping dataset based on paper-patent pairs provided in the embodiments of the present disclosure, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.

[0087] The disclosed embodiments screen out directly related paper and patent combinations in the target field to form a paper-patent pair dataset, and implement mapping of scientific taxonomies and technical taxonomies based on large-scale paper-patent pair data without relying on experts in the target field for manual judgment. This effectively overcomes the limitations of traditional taxonomy mapping dataset construction methods that rely too much on expert knowledge, significantly improves the objectivity and scalability of the taxonomy mapping dataset, and provides high-quality basic data support for artificial intelligence application scenarios such as large language model pre-training and cross-domain knowledge discovery.

[0088] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Figure 8 As shown, electronic device 4000 includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.

[0089] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logical blocks, modules and circuits described in connection with the present disclosure. The processor 4001 can also be a combination of implementing computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0090] The bus 4002 can include a path that transmits information between the above-mentioned components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, and the like. For convenience of representation, Figure 8 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0091] The memory 4003 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, and the like), a magnetic disk storage medium, other magnetic storage device, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation.

[0092] The memory 4003 is configured to store a computer program for implementing the embodiments of the present disclosure, and the processor 4001 is configured to control the execution of the computer program stored in the memory 4003. The processor 4001 is configured to execute the computer program stored in the memory 4003 to implement the steps of the foregoing method embodiments.

[0093] The electronic device can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like. Figure 8 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0094] The embodiments of the present disclosure provide a computer readable storage medium having stored thereon a computer program, which, when executed by a processor, can implement the steps and corresponding contents of the foregoing method embodiments.

[0095] It should be noted that the computer readable medium of the present disclosure described above can be a computer readable signal medium or a computer readable medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more conductive wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, an optical cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0096] The embodiments of the present disclosure further provide a computer program product comprising a computer program, which, when executed by a processor, can implement the steps and corresponding contents of the foregoing method embodiments. Compared with the prior art, the following technical effects can be achieved: The terms "first", "second", "third", "fourth", "1", "2", and the like (if any) in the description, claims, and drawings of the present disclosure, and the above-described drawings are used to distinguish similar objects, and do not have to be used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described.

[0097] It should be understood that, although the flowcharts of the embodiments of the present disclosure indicate the respective operation steps by arrows, the implementation order of the steps is not limited to the order indicated by the arrows. Unless otherwise specified herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be executed in other orders as required. In addition, part or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on the actual implementation scenario. Part or all of these sub-steps or stages can be executed at the same time, and each of these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.

[0098] The above is only an optional implementation manner of some implementation scenarios of the present disclosure, and it should be pointed out that, for ordinary skilled persons in the technical field, other similar implementation manners based on the technical concept of the present disclosure without departing from the technical concept of the present disclosure also belong to the protection scope of the embodiments of the present disclosure.

Claims

1. A method for constructing a taxonomy mapping dataset based on paper-patent pairs, characterized by: The method comprises: Determine a paper dataset and a patent dataset in the target field; the paper dataset includes multiple academic papers, and the patent dataset includes multiple patents; Determine a paper-patent pair dataset based on the patent dataset and the paper dataset; the paper-patent pair dataset includes multiple paper-patent pairs, each paper-patent pair includes an academic paper and a patent; Based on scientific and technical classifications, supplement the scientific classification labels corresponding to academic papers and the technical classification labels corresponding to patents in each paper-patent pair; For each paper-patent pair, determine the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair; based on the mapping relationship associated with all paper-patent pairs, obtain a classification mapping dataset; the classification mapping dataset is determined based on the scientific classification and the technical classification.

2. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to claim 1, characterized in that: Each academic paper in the paper dataset includes a first data field; Determining a paper-patent pair dataset based on the patent dataset and the paper dataset includes: For each patent in the patent dataset, determining a cited document of the patent and a second data field included in the cited document, and determining whether the cited document is a scientific non-patent citation based on the second data field; If the cited document is a scientific non-patent citation, matching the content of the second data field with the content of the first data field of each academic paper in the paper dataset to determine the academic paper in the paper dataset that matches the cited document; wherein the matching degree between the cited document and the matching academic paper is greater than or equal to a first threshold; Determine the authors of the academic papers and the inventors of the patents that match the cited documents; if the authors of the academic papers are the same or partially the same as the inventors of the patents, construct a paper-patent pair based on the papers and the patents; Based on all the paper-patent pairs associated with the patents in the patent dataset, a paper-patent pair dataset is determined.

3. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to claim 1, characterized in that: Determining a paper-patent pair dataset based on the patent dataset and the paper dataset includes: Determine the inventor of each patent in the patent dataset, and determine the author of each academic paper in the paper dataset; Match the inventors of the patents with the authors of the papers, and determine the academic inventors based on the matching results; Determine the academic papers published by the academic inventor in the paper dataset, and determine the patents published by the academic inventor in the patent dataset; The similarity between academic papers and patents published by the academic inventor is determined, and based on the similarity, a paper-patent pair dataset is determined.

4. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to any one of claims 1 to 3, characterized in that: The determining of the paper-patent pair dataset based on the patent dataset and the paper dataset further includes: For each paper-patent pair in the paper-patent pair dataset, determining the confidence of the paper-patent pair based on the similarity between the patent author and the paper inventor in the paper-patent pair, and the textual similarity between the patent and the paper; The confidence level is used to characterize the correlation between the paper and the patent in the paper-patent pair.

5. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to claim 4, characterized in that: The determining of the mapping relationship between the scientific classification labels and the technical classification labels in the paper-patent pair also includes: Filtering out paper-patent pairs in the paper-patent pair dataset whose confidence is less than or equal to a second threshold; and / or, Filtering out paper-patent pairs that are missing scientific classification labels and / or technical classification labels from the paper-patent pair dataset; and / or; Filter out paper-patent pairs lacking novelty in the paper-patent pair dataset; the difference between the publication time of the paper and the patent application time of the paper-patent pairs lacking novelty exceeds the grace period for not losing novelty.

6. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to claim 5, characterized in that: For each paper-patent pair, determining the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair includes: Determine a hierarchical structure of the scientific classification labels corresponding to the academic papers and the technical classification labels corresponding to the patents in each paper-patent pair; the hierarchical structure includes at least one level, and each sublevel has only one parent level; Based on the hierarchical structure, determining a mapping relationship between the scientific taxonomy label of the leaf category and the technical taxonomy label of the leaf category; Based on the hierarchical structure, the mapping relationship between the scientific taxonomy label of the leaf category and the technical taxonomy label of the leaf category is aggregated upward step by step to determine the mapping relationship between the scientific taxonomy label of each level and the technical taxonomy label of the corresponding level; Based on the mapping relationship between the scientific classification labels of each level and the technical classification labels of the corresponding level, the mapping relationship between the scientific classification labels corresponding to the academic papers and the technical classification labels corresponding to the patents in the paper-patent pair is determined.

7. The method for constructing a taxonomy mapping dataset based on paper-patent pairs according to claim 5, characterized in that: The step of determining, for each paper-patent pair, a mapping relationship between a scientific classification label corresponding to the academic paper and a technical classification label corresponding to the patent in the paper-patent pair further includes: Determine a hierarchical structure of scientific classification labels corresponding to academic papers and technical classification labels corresponding to patents in each paper-patent pair; the hierarchical structure includes at least one level, and a sub-level corresponds to multiple parent levels; Based on the hierarchical structure, determining scientific classification labels of each level corresponding to the academic paper and technical classification labels of each level corresponding to the patent in the paper-patent pair; Perform full-connect mapping on the scientific taxonomy labels of each level and the technical taxonomy labels of the corresponding level in turn; obtain the mapping relationship between the scientific taxonomy labels of each level and the technical taxonomy labels of the corresponding level; Based on the mapping relationship between the scientific classification labels of each level and the technical classification labels of the corresponding level, the mapping relationship between the scientific classification labels corresponding to the academic papers and the technical classification labels corresponding to the patents in the paper-patent pair is determined.

8. A device for constructing a scientific and technological classification mapping dataset based on paper-patent pairs, characterized in that: The device comprises: A first processing module is used to determine a paper dataset and a patent dataset in a target field; the paper dataset includes multiple academic papers, and the patent dataset includes multiple patents; A second processing module is configured to determine a paper-patent pair dataset based on the patent dataset and the paper dataset; the paper-patent pair dataset includes a plurality of paper-patent pairs, each paper-patent pair includes an academic paper and a patent; The third processing module is used to supplement the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in each paper-patent pair based on the scientific classification and technical classification; The fourth processing module is used to determine, for each paper-patent pair, the mapping relationship between the scientific classification label corresponding to the academic paper and the technical classification label corresponding to the patent in the paper-patent pair; based on the mapping relationship associated with all paper-patent pairs, obtain a classification mapping data set; the classification mapping data set is determined based on the scientific classification and the technical classification.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Portable underwater mass spectrometer

    US6727498B2

  • Hormonally up-regulated, neu-tumor-associated kinase

    US7368113B2

Cited By

  • Patent analysis program, patent analysis system, and patent analysis method

    JP7916055B1