Talent pool construction method and system based on patent application

By analyzing and classifying patent data, identifying and marking inventor information, calculating influence indexes and building a cooperative relationship chart, the problem of traditional talent pool construction methods lacking flexibility and depth in the rapidly changing technology field is solved, and the precise identification and utilization of key talents is achieved, real-time updates and efficient data management are supported, and the accuracy of talent decision-making and technical trend prediction in scientific and technological innovation is improved.

CN120011374AInactive Publication Date: 2025-05-16GUANGDONG ZHIDELI NETWORK TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411908265.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional method of building a talent pool based on patent applications lacks flexibility and depth in the rapidly changing technical fields, and cannot fully utilize multi-source data for comprehensive analysis, resulting in timeliness and accuracy of talent pool information, difficulty in real-time update, and in-depth mining and management of inventor information, resulting in insufficient identification and utilization of key talents, and the inability to effectively support talent decision-making and prediction of technological development trends in scientific and technological innovation.

Method used

By analyzing text and image information in patent data, identifying basic information of patents and classifying them, generating patent information extraction results; identifying inventor information and marking, calculating inventor influence index and technical contribution scores, evaluating the activity of the technical field and matching the search tags, building a cooperative relationship diagram and talent database among inventors, and monitoring and updating patent and talent database information in real time.

Benefits of technology

It enhances the accuracy and consistency of inventor identity data, improves the ability to identify key talents, monitors patent application status in real time and associates with talent databases, optimizes the timeliness of data, and supports more accurate talent decision-making and technical trend prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011374A_ABST
    Figure CN120011374A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data warehouses, in particular to a patent application-based talent pool construction method and system, and the method comprises the following steps: analyzing text and image information in patent data based on a patent information database, recognizing multiple basic information of patents, and carrying out the classification according to application regions, technical fields and application enterprises, and generating a patent information extraction result. According to the invention, by analyzing text and image information in patent data, multi-dimensional classification of patent information and inventor identity marking based on credit codes and organization codes are realized, the accuracy and consistency of the inventor identity data are enhanced, and by calculating the influence index and the technical contribution score of the inventor, the accuracy and the consistency of the inventor identity data are improved. The patent application state is monitored in real time and associated with a talent database through a retrieval label matching system formed by combining the influence score of the inventor with the information of the technical field and the application unit, and the timeliness of data is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data warehouse, and in particular to a method and system for constructing a talent pool based on a patent application. Background Art

[0002] The field of data warehouse technology focuses on the design, implementation and maintenance of large-scale storage systems for data analysis, aiming to achieve effective storage and management of batch data, optimize the performance of data retrieval and analysis, process complex data queries and support in-depth data mining activities. By using data integration, data is extracted, converted and loaded from multiple sources, and stored in a consistent format in the warehouse. Through data cleaning, the accuracy and consistency of the data are ensured, and data consistency analysis is used to ensure that the information in the warehouse accurately corresponds to the actual situation. Combined with the management of time dimension data, it tracks and records data changes over time, supports in-depth analysis of historical data and prediction of future trends, and is applied to business intelligence, market analysis, customer relationship management and decision support.

[0003] Among them, the talent pool construction method focuses on building a specialized talent pool for patent inventors. By systematically collecting and organizing inventor data in patent applications, key talents are identified and their influence and innovation capabilities in the technology field are evaluated. By integrating and analyzing patent data, the inventors' cooperation networks, technical expertise and innovation processes are revealed, providing valuable insights into technology trends. This is of great significance for technology companies, research institutions, and government departments to formulate relevant talent strategies and technology policies, helping to optimize the allocation of talent resources and promote strategic decision-making on technological innovation and industry development.

[0004] The traditional talent pool construction method based on patent applications lacks sufficient flexibility and depth when facing the rapidly changing technology field and complex talent assessment needs. In terms of technical talent assessment and talent pool construction, it cannot fully utilize multi-source data for comprehensive analysis, resulting in timeliness and accuracy issues in the information in the talent pool. When processing batches of dynamically changing patent data, it is difficult to update in real time, which limits the application efficiency of the talent pool in a rapidly changing market environment. It cannot deeply mine and systematically manage the inventor information in the patent data, resulting in inaccurate identification and utilization of key talents, and cannot effectively support enterprises and research institutions in making accurate talent decisions in scientific and technological innovation. The prediction of technology development trends is not accurate enough, making it difficult to provide effective support for science and technology policy formulation and talent strategy planning. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings in the prior art and to propose a talent pool construction method and system based on a patent application.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solution, a method for building a talent pool based on a patent application, comprising the following steps: S1: Based on the patent information database, analyze the text and image information in the patent data, identify multiple basic information of the patent, and classify it according to the application area, technical field, and applicant enterprise to generate patent information extraction results; S2: Based on the patent information extraction results, identify the names and affiliated institutions of the inventors of multiple patents, mark the inventors according to the type, unified credit code and organization code of the applicant, and generate inventor profile information; S3: Based on the inventor profile information, identify the number and citation frequency of multiple inventors' patents, calculate the influence index and technical contribution score of multiple inventors, and generate the inventor influence score; S4: using the inventor influence score, evaluating the technical fields in which multiple inventors are active, combining the inventor type and the applicant unit information, matching search tags for multiple inventors, and generating search tag matching records; S5: using the search tag matching records, constructing a cooperation relationship diagram and a talent database among inventors based on the inventors and patent citation information of multiple patents, and generating a relationship network analysis result; S6: Based on the relationship network analysis results, the patent application, approval progress and legal status update information in the patent database are monitored and analyzed in real time, and associated with the talent database to generate talent database update results.

[0007] As a further solution of the present invention, the patent information extraction result includes inventor identification data, patent technology field identification, and patent application information; the inventor profile information includes the inventor's professional background information, the applicant's unified credit code and organization code, the inventor's work unit and contact information; the inventor influence score includes the total number of the inventor's patents, the number of citations of the inventor's patents, and the activity analysis results of the technology field; the search tag matching record includes the inventor's patent technology classification label, the inventor's cooperation network label, and the inventor's application unit association label; the relationship network analysis result includes key node identification results, cooperation relationship diagram, and technology exchange path analysis results; the talent database update result includes patent status change information, legal status update information, and inventor information update record.

[0008] As a further solution of the present invention, based on the patent information database, the text and image information in the patent data is analyzed to identify multiple basic information of the patent, and the patent is classified according to the application area, technical field, and applicant enterprise. The specific steps of generating patent information extraction results are as follows: S101: Based on the patent information database, identify the title, abstract and graphic information according to the text and image information in multiple patents, record the inventor's name and application date, and generate basic data collection records; S102: Based on the basic data collection records, classify the patents according to the application regions and technical fields to generate classified patent data; S103: Based on the classified patent data, the patents are classified according to the applicant enterprise information, and patent lists of multiple enterprises and institutions are constructed, a patent library of the applicant entity is established, and patent information extraction results are generated.

[0009] As a further solution of the present invention, based on the patent information extraction results, the names and affiliated institutions of the inventors of multiple patents are identified, and the inventors are marked according to the type, unified credit code and organization code of the applicant. The steps of generating the inventor profile information are as follows: S201: Based on the patent information extraction result, according to the inventors and affiliated institutions information of multiple patents, identify the names and affiliated institutions of the inventors, and generate affiliated institution type information; S202: Based on the information on the type of the affiliated institution, the inventors are classified and recorded according to the type of the applicant, including enterprises, universities, and research institutes, and inventor classification data is generated; S203: Based on the inventor classification data, combined with the applicant's unified credit code and organization code, the inventors are marked, and identity identification codes are created for multiple inventors to generate inventor profile information.

[0010] As a further solution of the present invention, based on the inventor profile information, the number and citation frequency of multiple inventors' patents are identified, and the influence index and technical contribution score of multiple inventors are calculated. The steps of generating inventor influence scores are specifically as follows: S301: Based on the inventor profile information, identify and analyze the number of patents and the number of patent citations of multiple inventors to generate data statistics results; S302: Using the data statistical results, classify multiple patents of the inventor according to the technical field information, evaluate the influence of the inventor in multiple technical fields, and generate influence analysis results; S303: Based on the influence analysis results, according to the influence and number of patents in multiple technical fields, the overall influence index and technical contribution score of multiple inventors are calculated to generate the inventor influence score.

[0011] As a further solution of the present invention, the inventor influence score is used to evaluate the technical fields in which multiple inventors are active, and the search tags are matched for multiple inventors in combination with the inventor type and the applicant unit information. The steps of generating the search tag matching record are specifically as follows: S401: Based on the inventor influence score, by analyzing the technical field information of the patent, calculating the inventor's activity in multiple technical fields, and generating activity evaluation information; S402: According to the activity evaluation information, according to the patent application unit and application location, combined with the type of the inventor's affiliated unit, including enterprise, university, and research institute, matching search tags for multiple inventors, and generating tag configuration data; S403: Based on the tag configuration data, the tag data is used to configure and update the retrieval system of the talent database, including updating database query statements, optimizing the efficiency and accuracy of data retrieval, and generating retrieval tag matching records.

[0012] As a further solution of the present invention, the specific formula for calculating the inventor's activity in multiple technical fields is: ; in, On behalf of the inventor Activity scores in each technology area, Indexes for specified technology areas. is the total number of patents considered, is the index of a single patent, For the The weight corresponding to each patent, is the base of natural logarithms, is the decay rate, which determines the speed at which the weight of past patents decreases. The larger the value, the greater the impact of time and the more important the contribution of recent data. is the current year, used as a reference point for the assessment, For the The application year of each patent is used to calculate the length of time from the patent to the present.

[0013] As a further solution of the present invention, the search tag matching records are used to construct a cooperation relationship diagram and a talent database among inventors based on the inventors and patent citation information of multiple patents, and the steps of generating the relationship network analysis results are specifically as follows: S501: Based on the search tag matching record, by analyzing the applicant and inventor information of multiple patents, identifying the number of common patents and technical fields of multiple inventors, and generating a cooperation relationship data set; S502: Using the cooperation relationship dataset, defining nodes as inventors, defining edges as the number of common patents and technical fields, and weighting nodes and edges according to the number of common patents and the importance of the fields, to construct an inventor cooperation relationship graph; S503: Analyze the inventor cooperation relationship diagram, analyze and identify central nodes and intensive cooperation areas, identify key cooperation groups, and build a talent database based on the inventor information to generate relationship network analysis results.

[0014] As a further solution of the present invention, based on the relationship network analysis results, the patent application, approval progress and legal status update information in the patent database are monitored and analyzed in real time, and associated with the talent database. The steps of generating the talent database update results are specifically as follows: S601: Based on the relationship network analysis results, real-time monitoring of patent application and approval progress information in the patent database, recording patent status change events, and generating patent status monitoring records; S602: Based on the patent status monitoring record, identify the patent information in multiple status update events, and update the patent status information of the inventor by associating it with the inventor data in the talent database to generate status update data; S603: Based on the status update data, the talent database is updated and timestamp information is recorded, and in combination with regular integrity and accuracy verification of the data in the talent database, a talent database update result is generated.

[0015] A talent pool construction system based on patent applications, the talent pool construction system based on patent applications is used to execute the talent pool construction method based on patent applications, the system comprises: The data extraction module is based on the patent information database. By analyzing the text and image information in multiple patents, the module identifies the patent title, abstract, technical field and chart information, and classifies the patents according to the application region, technical field and application enterprise, and generates patent information summary data. The inventor identification module identifies the names and affiliated institutions of the inventors of multiple patents based on the patent information summary data, marks the inventors in combination with the type, unified credit code, and organization code of the applicant, and generates inventor marking information; The influence evaluation module calculates the influence index and technical contribution score of the inventors based on the inventor marking information and generates an influence analysis result by analyzing the number of patents and citation frequencies of multiple inventors; The tag matching module analyzes the technical fields in which multiple inventors are active based on the influence analysis results, matches search tags for inventors based on the inventor types and applicant unit information, and generates a tag information data set; The data update module constructs a cooperation relationship diagram among multiple inventors based on the label information data set, monitors the update information of the patent database in real time, associates the information with the talent database, and generates the talent database update result.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by analyzing the text and image information in the patent data, multi-dimensional classification of patent information and inventor identity marking based on credit codes and organizational codes are realized, the accuracy and consistency of the inventor identity data are enhanced, and the identification ability of key talents is enhanced by calculating the inventor's influence index and technical contribution score. By combining the inventor's influence score with the technical field and the applicant unit information, a retrieval tag matching system is formed, which monitors the patent application status in real time and associates it with the talent database, thereby optimizing the timeliness of the data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the workflow of the present invention; Figure 2 This is a detailed flow chart of S1 of the present invention; Figure 3 This is a detailed flow chart of S2 of the present invention; Figure 4 This is a detailed flow chart of S3 of the present invention; Figure 5 This is a detailed flow chart of S4 of the present invention; Figure 6 This is a detailed flow chart of S5 of the present invention; Figure 7 This is a detailed flow chart of S6 of the present invention; Figure 8 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0020] See also Figure 1The present invention provides a technical solution, a method for building a talent pool based on a patent application, comprising the following steps: S1: Based on the patent information database, analyze the text and image information in the patent data, identify multiple basic information of the patent, and classify it according to the application area, technical field, and applicant enterprise to generate patent information extraction results; S2: Based on the patent information extraction results, identify the names and affiliated institutions of the inventors of multiple patents, mark the inventors according to the type of applicant, unified credit code and organization code, and generate inventor profile information; S3: Based on the inventor profile information, identify the number and citation frequency of multiple inventors’ patents, calculate the influence index and technical contribution score of multiple inventors, and generate the inventor influence score; S4: Use the inventor influence score to evaluate the technical fields in which multiple inventors are active, combine the inventor type and applicant unit information, match search tags for multiple inventors, and generate search tag matching records; S5: Use search tags to match records, build a cooperation relationship diagram and talent database among inventors based on the inventors and patent citation information of multiple patents, and generate relationship network analysis results; S6: Based on the relationship network analysis results, real-time monitoring and analysis of patent applications, approval progress and legal status update information in the patent database are performed, and the information is associated with the talent database to generate talent database update results.

[0021] The results of patent information extraction include inventor identification data, patent technology field identification, and patent application information. The inventor's profile information includes the inventor's professional background information, the applicant's unified credit code and organizational code, the inventor's work unit and contact information. The inventor's influence score includes the total number of the inventor's patents, the number of citations of the inventor's patents, and the results of the activity analysis in the technology field. The search tag matching records include the inventor's patent technology classification label, the inventor's cooperation network label, and the inventor's application unit association label. The relationship network analysis results include key node identification results, cooperation relationship diagrams, and technology exchange path analysis results. The talent database update results include patent status change information, legal status update information, and inventor information update records.

[0022] See also Figure 2 Based on the patent information database, the text and image information in the patent data are analyzed to identify multiple basic information of the patent, and the patent is classified according to the application area, technical field, and application enterprise. The specific steps for generating patent information extraction results are as follows: S101: Based on the patent information database, identify the title, abstract and graphic information according to the text and image information in multiple patents, record the inventor's name and application date, and generate basic data collection records; In sub-step S101, the patent information database is used to apply optical character recognition technology and use the Tesseract OCR engine to automatically identify text data, including titles, abstracts, and graphic information, to ensure accurate text extraction from high-resolution images. The system identifies and records the inventor's name and application date through regular expression matching and date parsing algorithms. The date format is uniformly adjusted to ISO standards, and the information is saved in JSON format to facilitate subsequent data calls and analysis. The generated basic data collection records are stored in the form of a two-dimensional data table, including patent titles, abstract content, key elements of graphics, inventors, and application dates. The data set provides a data basis for subsequent classification and analysis.

[0023] S102: Based on the basic data collection records, classify the patents according to the application regions and technical fields to generate classified patent data; In sub-step S102, after the collected basic data is processed, the application area is located through the geographic information identification system (GIS). The built-in regional identification module of the system can identify geographical identifiers around the world, and the support vector machine in the machine learning algorithm is used to automatically classify the patent technology field. During the classification process, natural language processing technology is used to extract key terms and concepts from the patent text, and classification labels are generated based on the frequency and relevance of the terms. The classified data includes patent number, application area, and technology field label. The data is saved in a standardized CSV format to facilitate subsequent data processing and machine learning training, ensuring the integrity and availability of the data. The classified patent data provides a basis for the construction of the enterprise patent library in the next stage.

[0024] S103: Based on the classified patent data and the applicant enterprise information, the patents are classified, and patent lists of multiple enterprises and institutions are constructed, a patent library of the applicant entity is established, and patent information extraction results are generated; In sub-step S103, based on the classified patent data, the K-means clustering algorithm is used to classify the patents. According to the applicant enterprise information, a patent list for each enterprise is constructed. In the process, the enterprise names are fuzzy matched and synonyms are replaced to ensure data consistency. The association rule mining technology Apriori algorithm is used to analyze the patent application patterns among different enterprises and construct an enterprise patent library. The data in the library includes enterprise identification, number of patents, and patent type information. The data is stored in the form of a relational database. The generated patent information extraction results provide data support for the enterprise's R&D and market strategies, and provide an accurate data source for subsequent industry analysis and enterprise positioning.

[0025] See also Figure 3 Based on the patent information extraction results, the names and affiliated institutions of the inventors of multiple patents are identified, and the inventors are marked according to the applicant's type, unified credit code and organization code. The specific steps for generating inventor profile information are as follows: S201: Based on the patent information extraction results, according to the inventors and affiliated institutions information of multiple patents, identify the names and affiliated institutions of the inventors, and generate affiliated institution type information; In sub-step S201, text analysis technology is used to extract the names of inventors and information on affiliated institutions from patent documents using the NER module of the entity recognition model SpaCy, ensuring accurate extraction of key information from complex text data. Through the institutional classification algorithm, the system distinguishes different types of affiliated institutions, including enterprises, universities or research institutes. The classification is based on keywords in the institution name and classification information in the existing database. Each affiliated institution information is then labeled with a type and stored in a structured database. The information includes the institution name and type identifier, and the affiliated institution type information is generated. The information serves as the basis for identifying and classifying inventors, providing a foundation for subsequent data processing and analysis.

[0026] S202: Based on the information of the type of affiliated institution, the inventors are classified and recorded according to the type of applicant, including enterprises, universities, and research institutes, and inventor classification data is generated; In sub-step S202, a classification processing method is adopted to classify inventors according to the previously determined information on the type of affiliated institutions. A decision tree classification algorithm is used, and the DecisionTreeClassifier in Python's Scikit-Learn library is used. Based on the data attributes of the institutional type, including enterprises, universities, and research institutes, inventors are classified into corresponding categories. The record of each inventor includes the name, affiliated category, and associated patent information. The classified data is saved in an SQL database, including the inventor ID, name, classification label, and the number of associated patents, to generate inventor classification data. The data provides an accurate classification basis for further statistical analysis and creation of inventor archives.

[0027] S203: Based on the inventor classification data, combined with the applicant's unified credit code and organization code, the inventor is marked, and identity identification codes are created for multiple inventors to generate inventor profile information; In sub-step S203, in combination with the applicant's unified credit code and organizational code, the system verifies and associates the identity information of each inventor through a code matching algorithm and hash matching technology, creates a unique identification code for each inventor, and uses the identification code generation algorithm UUID to ensure that the information of each inventor is unique and traceable. The information is aggregated and stored in a centralized information management system, including the inventor's name, affiliated organization, identification code, associated patent ID, and generated inventor profile information.

[0028] See also Figure 4 Based on the inventor profile information, the number and citation frequency of multiple inventors’ patents are identified, and the influence index and technical contribution score of multiple inventors are calculated. The specific steps for generating inventor influence scores are as follows: S301: Based on the inventor profile information, identify and analyze the number of patents and the number of patent citations of multiple inventors, and generate data statistics results; In sub-step S301, the SQL query language is used to retrieve the inventor's archival information from the central database, including the patents and citation times of each inventor. The data is summarized through SQL aggregation functions to calculate the total number of patents and the total number of citations of each inventor. The Pandas library is used to process large-scale data sets during data processing to ensure the efficiency and accuracy of data processing. After the statistics are completed, the data is converted into visual information, and the Matplotlib library is used to generate bar charts and scatter plots to intuitively display the number of patents and citations of each inventor. The generated data statistics show the individual's R&D activities and reflect the market and academic influence of their patents.

[0029] S302: Using the statistical results, classify multiple patents of the inventor according to the technical field information, evaluate the influence of the inventor in multiple technical fields, and generate influence analysis results; In sub-step S302, based on the statistical results, natural language processing technology is used, and the TF-IDF algorithm is adopted to extract keywords from the description of each patent, and the patents are classified into technical fields. After the classification is completed, network analysis tools, including NetworkX, are used to build a patent network for each inventor. Nodes represent patents, and edges represent citation relationships between patents. According to the classification and citation of patents, the influence of each inventor in each technical field is evaluated, and the influence score is calculated. The score is weighted based on the number of patent citations and the importance of the technical field to which it belongs. The generated influence analysis results are displayed in the form of charts, providing the relative influence and market contribution of the inventor in different technical fields.

[0030] S303: Based on the results of the influence analysis, the overall influence index and technical contribution score of multiple inventors are calculated according to the influence and number of patents in multiple technical fields to generate the inventor influence score; In the above content, based on the results of influence analysis, the inventor's overall influence index and technical contribution score are calculated according to the formula ; In the formula, Represents the overall influence index, Representative The number of patents in the technical field, Representative Number of patent citations in the technical field, and are the weight coefficients of the number of patents and the number of citations, respectively; Detailed explanation of the formula and the process of formula calculation and derivation: Assume that the number of patents in each technology field They are 5, 10, and 8 respectively, and the corresponding number of citations The weight coefficient is 15, 5, 20. , ,calculate : ; The result 33.2 shows that the target inventor has a high influence in the technical field involved. The calculation process is used to evaluate the inventor's market and academic contributions to obtain the inventor's influence score.

[0031] See also Figure 5 , using the inventor influence score, evaluate the technical fields where multiple inventors are active, combine the inventor type and applicant unit information, match search tags for multiple inventors, and generate search tag matching records in the following steps: S401: Based on the inventor influence score, by analyzing the technical field information of the patent, the inventor's activity in multiple technical fields is calculated to generate activity evaluation information; The specific formula for calculating the inventor’s activity in multiple technical fields is: ; in, On behalf of the inventor Activity scores in each technology area, Indexes for specified technology areas. is the total number of patents considered, is the index of a single patent, For the The weight corresponding to each patent, is the base of natural logarithms, is the decay rate, which determines the speed at which the weight of past patents decreases. The larger the value, the greater the impact of time and the more important the contribution of recent data. is the current year, used as a reference point for the assessment, For the The application year of each patent is used to calculate the length of time from the patent to the present.

[0032] formula: ; Detailed explanation of the formula and the process of formula calculation and derivation: The formula is used to calculate the inventor's activity score in the target technology field. The score is based on the weight associated with the number of each patent and its application time, providing a comprehensive impact of the inventor's current and historical contributions to the field.

[0033] Parameter meaning and setting value: For inventors in the field of technology Activity score of For the Patents in the field , assuming that each patent in the pharmaceutical field by the inventor contributes 1 to activity; is the time decay coefficient, which is set to 0.1. The value reflects the impact of time on the importance of the patent, indicating that the importance of the target patent decays to 90% of the original value every year; is the current year, assuming it is 2024; For the The application year of the patents, For the total number of patents considered, assume that the inventor has 5 patents in the pharmaceutical field, all applied for in 2020.

[0034] is the base of the natural logarithm, approximately 2.71828.

[0035] Substitute the parameters into the formula for calculation: Assume that all five patents in the pharmaceutical field were applied for in 2020: ; result It shows that the comprehensive activity score of the inventor in the pharmaceutical field is 3.3515. The score reflects the inventor's activity in the target technology field. A higher score indicates that the inventor has significant activity in the pharmaceutical field in the past four years. This calculation process quantifies the relative activity of each inventor in their respective technology fields, providing a standardized evaluation tool for the talent database.

[0036] S402: According to the activity evaluation information, the patent application unit and application location, and the type of the inventor's affiliated unit, including enterprises, universities, and research institutes, multiple inventors are matched with search tags to generate tag configuration data; In sub-step S402, the data label configuration is refined according to the generated activity evaluation information. The classification algorithm is used to assign labels to the inventors by combining the patent application unit and application location information of each inventor. The inventors are grouped using the K-means clustering algorithm. The clustering basis includes the inventor's activity score, unit type and geographic location information. The clustering analysis is performed using Python's Sklearn library. The generated label configuration data includes the name of each inventor, the clustered label, the unit type and its geographic location. The information is stored in a SQL database to support more accurate talent search and resource matching.

[0037] S403: Based on the tag configuration data, the tag data is used to update the configuration of the search system of the talent database, including updating the database query statement, optimizing the efficiency and accuracy of data search, and generating search tag matching records; In sub-step S403, based on the generated label configuration data, the retrieval system of the talent database is updated, and the SQL language is used to optimize the database query statements, including writing the WHERE clause to improve the accuracy and efficiency of the retrieval, and reducing the query response time and improving the data retrieval performance by adjusting the index and query optimizer parameter settings. The system tests various query statements to ensure that the newly configured labels can be correctly applied to the query. The optimized retrieval system can quickly and accurately match and retrieve the inventor information. The generated retrieval label matching records include the execution time of each query, the accuracy evaluation of the query results and the system resource usage, providing reference data for subsequent system maintenance and upgrades.

[0038] See also Figure 6 , using search tags to match records, and based on the inventors and patent citation information of multiple patents, construct a cooperation relationship diagram and talent database among inventors. The specific steps to generate the relationship network analysis results are: S501: Based on the search tag matching records, by analyzing the applicant and inventor information of multiple patents, identifying the number of common patents and technical fields of multiple inventors, and generating a cooperation relationship dataset; In sub-step S501, the system first uses SQL query to extract the applicant and inventor information of all patents from the database, uses Python's Pandas library to process the target data, merges the patents involved by each inventor, identifies pairs of inventors with common patents, and calculates the number of common patents between pairs. For each pair of inventors, the technical field of the common patents is analyzed, and the target information is combined to generate a cooperation relationship dataset including the inventor name, the number of common patents, and the technical field. The dataset provides a detailed view of the cooperation between inventors and provides a data basis for network analysis. The generated cooperation relationship dataset is saved in a table form, which provides basic data for constructing a network diagram and subsequent analysis.

[0039] S502: using the cooperation relationship dataset, defining nodes as inventors, edges as the number of common patents and technical fields, and weighting nodes and edges according to the number of common patents and the importance of the fields, to construct an inventor cooperation relationship graph; In the above content, when constructing the inventor cooperation relationship graph, the weight of the node is calculated by the formula Perform calculations; In the formula, represents the weight of the node, represents the number of joint patents, Represents the importance score of the technical field covered by the common patent, and are weight coefficients, which adjust the impact of the number of common patents and the importance of the technological field respectively; Detailed explanation of the formula and the process of formula calculation and derivation: Considering that the number of joint patents and the importance of technical fields have a significant impact on the structure of the inventor collaboration graph, it is assumed that the weight coefficient , The inventor represented by a certain node has 5 patent collaborations in total, and the importance score of the technical field to which the patent belongs is 85 points, according to the formula: ; Result 29 shows that this inventor is very important in the collaboration graph and has extensive collaboration in multiple important technological fields. The score reflects his central position and influence in the technology network. The calculation process is used to evaluate and visualize the depth and breadth of technological collaboration, providing a quantitative basis for subsequent talent database construction and relationship network analysis.

[0040] S503: Analyze the inventor cooperation relationship diagram, analyze and identify the central nodes and intensive cooperation areas, identify key cooperation groups, and build a talent database based on the inventor information to generate relationship network analysis results; In sub-step S503, the constructed inventor cooperation relationship graph is analyzed. By focusing on the central nodes and intensive cooperation areas in the graph, the centrality analysis methods in graph theory, including degree centrality and betweenness centrality, are used to identify inventors who occupy a core position in the network, and a talent database is constructed, in which the detailed information of each inventor and his position in the network are recorded. The generated relationship network analysis results provide a scientific basis for human resource management and team building, and provide enterprises or research institutions with strategies for optimizing resource allocation and enhancing innovation capabilities.

[0041] See also Figure 7 Based on the relationship network analysis results, the patent application, approval progress and legal status update information in the patent database are monitored and analyzed in real time, and associated with the talent database. The specific steps for generating talent database update results are as follows: S601: Based on the relationship network analysis results, real-time monitoring of patent application and approval progress information in the patent database, recording patent status change events, and generating patent status monitoring records; In sub-step S601, the system monitors patent application and approval progress information in real time by setting triggers and event listeners of the patent database, uses SQL triggers to automatically detect changes in patent status, and records change events. The trigger is designed to automatically execute record operations when the patent status field is updated to capture changes, such as the conversion of patents from application to approval or rejection. The recorded information includes patent ID, status before and after the change, and change time. The generated patent status monitoring records are saved in the form of a time series database, which optimizes the reading and writing efficiency of timestamp data and ensures the real-time and accuracy of the monitoring records.

[0042] S602: Based on the patent status monitoring record, identify the patent information in multiple status update events, and update the patent status information of the inventor by associating it with the inventor data in the talent database to generate status update data; In sub-step S602, the system uses Python's Pandas library to query and process the monitoring records based on the generated patent status monitoring records, identifies each status update event in the records, and associates it with the inventor data in the talent database through SQL queries to update the inventor's patent status information. In the process, the current patent list of each inventor is checked and the new status information is updated to the corresponding records. The data fields involved in the update operation include patent status and last update time. The generated status update data is saved in the master database to ensure that the inventor's patent records are consistent with the actual patent status, providing accurate data support for subsequent analysis and decision-making.

[0043] S603: Based on the status update data, the talent database is updated and the timestamp information is recorded, and the integrity and accuracy of the data in the talent database are checked regularly to generate a talent database update result; In sub-step S603, the system updates the talent pool database based on the status update data, including writing the latest status update data into the database and recording the relevant timestamp information, using MySQL's transaction processing mechanism to ensure the atomicity and consistency of data updates, and the system regularly performs data integrity and accuracy checks. It uses data verification algorithms, including Checksum and redundancy checks, to ensure that the information in the database has not been tampered with or damaged during transmission or storage. The generated talent database update results reflect the latest patent and inventor information, and also provide an effectiveness evaluation of the system's data processing and protection measures.

[0044] See also Figure 8 A talent pool construction system based on patent applications, the talent pool construction system based on patent applications is used to execute the talent pool construction method based on patent applications, the system includes: The data extraction module is based on the patent information database. By analyzing the text and image information in multiple patents, the module identifies the patent title, abstract, technical field and chart information, and classifies the patents according to the application region, technical field and application enterprise, and generates patent information summary data. The inventor identification module identifies the names and affiliated institutions of inventors of multiple patents based on the patent information summary data, marks the inventors in combination with the applicant's type, unified credit code, and organizational code, and generates inventor marking information; The impact assessment module is based on the inventor tag information. By analyzing the number of patents and citation frequencies of multiple inventors, the impact index and technical contribution score of the inventor are calculated to generate the impact analysis results. The label matching module analyzes the technical fields in which multiple inventors are active based on the results of influence analysis, combines the inventor type and applicant unit information, matches the search labels for the inventors, and generates a label information dataset; The data update module builds a cooperation relationship diagram among multiple inventors based on the label information data set, monitors the update information of the patent database in real time, associates the information with the talent database, and generates the talent database update results.

[0045] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A method for building a talent pool based on patent applications, characterized in that: The following steps are involved: Based on the patent information database, the text and image information in the patent data are analyzed to identify multiple basic information of the patent, and the patent is classified according to the application area, technical field, and applicant enterprise to generate patent information extraction results; Based on the patent information extraction results, identify the names and affiliated institutions of the inventors of multiple patents, mark the inventors according to the type, unified credit code and organization code of the applicant, and generate inventor profile information; Based on the inventor profile information, identify the number and citation frequency of multiple inventors' patents, calculate the influence index and technical contribution score of multiple inventors, and generate inventor influence scores; Using the inventor influence score, evaluate the technical fields in which multiple inventors are active, combine the inventor type and applicant unit information, match search tags for multiple inventors, and generate search tag matching records; Using the search tag matching records, a cooperation relationship diagram and a talent database among inventors are constructed based on the inventors and patent citation information of multiple patents, and a relationship network analysis result is generated; Based on the relationship network analysis results, the patent application, approval progress and legal status update information in the patent database are monitored and analyzed in real time, and associated with the talent database to generate talent database update results.

2. The method for building a talent pool based on patent applications according to claim 1, characterized in that: The patent information extraction results include inventor identification data, patent technology field identification, and patent application information; the inventor profile information includes the inventor's professional background information, the applicant's unified credit code and organization code, the inventor's work unit and contact information; the inventor influence score includes the total number of the inventor's patents, the number of citations of the inventor's patents, and the activity analysis results of the technology field; the search tag matching records include the inventor's patent technology classification label, the inventor's cooperation network label, and the inventor's application unit association label; the relationship network analysis results include key node identification results, cooperation relationship diagrams, and technology exchange path analysis results; the talent database update results include patent status change information, legal status update information, and inventor information update records.

3. The method for building a talent pool based on patent applications according to claim 1, characterized in that: Based on the patent information database, the text and image information in the patent data is analyzed to identify multiple basic information of the patent, and the patent is classified according to the application area, technical field, and applicant enterprise. The specific steps to generate the patent information extraction results are as follows: Based on the patent information database, identify the title, abstract and illustration information according to the text and image information in multiple patents, record the inventor's name and application date, and generate basic data collection records; Based on the basic data collection records, classify the patents according to the application regions and technical fields to generate classified patent data; Based on the classified patent data, the patents are classified according to the applicant enterprise information, and a patent list of multiple enterprises and institutions is constructed, a patent library of the applicant entity is established, and a patent information extraction result is generated.

4. The method for building a talent pool based on patent applications according to claim 1, characterized in that: Based on the patent information extraction results, the names and affiliated institutions of the inventors of multiple patents are identified, and the inventors are marked according to the type, unified credit code and organization code of the applicant. The specific steps of generating the inventor profile information are as follows: Based on the patent information extraction results, according to the inventors and affiliated institutions information of multiple patents, the names and affiliated institutions of the inventors are identified, and the affiliated institution type information is generated; Based on the information on the type of institution, inventors are classified and recorded according to the type of applicant, including enterprises, universities, and research institutes, to generate inventor classification data; Based on the inventor classification data, combined with the applicant's unified credit code and organizational code, the inventors are marked, and identity identification codes are created for multiple inventors to generate inventor profile information.

5. The method for building a talent pool based on patent applications according to claim 1, characterized in that: Based on the inventor profile information, the number and citation frequency of multiple inventors' patents are identified, and the influence index and technical contribution score of multiple inventors are calculated. The steps of generating the inventor influence score are specifically as follows: Based on the inventor profile information, identify and analyze the number of patents and the number of patent citations of multiple inventors to generate statistical results; Using the statistical results of the data, multiple patents of the inventor are classified according to the information of the technical fields, and the influence of the inventor in multiple technical fields is evaluated to generate influence analysis results; Based on the influence analysis results, the overall influence index and technical contribution score of multiple inventors are calculated according to the influence and number of patents in multiple technical fields to generate the inventor influence score.

6. The method for building a talent pool based on patent applications according to claim 1, characterized in that: The inventor influence score is used to evaluate the technical fields in which multiple inventors are active, and the search tags are matched for multiple inventors in combination with the inventor type and the applicant unit information. The specific steps of generating the search tag matching record are as follows: Based on the inventor influence score, by analyzing the technical field information of the patent, the inventor's activity in multiple technical fields is calculated to generate activity evaluation information; According to the activity evaluation information, based on the patent application unit and application location, combined with the type of the inventor's affiliated unit, including enterprises, universities, and research institutes, multiple inventors are matched with search tags to generate tag configuration data; Based on the tag configuration data, the tag data is used to configure and update the retrieval system of the talent database, including updating database query statements, optimizing the efficiency and accuracy of data retrieval, and generating retrieval tag matching records.

7. The method for building a talent pool based on patent applications according to claim 6 is characterized in that: The specific formula for calculating the inventor's activity in multiple technical fields is: ; in, On behalf of the inventor Activity scores in each technology area, Indexes for specified technology areas. is the total number of patents considered, is the index of a single patent, For the The weight corresponding to each patent, is the base of natural logarithms, is the decay rate, which determines the speed at which the weight of past patents decreases. The larger the value, the greater the impact of time and the more important the contribution of recent data. is the current year, used as a reference point for the assessment, For the The application year of each patent is used to calculate the length of time from the patent to the present.

8. The method for building a talent pool based on patent applications according to claim 1, characterized in that: The steps of using the search tag matching records to construct a cooperation relationship diagram and a talent database among inventors based on the inventors and patent citation information of multiple patents and generating the relationship network analysis results are as follows: Based on the search tag matching records, by analyzing the applicant and inventor information of multiple patents, identifying the number of common patents and technical fields of multiple inventors, and generating a cooperation relationship dataset; Using the cooperation relationship dataset, defining nodes as inventors, defining edges as the number of common patents and technical fields, and weighting nodes and edges according to the number of common patents and the importance of the fields, to construct an inventor cooperation relationship graph; Analyze the inventor cooperation relationship diagram, analyze and identify central nodes and intensive cooperation areas, identify key cooperation groups, build a talent database based on the inventor information, and generate relationship network analysis results.

9. The method for building a talent pool based on patent applications according to claim 1, characterized in that: Based on the relationship network analysis results, the patent application, approval progress and legal status update information in the patent database are monitored and analyzed in real time, and associated with the talent database. The steps of generating the talent database update results are as follows: Based on the relationship network analysis results, real-time monitoring of patent application and approval progress information in the patent database is performed, patent status change events are recorded, and patent status monitoring records are generated; Based on the patent status monitoring record, identifying the patent information in multiple status update events, updating the patent status information of the inventor by associating with the inventor data in the talent database, and generating status update data; Based on the status update data, the talent database is updated and the timestamp information is recorded, and in combination with regular integrity and accuracy verification of the data in the talent database, a talent database update result is generated.

10. A talent pool construction system based on patent applications, characterized in that: According to the method for building a talent pool based on a patent application according to any one of claims 1 to 9, the system comprises: The data extraction module is based on the patent information database. By analyzing the text and image information in multiple patents, the module identifies the patent title, abstract, technical field and chart information, and classifies the patents according to the application region, technical field and application enterprise, and generates patent information summary data. The inventor identification module identifies the names and affiliated institutions of the inventors of multiple patents based on the patent information summary data, marks the inventors in combination with the type, unified credit code, and organization code of the applicant, and generates inventor marking information; The influence evaluation module calculates the influence index and technical contribution score of the inventors based on the inventor marking information and generates an influence analysis result by analyzing the number of patents and citation frequencies of multiple inventors; The tag matching module analyzes the technical fields in which multiple inventors are active based on the influence analysis results, matches search tags for inventors based on the inventor types and applicant unit information, and generates a tag information data set; The data update module constructs a cooperation relationship diagram among multiple inventors based on the label information data set, monitors the update information of the patent database in real time, associates the information with the talent database, and generates the talent database update result.

Citation Information

Cited By

  • Pod label marking method, device, equipment, medium and product

    CN120804761A

  • Patent analysis program, patent analysis system, and patent analysis method

    JP7916055B1