Information processing method, electronic device, and computer-readable storage medium

Through dynamic academic data, the expert contribution and project requirements are evaluated, and the reconstructed resume is generated and combined with semantic correlation and target models, the coverage and accuracy of traditional talent discovery methods are solved, and efficient and accurate expert matching is achieved.

CN120256737BActive Publication Date: 2025-08-22HANGZHOU WEIMING XINKE TECH CO LTD +1

Patent Information

Application Number
CN202510696979.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-22
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Traditional talent discovery methods rely on expert recommendations and academic conferences, which have limited coverage, strong subjectivity, difficulty in quantitative evaluation, and inaccurate project requirements decomposition, resulting in high inaccuracy in expert matching.

Method used

An information processing method based on dynamic academic data is adopted to evaluate the expert contribution degree through the citation decay factor, an initial resume is generated and reconstructed, and a project requirement is processed by semantic association and target model to realize expert screening.

Benefits of technology

It improves the timeliness and accuracy of expert information mining, ensures that the recommendation results match project requirements, reduces subjective interference, and realizes quantitative evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256737B_ABST
    Figure CN120256737B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides an information processing method, an electronic device and a computer-readable storage medium. Relating to the field of information processing, the method includes: determining the field contribution information corresponding to a plurality of candidate experts based on dynamic academic data; using a first target model for processing based on the dynamic academic data to generate initial resumes corresponding to a plurality of candidate experts; screening and reconstructing the initial resumes according to a predetermined query template to obtain a reconstructed resume; performing semantic association based on the field contribution information corresponding to a plurality of candidate experts and the reconstructed resumes to obtain a talent portrait; using a second target model for processing based on the project requirement information of the target project to obtain a query corpus; based on the query corpus, screening the talent portraits corresponding to a plurality of candidate experts to determine the target experts required for the target project. The present application solves the technical problem in the related art that the matching degree of the mined expert information is not ideal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and in particular to an information processing method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Traditional talent discovery methods rely primarily on expert recommendations, academic conference discussions, and self-referrals. These methods suffer from limited coverage, strong subjectivity, and difficulty in quantitative evaluation. For example, expert recommendations can be subjective, leading to incomplete results. Similarly, academic conference discussions are limited by themes and participants, making it difficult to achieve large-scale, cross-disciplinary talent matching.

[0003] In recent years, with the development of big data technology, research on data-driven talent discovery has gradually gained momentum. Existing expert recommendation systems are mostly based on static academic data, such as fixed numbers of published papers and citations. This data fails to reflect experts' latest research findings and academic trends in a timely manner, potentially leading to a mismatch between recommended experts and actual project needs. Furthermore, in traditional project requirements processing, requirements decomposition often relies on manual labor, which can lead to inaccurate and incomplete decomposition. Due to an insufficient understanding of customer needs or a lack of a systematic decomposition method, project teams may experience deviations between the decomposed requirements and the actual project goals. The generated corpus may not fully cover all aspects of the project requirements, resulting in inaccurate or incomplete search results, further impacting the accuracy of expert matching.

[0004] To address the above problems, no solution has been proposed in the relevant technologies yet. Summary of the Invention

[0005] The embodiments of the present application provide an information processing method, an electronic device, and a computer-readable storage medium to alleviate or solve the technical problem in related technologies of unsatisfactory matching of mined expert information.

[0006] In a first aspect, an embodiment of the present application provides an information processing method, including:

[0007] Based on dynamic academic data, determine the field contribution information corresponding to multiple candidate experts. The dynamic academic data includes a citation decay factor, which is used to indicate that the number of citations of academic achievements decays over time. The field contribution information includes the contribution of the corresponding candidate experts in multiple fields.

[0008] Based on dynamic academic data, a first target model is used to process and generate initial resumes corresponding to multiple candidate experts. The first target model is trained based on a training set that includes public resumes and reference academic information. There is a corresponding relationship between the reference academic information and the public resumes.

[0009] According to the predetermined query template, the initial resume is screened and reconstructed to obtain the reconstructed resumes corresponding to multiple candidate experts;

[0010] Based on the semantic association of the field contribution information and reconstructed resumes of multiple candidate experts, the talent portraits corresponding to the multiple candidate experts are obtained;

[0011] Based on the project requirement information of the target project, the second target model is used to process and obtain the query corpus; the second target model is trained based on the historical requirement information and historical expert information of the reference project;

[0012] Based on the query corpus, the talent profiles corresponding to multiple candidate experts are screened, and the target experts required for the target project are determined from the multiple candidate experts.

[0013] In a second aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any method of the embodiment of the present application when executing the computer program.

[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any one of the embodiments of the present application is implemented.

[0015] Based on the information processing method of the first aspect mentioned above, this application has at least the following beneficial effects or advantages: using interpretable resume information as intermediate information, compared with the processing method of the black box model, each step of information processing and conversion is related to the resume information, from generating the initial resume based on dynamic academic data to reconstructing the initial resume to generate a semantically expressed talent portrait, making the entire expert screening process more transparent and understandable. And by combining dynamic academic data with the second target model, the efficiency of expert information mining is improved, and the latest research results of experts can be reflected in real time, overcoming the lag problem of traditional static data, and ensuring the timeliness and accuracy of the recommendation results. It can effectively solve the problem of unsatisfactory accuracy of traditional talent discovery methods and match highly adaptable experts to target projects.

[0016] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0018] Figure 1 A flowchart showing an information processing method according to an embodiment of the present application is shown;

[0019] Figure 2 A schematic diagram showing the distribution of topics in the information processing method according to an embodiment of the present application is shown;

[0020] Figure 3 A schematic diagram showing author contribution allocation of the information processing method according to an embodiment of the present application is shown;

[0021] Figure 4 A schematic block diagram showing an information processing method according to an embodiment of the present application is shown;

[0022] Figure 5 A block diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0023] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0024] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.

[0025] It should be noted that the above-mentioned application scenarios or application examples provided in the embodiments of this application are for ease of understanding, and the embodiments of this application do not specifically limit the application of the technical solution. In addition, the user information (including but not limited to user device information, user personal information, project requirement information, expert-related information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, dynamic academic data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation portals for users to choose to authorize or refuse.

[0026] The solution provided by related technologies divides literature into disciplinary areas, then extracts academic keywords from these disciplinary areas. Using researchers' academic output data, researchers can filter out keywords for academic research topics. Keyword similarity calculations are then used to determine the researchers' academic research topics. Faced with massive amounts of academic data, the processing methods provided by related technologies struggle to efficiently cope with this large volume of data, resulting in inefficient data mining. Furthermore, these technologies primarily rely on keyword extraction and similarity calculations, lacking the ability to explore deeper relationships within the data. This makes it difficult to comprehensively assess researchers' overall capabilities and potential, while also failing to promptly reflect their latest developments and capture their latest research findings and directions.

[0027] The following describes in detail the technical solution of this application and how it solves the aforementioned technical problems using specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following describes the embodiments of this application in detail with reference to the accompanying drawings.

[0028] Figure 1 A flow chart of the information processing method according to an embodiment of the present application is shown. Figure 1 As shown, the method may include steps S101 to S105.

[0029] Step S101: Based on dynamic academic data, determining the field contribution information corresponding to each of the multiple candidate experts. The dynamic academic data includes a citation decay factor, which is used to indicate that the number of citations of academic achievements decays over time. The field contribution information includes the contribution of the corresponding candidate experts in multiple fields.

[0030] Step S102: Based on the dynamic academic data, a first target model is used to process and generate initial resumes corresponding to multiple candidate experts. The first target model is trained based on a training set including public resumes and reference academic information. There is a corresponding relationship between the reference academic information and the public resumes.

[0031] Step S103: screening and reconstructing the initial resume according to a predetermined query template to obtain reconstructed resumes corresponding to multiple candidate experts;

[0032] Step S104: semantically associate the field contribution information and reconstructed resumes of the candidate experts to obtain talent profiles corresponding to the candidate experts;

[0033] Step S105: Based on the project requirement information of the target project, a second target model is used to process the query corpus; the second target model is trained based on historical requirement information and historical expert information of the reference project;

[0034] Step S106: Based on the query corpus, the talent profiles corresponding to the multiple candidate experts are screened, and the target expert required for the target project is determined from the multiple candidate experts.

[0035] The aforementioned dynamic academic data is the core data source for this embodiment. It includes not only traditional academic achievement data (such as the number of papers published and citations), but also specifically incorporates a citation decay factor. This factor represents the decay of citations over time, more accurately reflecting the timeliness and current influence of academic achievements. In this way, this embodiment can dynamically assess the contributions of candidate experts across multiple fields, avoiding the lag associated with relying solely on static data.

[0036] This second-goal model, trained based on historical requirements information and expert insights from reference projects, can more accurately decompose and understand project requirements. Processing project requirements information with the second-goal model generates high-quality query corpora. This not only improves the accuracy and comprehensiveness of requirements decomposition but also avoids the potential biases that can arise from manual decomposition.

[0037] The above-mentioned academic achievements include at least papers, patents, academic works (including software works), research reports, technical standards and specifications formulated with the participation of experts, etc.

[0038] In this embodiment, dynamic academic data is used to determine the domain contribution information of candidate experts, overcoming the limitations of related technologies that rely on static data. Dynamic academic data can promptly reflect the latest research results and academic trends of experts. By introducing a citation decay factor, it can more accurately assess the influence and contribution of experts over time, thereby improving the timeliness and accuracy of recommendation results. Using dynamic academic data, processing is performed with the help of a first target model to generate initial resumes for multiple candidate experts. The first target model is trained using a training set consisting of publicly available resumes and reference academic information (with a corresponding relationship between the two). The generated initial resumes comprehensively reflect the experts' academic background and actual capabilities, providing more comprehensive and accurate basic information in an unstructured format for subsequent screening. The initial resumes are screened and reconstructed to remove irrelevant information and highlight key content, making the reconstructed resumes more concise and targeted. Based on the reconstructed resumes, semantic associations are made with the domain contribution information to generate talent profiles that facilitate subsequent matching and screening with the query corpus. The second target model is used to process project requirement information to generate the query corpus. The second target model is trained based on historical requirements and expert information from reference projects, enabling a more accurate breakdown of project requirements and avoiding the potential inaccuracies and incompleteness of manual decomposition. By searching the reconstructed resumes of candidate experts using query data, it is possible to quickly and accurately identify the expert who best matches the target project requirements from a large, cross-disciplinary pool of candidate experts, reducing subjective interference and enabling a quantitative assessment of talent matching.

[0039] It's important to note that incorporating field contribution information into the reconstructed resume can add a quantitative dimension. While resumes are mostly qualitative and difficult to compare horizontally, field contribution information can provide a quantifiable contribution ranking, such as "contributions in the field of neural networks exceed 90% of peers." Furthermore, because the primary objective model is trained on publicly available resumes, the resulting reconstructed resumes tend to be biased towards positive descriptions. Therefore, field contribution information is needed to mitigate the effects of resume packaging. This information can be used to provide relevant information that doesn't mention failed projects or non-lead work, de-emphasizing non-core field experience. For example, if an expert's field contributions show that they primarily published papers in low-impact conferences, this association can weaken the positive semantics of these technical areas in the reconstructed resume. This semantic association can also be used through co-occurrence analysis (e.g., Expert A is cited in both the energy and materials fields) to identify cross-disciplinary innovation potential not explicitly mentioned in the resume, helping project teams identify more suitable, multidisciplinary talent.

[0040] Exemplarily, the query corpus can be obtained in a variety of ways, such as using the FASTopic model. The core of FASTopic is to extract the implicit semantic features of the requirement text through topic modeling. The project requirement information can be expressed in at least one of the following forms: text modality, voice modality, image modality, etc. One or more of the above modalities can be used to represent the requirements and functional descriptions. The multimodal fusion strategy can be to convert speech to text and merge it with the requirements in text modality, and then unify the modeling through FASTopic. Alternatively, OCR (Optical Character Recognition) can be used to extract text from the image, which is then spliced ​​with the text requirements and input into the model for processing.

[0041] Before inputting the model, project requirement information can be preprocessed, such as by cleaning the text, removing punctuation, stop words, and standardizing terminology. After tokenization, the preprocessed project requirement information can be constructed into a vocabulary or word frequency matrix. Example input: "Develop an intelligent customer service system that supports multilingual real-time translation, user sentiment recognition, and automatic ticket generation during conversations. The technical solution must balance accuracy and responsiveness." After preprocessing, the following results are obtained: ["intelligent customer service", "multilingual translation", "real-time", "emotion recognition", "automatic ticket generation", "accuracy", "responsiveness"]. The FASTopic model can output multiple topics and their weights. For example: 1. Topic Name: Multilingual NLP, Keywords: Multilingual, Translation, Real-time, Accuracy, Weight: 0.5; 2. Topic Name: Sentiment Analysis, Keywords: Emotion Recognition, User Feedback, Sentiment Classification, Weight: 0.3; 3. Topic Name: Low-Latency System, Keywords: Response Speed, Ticket Generation, Distributed, Weight: 0.2.

[0042] The reference project and the target project are of the same type, providing some reference value. Historical demand information indicates that the reference project previously matched experts with the research topic of multilingual NLP (Natural Language Processing). Therefore, the second target model increases the weight of Topic 1, which can be considered a priori knowledge. For example, to calculate the average match score (normalized to a range of 0 to 1) for a specific area in the reference project, the weight increase formula can be: base weight × (1 + historical score gain). This historical score gain is quantified based on historical matching topics indicated by historical expert information.

[0043] Based on the predetermined domain mapping relationship, Topics 1, 2, and 3 can be mapped to specific domains. For example, the Multilingual NLP topic corresponds to the following domains: "Machine Translation," "Speech Recognition," and "Cross-Language Modeling." The Sentiment Analysis topic corresponds to the following domains: "Affective Computing," "Psychology + AI," and "User Behavior Analysis." The Low-Latency System topic corresponds to the following domains: "Real-Time Computing," "Edge Computing," and "High-Concurrency Architecture." The query corpus can then be: "Searching for experts in at least two of the following domains: machine translation, affective computing, and real-time computing, with at least three years of productization experience." Based on the query corpus, prompts can be generated, including: speech recognition, speech recognition, psychology + AI, user behavior analysis, edge computing, and high-concurrency architecture.

[0044] For example, the above-mentioned dynamic academic data can be obtained in various ways. For example, the data collection module can connect to academic and patent databases, automatically collect journal literature and patent data through API interfaces and crawler technology, including paper titles, authors, keywords, abstracts, publication dates, publishing institutions, contact information, journals, citations, and other information, as well as patent titles, inventors, abstracts, claims, classification numbers, and other information. The collected data is stored in a distributed database to facilitate subsequent processing and analysis.

[0045] To improve the quality of dynamic academic data, the collected data is cleaned to remove duplicate, erroneous, and irrelevant information. Preprocessing operations such as word segmentation, denoising, and stemming are performed on the text data to improve data quality and the accuracy of subsequent analysis.

[0046] For example, academic achievements come in various forms. For example, using papers and patents as an example, we can integrate paper and patent data to establish a unified data model and standards. This involves standardizing terminology, talent data, and institutional information to achieve data disambiguation and normalization. For example, we can match and integrate author information, institutional information, contact information, email addresses, school background, and work experience from papers and patents to form a standardized expert and team data model, providing a foundation for subsequent analysis and mining.

[0047] The citation decay factor represents the decay of citations over time, more accurately reflecting the timeliness and influence of academic achievements. Over time, some earlier research may be gradually superseded by newer findings, and its citation counts will decrease accordingly. By introducing this decay factor, the actual impact of academic achievements over different time periods can be more rationally assessed. Compared to traditional static academic data (such as fixed numbers of published papers and citations), the citation decay factor dynamically reflects an expert's latest research findings and academic trends. It avoids the lag caused by relying solely on historical data and can more promptly capture the activity and influence of experts in the current field, thereby improving the timeliness and accuracy of expert recommendations.

[0048] It should be noted that citation behavior and citation half-lives vary across disciplines. The citation decay factor can be adjusted based on the characteristics of the discipline to better adapt to the citation patterns of different disciplines, thereby more accurately assessing the contributions of experts in their respective fields.

[0049] According to some embodiments provided in this application, in order to determine the citation attenuation factor, the method may further include the following specific steps:

[0050] Obtain the first citation count of academic achievements within a predetermined time range, and the second citation count of academic achievements outside the predetermined time range;

[0051] According to a predetermined attenuation ratio, the preset first weight is attenuated and adjusted in each predetermined attenuation period to obtain a second weight outside the predetermined time range;

[0052] A citation count attenuation factor is obtained by weighting based on the first weight, the second weight, the first citation count, and the second citation count.

[0053] In the embodiments provided herein, the first citation count represents the number of citations received by the academic achievement within a predetermined timeframe, reflecting the academic achievement's influence in the earlier period. The second citation count represents the number of citations received by the academic achievement outside the predetermined timeframe, reflecting the academic achievement's influence in the recent period. In this way, a comprehensive assessment of the academic achievement's citations over different timeframes can be achieved. By introducing a predetermined decay ratio and adjusting the preset first weight to decay according to a predetermined decay period, a second weight outside the predetermined timeframe is obtained. As the influence of an academic achievement gradually weakens over time, this temporal change is reflected by the decay weight. Specifically, the first weight is gradually reduced according to the decay ratio and decay period to obtain the second weight. A weighted calculation is performed based on the first weight, the second weight, the first citation count, and the second citation count to obtain a citation count decay factor. This factor comprehensively reflects the academic achievement's citations over different timeframes and, through the decay adjustment of the weight, more accurately reflects the timeliness and current influence of the academic achievement. In this way, an expert's academic contribution can be more scientifically assessed and applied to the talent discovery process to ensure that the recommended experts' capabilities are up-to-date and better match project needs.

[0054] For example, for a particular academic achievement, the weight for citations in the past five years is set to 1 (i.e., the first weight), indicating that the number of citations in the past five years contributes most to the weight, reflecting the recent impact of the paper. The decay rate is then set to 20%, and the decay cycle is set to annually. After five years, the weight of the number of citations each year (i.e., the second weight) decays by 20%. That is, the weight in the sixth year is 0.8, the seventh year is 0.6, and so on, reflecting the gradual weakening of the influence of academic achievements over time. The above specific numerical settings are for illustration only and are not limiting.

[0055] According to some embodiments provided by the present application, in step 101: determining the field contribution information corresponding to a plurality of candidate experts based on dynamic academic data may include the following specific steps:

[0056] Determine the subject to which each scholarly output included in the dynamic scholarly data belongs;

[0057] According to the domain mapping relationship and the subject of each academic achievement, the domain information corresponding to each academic achievement is determined. The domain mapping relationship represents the corresponding relationship between multiple fields and subjects.

[0058] Based on the field information corresponding to each academic achievement, the contribution statistics of the academic achievements of multiple candidate experts in different fields are collected to obtain the field contribution information corresponding to the multiple candidate experts.

[0059] In the embodiment provided in the present application, each academic achievement in the dynamic academic data is analyzed to determine the subject to which it belongs. According to the preset field mapping relationship, the subject to which each academic achievement belongs is mapped to the corresponding field information. The field mapping relationship defines the correspondence between multiple fields and subjects. Through this mapping, the specific field to which each academic achievement belongs can be clarified. Based on the field information corresponding to each academic achievement, the contribution statistics of the academic achievements of multiple candidate experts in different fields are performed. Through statistical analysis, the contribution of each candidate expert in various fields is obtained, and field contribution information is finally formed. Through the above-mentioned processing, the contribution of the same expert in different fields can be fully quantified, so that it can be widely used in cross-domain talent discovery scenarios and meet the needs of different projects for multi-domain experts.

[0060] For example, there are various ways to set the above-mentioned domain mapping relationships, such as statistically determined methods or methods based on a pre-set database. Alternatively, hierarchical clustering combined with the DBSCAN algorithm can be used to cluster topics and keywords to form multiple technology field clusters. By continuously merging or splitting data points, hierarchical clustering can effectively handle data distributions of varying sizes and shapes. DBSCAN is a density-based clustering algorithm that can identify clusters of arbitrary shapes and is robust to noise. By setting appropriate radius and minimum point count parameters, DBSCAN can group densely connected points into the same cluster while labeling isolated points as noise. An example is provided below. For example, a hierarchical clustering algorithm can be used to perform preliminary clustering of topics and keywords (or keywords) to form multiple preliminary technology field clusters. The DBSCAN algorithm is then used to further refine the clustering results. After clustering, the contextual understanding and reasoning capabilities of the large model are used to analyze the hierarchical relationships between keywords within each cluster. A hierarchical technology labeling system is constructed, from industry fields to research directions, and then to specific topics and keywords. By constructing the above-mentioned hierarchical relationship, technical labels are combined with the academic achievements of experts to form a complete production and technology map (that is, the field mapping relationship).

[0061] To facilitate understanding, we use some of the subject hierarchical structures in the field of intelligent Internet of Things as examples:

[0062] {

[0063] "Topic Number": 0,

[0064] "Topic": "Wireless Communication",

[0065] "Subject Keywords": "routing allocation, satellite, throughput, wireless data packets, terrestrial network, coverage, backhaul network, resource allocation, access network, communication, transmission delay, relay, cellular network, aerial node, satellite network, relay deployment, latency, radio connectivity",

[0066] "Research Direction": "Future Networks"

[0067] },

[0068] {

[0069] "Topic Number": 1,

[0070] "Topic": "Optimization Algorithm",

[0071] "Subject Keywords": "evolutionary algorithm, swarm intelligence, genetic algorithm, goal search, heuristic optimization, swarm algorithm, particle swarm optimization (PSO), metaheuristic algorithm, mutation, multi-objective optimization, scheduling problem, ant colony algorithm, Pareto optimizer, NSGA, PSO, multi-objective scheduling, wolf pack algorithm, heuristic algorithm, problem solving",

[0072] "Research Direction": "Machine Learning"

[0073] },

[0074] {

[0075] "Topic Number": 2,

[0076] "Topic": "Multi-Agent",

[0077] "Subject Keywords": "Multi-agent systems, distributed control, event triggering mechanism, consensus algorithm, random system, cooperation strategy, deception attack, networked system",

[0078] "Research Direction": "Machine Learning"

[0079] }

[0080] And take some of the subject hierarchical structures in the field of life and health as examples:

[0081] {

[0082] "Topic Number": 1,

[0083] "Topic": "Biological Evolution",

[0084] "Subject Keywords": "species, phylogeny, diversity, evolution, conservation",

[0085] "Research Direction": "Future Agriculture"

[0086] },

[0087] {

[0088] "Topic Number": 2,

[0089] "Topic": "Respiratory Diseases",

[0090] "Subject Keywords": "asthma, chronic obstructive pulmonary disease, allergic diseases",

[0091] "Research Direction": "Biopharmaceuticals"

[0092] },

[0093] {

[0094] "Topic Number": 3,

[0095] "Topic": "Arrhythmia",

[0096] "Subject Keywords": "arrhythmia, atrial fibrillation, ablation, catheter",

[0097] "Research Direction": "Medical Equipment"

[0098] },

[0099] By mapping the patent data of papers to the corresponding subject-level clusters through the above examples, we can achieve a precise match between the patent data of papers and subject areas and research directions. This can clearly display the subject distribution and research directions of different technical fields, providing visual support for talent discovery and project matching.

[0100] According to some embodiments provided by this application, the subject to which each academic achievement included in the dynamic academic data belongs can be determined by the following processing method, including the following specific steps:

[0101] Based on the feature embedding processing of each academic achievement document, document embedding features, topic embedding features, and word embedding features are obtained; document embedding features are used to represent the semantic features of each academic achievement document, topic embedding features are used to represent the semantic features of each topic, and word embedding features are used to represent the semantic features of each word included in each academic achievement document;

[0102] Determine a first mapping relationship between documents and topics based on the distribution of documents of each academic achievement on different topics;

[0103] Determining a second mapping relationship between the word and the topic based on the distribution of each word on different topics;

[0104] Based on the first mapping relationship and the second mapping relationship, the subject to which each academic achievement belongs is determined.

[0105] In the embodiment provided by this application, academic achievements are in the form of documents, and feature embedding processing is performed on the documents of each academic achievement to obtain three embedding features: document embedding features, topic embedding features, and word embedding features. Document embedding features represent the semantic features of the documents of each academic achievement, which can reflect the core content and thematic tendency of the entire document. Topic embedding features represent the semantic features of each topic, which are used to describe the semantic connotations of different topics. Word embedding features represent the semantic features of each word contained in the documents of each academic achievement, which can capture the semantic information of a single word. Based on the distribution of the documents of each academic achievement on different topics, the probability distribution of the documents on each topic is calculated, and the corresponding relationship between the documents and the topics is determined, thereby obtaining the first mapping relationship between the documents and the topics. Based on the above-mentioned first mapping relationship and second mapping relationship, the corresponding relationship between the documents and the topics and the corresponding relationship between the words and the topics are comprehensively considered to finally determine the topic to which each academic achievement belongs. Through document embedding features, topic embedding features and word embedding features, the semantic features of academic achievements can be captured from different levels, providing richer semantic information. Combining the first mapping relationship between documents and topics and the second mapping relationship between words and topics, the topic to which the academic achievement belongs can be determined more accurately, avoiding the errors that may be caused by single-level analysis.

[0106] For example, the document embedding feature (denoted as E d ), topic embedding features (denoted as E t ), and word embedding features (denoted as E w There are many ways to do this, such as using pre-trained Transformer models (such as the BERT model), Word2Vec, and GloVe algorithms. BERT is a pre-trained language model based on Transformer that can capture contextual information and generate context-related word embeddings. It provides powerful language representation capabilities and is suitable for a variety of natural language processing tasks. Word2Vec is a word embedding method based on neural networks. It maps each word to a fixed-length vector space by training a two-layer neural network. GloVe is a word embedding method based on global lexical statistical information. It learns the vector representation of each word by minimizing the weighted square error of the co-occurrence matrix between words.

[0107] Specifically, the FASTopic topic model can be adopted, and the fast Sinkhorn algorithm can be used to solve the optimal transmission between documents, topics and word embeddings, generate the topic distribution of documents, and construct a topic word co-occurrence matrix to capture the semantic associations between topic words, thereby mapping academic achievement data to corresponding research topics.

[0108] FASTopic is an efficient, flexible, and transferable topic model for automatically discovering latent topics from large-scale text data. Topic modeling is achieved by modeling optimal transfer plans between document, topic, and word embeddings. In FASTopic, the Sinkhorn algorithm is used to compute optimal transfer plans between documents and topics, and between topics and words. The Sinkhorn algorithm is an iterative method for solving the optimal transfer problem, finding an approximate optimal transfer plan by applying iterative normalization steps to a given cost matrix.

[0109] For ease of understanding, the following example is given, assuming that the document collection is D ={ d 1, d 2,…, dN},in represents the number of documents. The optimal transfer between documents, topics, and word embeddings can be expressed as:

[0110] ;

[0111] in, is the document-topic transfer matrix, is the topic-word transfer matrix, and the Sinkhorn algorithm is used to solve the optimal transfer. The document-topic distribution is the first mapping relationship. di The topic distribution (i is used to identify different documents) can be expressed as: ,in, Is a vector of all 1s. The distribution of keywords is recorded as the second mapping relationship, the topic The word distribution can be expressed as: .

[0112] To optimize the above topic model, the following objective function can be minimized :

[0113]

[0114] Where M is the number of topics, represents the Euclidean norm, represents the embedding features of the i-th document, The embedding features of the jth topic, represents the transmission probability distribution between the jth topic and all words, represents the transmission probability distribution between the i-th document and all topics.

[0115] above Used to represent all document embedding features and optimal transmission plan The sum of squared distances between the mapped topic embedding features. This reflects the total difference between all document embedding features and the topic embedding features mapped by the optimal transfer plan. The smaller the total difference, the better the match between the document embedding features and the topic embedding features, that is, the more accurate the topic model's representation of the document.

[0116] Similarly, the above The sum of the squares of the distances between the embedding features used to represent the jth topic and its representation in the word space measures the total difference between all topic embedding features and the word embedding features mapped by the optimal transfer plan. It reflects the total difference between all topic embedding features and the word embedding features mapped by the optimal transfer plan. The smaller the total difference, the better the match between the topic embedding features and the word embedding features, that is, the more accurate the representation of the word by the topic model.

[0117] For example, visualization can be performed based on the distribution of topics in the field. Taking some keywords in the field of life and health as an example, Figure 2 A schematic diagram showing the distribution of topics in the information processing method according to an embodiment of the present application is shown in FIG. Figure 2 The figure below shows a keyword bar chart for four different topics (Topic 29, Topic 108, Topic 79, and Topic 195). Each bar chart represents the top-weighted keyword within that topic and its relative importance (indicated by the length of the bar). These keywords were extracted from a large number of documents using a text analysis technique (such as topic modeling) and represent the core content of each topic.

[0118] The keywords of Topic 29 include hearing, speech, auditory, loss, and noise, indicating that it is related to hearing health or hearing loss.

[0119] Keywords in Topic 108 include remission, disease, severe, corticosteroids, and inhibitors, indicating relevance to the treatment and management of a disease.

[0120] Keywords for Topic 79 include glioma, glioblastoma, gbm, gliomas, and tumors, indicating a relevance to brain tumors.

[0121] The keywords of Topic 195 are antioxidant, flavonoids, curcumin, quercetin and polyphenols, indicating that it is related to nutrition, health food or disease prevention.

[0122] The numerical value next to each keyword (e.g., 0.06, 0.04, 0.15, etc.) indicates the weight or frequency of the keyword in the corresponding topic. The larger the numerical value, the more important or common the keyword is in the topic. The above visualization method helps to quickly understand the main focus and research areas of each topic.

[0123] In order to make feature embedding more adaptable to different fields, according to the embodiment provided by this application, before performing feature embedding processing based on the document of each academic achievement, the following steps are also included:

[0124] Obtain domain corpus data corresponding to multiple fields;

[0125] Performing semantic representation adjustments on the initial models corresponding to the multiple fields on the domain corpus data of the corresponding fields, thereby obtaining adjusted initial models corresponding to the multiple fields;

[0126] The adjusted initial models corresponding to the multiple fields are used to perform feature embedding processing.

[0127] In an embodiment of the present application, domain corpus data corresponding to multiple fields (such as life health, smart Internet of Things, etc.) are obtained. The above-mentioned corpus data can cover various text forms in various fields, such as papers, patents, technical reports, etc., to ensure that the model can learn the specific semantics and word usage habits of each field. The initial model is semantically represented and adjusted based on the domain corpus data of each field. The purpose is to make the model better adapt to the semantic characteristics of the specific field. The fine-tuned model can more accurately capture the semantic information of documents, topics and words, thereby improving the quality of feature embedding.

[0128] For example, the collected domain data is fine-tuned according to different fields to improve the representation capabilities of domain terms. Specifically, using BGE (BAAI General Embedding) as the base model (i.e., the initial model), fine-tuning on the domain corpus improves the semantic representation capabilities of professional terms (such as "immune checkpoint inhibitors" and "single-cell sequencing"), which can be represented in the following ways:

[0129]

[0130] in, Represents the initial model, such as Sentence-BERT or RoBERTa. Sentence-BERT is a variant of BERT that generates sentence embeddings. RoBERTa is also a pre-trained model of BERT. By improving training strategies and optimization methods, it improves the model's performance on various natural language processing tasks. Represents domain corpus data. Represents the learning rate, which is used to control the magnitude of model parameter updates. It represents the fine-tuning process, which updates the model parameters to adapt to the semantic information of a specific domain by training on the domain corpus. Represents a fine-tuned model that has improved its ability to represent domain terms.

[0131] In the case where the academic achievements are papers, according to some embodiments provided in this application, based on the field information corresponding to each academic achievement, the contribution statistics of the academic achievements of multiple candidate experts in different fields are collected to obtain the field contribution information corresponding to the multiple candidate experts, including the following specific steps:

[0132] In the case where the type of academic achievement is a paper, the first contribution weight of each paper is determined according to the product of the journal impact factor of each paper indicated in the dynamic academic data, the author ranking coefficient of the academic achievement, and the citation number attenuation factor;

[0133] For each of the multiple candidate experts, statistics are performed based on the field information corresponding to each paper and the first contribution weight of each paper to determine the single-field contribution of each expert in the multiple fields.

[0134] Based on the single-field contribution and the number of single-field papers of each expert in multiple fields, the field contribution information corresponding to multiple candidate experts is determined.

[0135] In the embodiment provided in this application, the first contribution weight of each paper is determined by comprehensively considering the journal impact factor, author ranking coefficient and citation count decay factor, and the contribution of candidate experts in various fields is calculated accordingly. The journal impact factor reflects the academic influence of the journal and is an important indicator for measuring the quality of a paper. The author ranking coefficient is determined based on the author's ranking in the paper. Generally, the first author has the greatest contribution, and the higher the ranking, the higher the coefficient. The citation count decay factor reflects the decay of the number of citations of the paper over time and is used to evaluate the timeliness of the paper. For each of the multiple candidate experts, their papers in various fields are collected, and based on the first contribution weight of each paper, the total contribution weight of the expert in each field is calculated. The single-field contribution of each expert's papers in each field is calculated, that is, the sum of the first contribution weights of all papers in the field. For each expert, a comprehensive indicator of their single-field contribution and the number of papers in each field is calculated. It can be a weighted sum of the single-field contribution and the number of papers. The weights can be adjusted according to actual needs, thereby obtaining each expert's field contribution information in each field for subsequent talent matching and evaluation.

[0136] Specifically, the first contribution weight can be expressed in the following manner:

[0137] First contribution weight = journal impact factor × author ranking coefficient Wi × citation count attenuation factor.

[0138] Among them, the author ranking coefficient It reflects the author's contribution to the paper and is determined by the author's ranking in the paper. For example, the author ranking coefficient can be obtained in a variety of ways, such as assigning contribution weights based on the author's ranking in the paper, or using linear decreasing method, exponential decreasing method, custom weight assignment, etc. Linear decreasing method: Given a decreasing weight sequence, such as 1, 0.8, 0.6..., directly assign weights based on the author's ranking The exponential decrease method uses an exponential function to determine the weight. The following example illustrates how to assign contribution weights based on the author's ranking in the paper. The specific calculation formula is as follows:

[0139]

[0140] Here, i represents the author's ranking position (counting starts from 1), and n represents the total number of authors in the paper, emphasizing that authors with higher rankings have greater contributions. Figure 3 The author contribution distribution diagram of the information processing method of the embodiment of the present application is shown. Let n=4, and the author contribution curve can be seen as follows: Figure 3 shown.

[0141] Here, n=4 indicates there are four authors. The horizontal axis represents the author number, from 1 to 4, with number 1 representing the first author, number 2 representing the second author, and so on. The vertical axis represents the contribution distribution, that is, the weight of each author's contribution to the paper.

[0142] Figure 3 The dotted lines and circles in the figure indicate that each author's contribution weight decreases as their author number increases. For example, the first author (number 1) has the highest contribution weight, close to 0.5. This weight decreases as the author number increases. The second author (number 2) has a contribution weight of approximately 0.25, the third author (number 3) has a further decrease, around 0.15, and the fourth author (number 4) has the lowest contribution weight, close to 0.1.

[0143] In the case where the academic achievements are patents, according to some embodiments provided in this application, based on the field information corresponding to each academic achievement, the contribution statistics of the academic achievements of multiple candidate experts in different fields are collected to obtain the field contribution information corresponding to the multiple candidate experts, including the following specific steps:

[0144] In the case where the type of academic achievement is a patent, the second contribution weight of each patent is determined according to the product of the classification number coverage, status coefficient, and citation number attenuation factor of each patent indicated in the dynamic academic data; the status coefficient includes the quantitative coefficient of the authorization status, and the classification number coverage indicates the scarcity of the technology corresponding to the patent;

[0145] For each of the multiple candidate experts, according to the field information corresponding to each patent and based on the first contribution weight of each patent, determine the patent single field contribution of each expert corresponding to the multiple fields;

[0146] Based on the patent field contribution and the number of patent fields corresponding to each expert in multiple fields, the field contribution information corresponding to multiple candidate experts is determined.

[0147] In the embodiment provided in the present application, for academic achievements of patent types, the second contribution weight of the patent is determined by considering the classification number coverage, status coefficient and citation number attenuation factor of the patent, and the patent contribution of the candidate experts in different fields is calculated accordingly. The classification number coverage indicates the scarcity of the technology corresponding to the patent. The lower the coverage, the scarcer the patent technology is, and the greater its contribution weight may be. The status coefficient includes a quantitative coefficient of the authorization status, which reflects the authorization status or public status of the patent, and to a certain extent indicates the possibility of implementation of the patent. The counting method of the status coefficient can be set as needed. For example, the higher the status coefficient, the greater the implementation and application potential of the patent. For each expert, collect their patents in various fields. According to the second contribution weight of each patent, count the total contribution weight of the expert in each field. Calculate the patent single-field contribution of each expert in each field, that is, the sum of the second contribution weights of all patents in the field. For each expert, calculate the comprehensive index of his / her patent contribution in each field and the number of patents. Similar to the paper, it can be the weighted sum of the patent contribution in each field and the number of patents. The weight can be adjusted according to actual needs, and finally obtain the field contribution information of each expert in each field for subsequent talent matching and evaluation.

[0148] Exemplarily, the above-mentioned status coefficient setting method can be for an authorized patent in a public state, with a status coefficient set to 1.0, which means that the patent has passed the review and obtained authorization, and has high effectiveness and value. The coefficient is set to 0.5 for a public unauthorized patent. Furthermore, for the unauthorized status, it can be further divided according to the application stage. For example, if it is still in the substantive examination stage, the coefficient is set to 0.8, indicating that the examination result has not yet been determined, but it has passed the preliminary examination and has certain potential value. The coefficient is set to 0.6 for the review stage, indicating that the patent application enters the review stage after being preliminarily rejected. The applicant has the opportunity to refute the reasons for rejection, but there is still uncertainty. The coefficient of the rejected or invalid status is set to 0.5, indicating that the patent application has been formally rejected or declared invalid, and its effectiveness and value are low.

[0149] Specifically, when calculating the second contribution weight of a patent, the above-mentioned status coefficient can be combined with the classification number coverage and the citation count attenuation factor as shown below:

[0150] Second contribution weight = classification number coverage × status coefficient × citation number attenuation factor

[0151] In this way, the value and influence of patents can be evaluated more accurately, and the single-field contribution and field contribution information of candidate experts in different fields of patents can be determined.

[0152] According to an embodiment of the present application, after executing step S101: based on dynamic academic data, determining the field contribution information corresponding to multiple candidate experts, the method also includes making talent portraits of the candidate experts based on the above-mentioned field mapping relationship (i.e., talent map). Talent portrait is a multi-dimensional, visual description of personal abilities and characteristics formed by integrating information from multiple dimensions such as a person's professional skills, educational background, academic achievements, and research directions. The above-mentioned field mapping relationship clusters the technical fields to obtain field clusters, which can be hierarchically divided according to themes and research directions. Dynamically associating paper authors and patent inventors with technical subject hierarchies specifically includes the following steps:

[0153] Count the number of first academic achievements of each candidate expert in each topic and the number of second academic achievements in each field cluster;

[0154] According to the correspondence between each candidate expert and academic achievements, determine the subject distribution coefficient of each candidate expert in different topics in a single field;

[0155] Based on the number of first academic achievements, the number of second academic achievements, the subject distribution coefficient, and the field contribution information, the talent label of each candidate expert is determined.

[0156] In an embodiment of the present application, by counting the number of first academic achievements of each candidate expert in each topic and the number of second academic achievements in each field cluster, the contribution and influence of the expert in a specific technical topic and field can be evaluated more finely, which is conducive to distinguishing the depth and breadth of the expert's research on different technical topics. Determining the subject distribution coefficient of each candidate expert in different topics in a single field can quantify the research interests and professional field distribution of the candidate expert, which helps to show whether the expert's research is concentrated on a specific topic or across multiple related topics. The correspondence between each candidate expert and the academic achievement includes dynamically associating the paper author and the patent inventor with the technical topic level, which can realize real-time tracking and updating of the expert's research results. The dynamic association can ensure the timeliness and accuracy of the talent portrait, so that it can reflect the expert's latest research results and research direction. Providing real-time data-driven expert evaluation and talent labeling for decision makers can enhance decision support in talent selection, project allocation, resource allocation, etc.

[0157] For example, the method of determining talent labels can be to divide candidate experts into talent gradients, for example, by analyzing indicators such as the number of academic achievements, field contributions, and influence of candidate experts in various research fields, to identify experts with outstanding performance in specific fields.

[0158] If a candidate expert has published ≥ the first threshold number of papers (e.g., 30 papers) or has ≥ the second threshold number of authorized patents (e.g., 10 patents) in a certain research direction, and his / her contribution and influence are greater than the first predetermined contribution threshold, the talent label will be marked as "field expert";

[0159] If a candidate expert has published ≥ the third threshold (e.g., 5 papers) or has ≥ the fourth threshold (e.g., 3 authorized patents) in a certain research direction, and his / her contribution and influence are greater than the second predetermined contribution threshold, he / she will be marked as "key citation";

[0160] If the candidate expert's contribution and influence in a certain research direction is greater than the third predetermined contribution threshold, he / she will be marked as a "preliminary screening talent";

[0161] If the candidate expert's contribution and influence in multiple research fields is greater than the fourth predetermined contribution threshold, he or she will be marked as a "compound talent".

[0162] After determining the labels, all labels corresponding to each candidate expert are integrated to determine the talent profile of the candidate expert.

[0163] After executing step S102, the method further includes:

[0164] When it is detected that there is incremental data in the dynamic academic data, the incremental data is input into the first target model for processing to obtain updated data;

[0165] Update the updated data to the initial resume according to the scheduled sequence.

[0166] In the embodiments of this application, dynamic academic data changes over time, and incremental data such as new research results and citation counts are constantly generated. Timely inputting incremental data into the first target model for processing and updating the initial resume ensures that the information in the initial resume keeps pace with academic trends and always reflects the candidate's latest academic contributions and capabilities, avoiding inaccurate expert evaluations due to outdated data.

[0167] In the embodiment provided in the present application, the first target model is trained based on a training set of public resumes and reference academic information, but as time goes by and the academic environment changes, the training set also needs to be updated in real time. On the one hand, new public resumes and corresponding reference academic information are continuously collected and added to the training set. Public resumes can be obtained from channels such as various academic websites and expert personal homepages through web crawler technology, and information retrieval and matching algorithms are utilized to find corresponding reference academic information, such as relevant academic achievement citations, cooperative project experience, etc. On the other hand, an online learning algorithm is adopted to carry out real-time training and optimization of the first target model. When new training data arrives, the model can adjust parameters in time to adapt to the new data distribution and characteristics, and improve the accuracy and real-time performance of generating the initial resume.

[0168] According to the embodiments provided herein, in step S103, the initial resume is screened and reconstructed according to a predetermined query template to obtain reconstructed resumes corresponding to multiple candidate experts. The query template can be generated, for example, by incorporating natural language processing (NLP) technology to construct a semantic understanding model (e.g., a shared second target model). The predetermined query template is no longer limited to fixed keyword matching but instead allows users to express their query requirements in natural language. For example, if a user enters "experts who have made significant breakthroughs in the field of artificial intelligence in the past five years," the semantic understanding model will parse the user's query and extract key semantic information, such as "timeframe: past five years," "field: artificial intelligence," and "nature of achievement: significant breakthrough." The initial resumes are then screened and reconstructed based on this semantic information. During this screening process, knowledge graph technology is combined to perform correlation analysis on information such as academic achievements and research fields to determine whether they meet the semantic requirements of "significant breakthrough," ultimately resulting in a reconstructed resume that meets the user's needs.

[0169] Through the above processing, the flexibility and convenience of query template customization are improved, the threshold for users to use query templates is lowered, and the limitation of mastering keyword grammar rules is removed.

[0170] For example, query templates are dynamically generated and adjusted based on different application scenarios and user behavior data. For example, historical user query behavior and usage scenario information (such as academic collaboration and talent recruitment) are recorded, and machine learning algorithms are used to analyze user preferences and demand patterns. For example, in talent recruitment scenarios, the focus is on experts' project experience and practical achievements. This means that the query template can be automatically adjusted to increase the screening weight of fields related to project experience and practical achievements. Furthermore, when new research hotspots are detected in an academic field, relevant screening criteria, such as emerging technology keywords and hot research directions, are automatically added to the query template, allowing for more targeted screening and reconstruction of the initial resume.

[0171] For example, a query template containing multiple dimensional screening conditions can be set to implement a multi-dimensional cross-query. Based on the multi-dimensional screening conditions, the initial resume is cross-screened and reconstructed, and the expert information is comprehensively evaluated from different angles to ultimately generate a reconstructed resume that meets the multi-dimensional requirements.

[0172] In some embodiments provided herein, the reconstructed resume can be weighted based on existing expert information. For example, some experts may prioritize their latest research achievements in a specific field, while others may highlight the growing trend of their academic influence. Furthermore, based on the expert's real-time social network information (such as the frequency of interaction with experts in other fields and their participation in discussions on popular academic topics), the display weight of relevant content in the reconstructed resume can be dynamically adjusted to better reflect the expert's real-time capabilities.

[0173] According to the embodiment provided by the present application, in step S105: based on the query corpus, talent profiles corresponding to multiple candidate experts are screened, and the target expert required for the target project is determined from the multiple candidate experts, including the following specific steps:

[0174] Generate prompt word information based on the query corpus, the prompt word information including at least one of an expanded word, a translated name, a synonym, and an abbreviation for the query corpus;

[0175] Generate screening conditions according to the prompt word information;

[0176] Determine the target profile that meets the screening criteria from the reconstructed resumes of multiple candidate experts;

[0177] Perform correlation analysis based on the prompt word information and determine the target node from multiple types of nodes in the predetermined graph, including personal information nodes, skill nodes, project experience nodes, academic achievement nodes, and talent profile nodes;

[0178] Determine the target expert based on the candidate experts corresponding to the target portrait and the candidate experts corresponding to the target node.

[0179] In the embodiments provided herein, by generating prompt word information containing the query's expanded terms, translated names, synonyms, and abbreviations, the scope of the query can be greatly expanded, avoiding the omission of relevant information due to differences in wording. For example, for the query "artificial intelligence," including "AI" (the abbreviation) and "Artificial Intelligence" (the English translation) as prompt words ensures that all relevant information is found. Using prompt word information to generate screening criteria allows for more accurate selection of target profiles from talent profiles. These screening criteria consider a variety of possible wordings, improving the accuracy and comprehensiveness of screening and reducing the possibility of misjudgments and missed detections. By performing correlation analysis and identifying target nodes based on the prompt word information among various node types in a predetermined graph, candidate experts can be matched across multiple dimensions (such as personal information, skills, and project experience). This helps to fully understand the candidate's abilities and background, and more accurately identify experts that closely match the target requirements. By combining the candidate experts corresponding to the target profile with the candidate experts corresponding to the target node to determine the target expert, information from different sources and dimensions is integrated, ensuring that the final target expert more closely meets the actual needs of the project and improving the matching between the expert and the project.

[0180] According to some embodiments provided by the present application, the method further includes: extracting data from the structured domain contribution information to obtain first entity data and first relationship data; generating a predetermined graph based on the first entity data and the first relationship data; or,

[0181] Data extraction is performed based on the reconstructed resumes corresponding to the multiple candidate experts to obtain second entity data and second relationship data; based on the second entity data and the second relationship data, a predetermined graph is generated, and the resumes are reconstructed into an unstructured data form.

[0182] In the embodiments provided in the present application, two optional graph generation methods are provided. One is that for structured field contribution information, entity data and relationship data can be directly extracted from it, making full use of the characteristics of structured data that are easy to analyze and process, and quickly and accurately providing key data for generating a predetermined graph. The other is that for unstructured reconstructed resumes, valuable information can be mined through data extraction technology, and unstructured data that was originally difficult to use directly can be converted into entity and relationship data that can be used to build a graph, broadening the data source and improving data utilization. As a structured knowledge representation method, the predetermined graph can clearly show the relationship between experts and between experts and various fields, projects, etc. Based on such a comprehensive and accurate graph, in the decision-making process of determining target experts, the ability, experience and relevance of experts can be analyzed more accurately, so as to make decisions that are more in line with actual needs and improve the quality and efficiency of decision-making.

[0183] According to the embodiment provided by the present application, after step S101, the method further includes the following specific steps:

[0184] When it is detected that there is incremental data in the dynamic academic data, the incremental data is input into the first target model for processing to obtain updated data;

[0185] Update the updated data to the initial resume according to the scheduled sequence.

[0186] In the embodiment provided by the present application, dynamic academic data is detected, and when incremental data is found, these incremental data are input into the first target model for processing. After being processed by the first target model, updated data is obtained, and the incremental data is processed by the first target model, and the new data can be converted into updateable content according to the rules and logic of the model, and the obtained updated data is updated to the initial resume according to a predetermined time sequence. The incremental data in the dynamic academic data can be captured in time, so that the initial resume can be updated in real time according to the newly generated data, ensuring that the resume content always reflects the latest academic situation and avoiding outdated resume information.

[0187] For example, a dynamic update mechanism can be established to regularly collect new patent and paper data and update the embedding model to capture emerging technologies. Automated scripts or data collection tools can be used to regularly collect new data from academic and patent databases. Data collection can be automated on a scheduled basis (e.g., daily, weekly, or monthly) or triggered based on specific conditions (e.g., database update notifications). New data can be used to fine-tune existing embedding models to adapt to new technological trends and data distribution. Incremental learning strategies can also be employed to enable the model to learn new knowledge without forgetting previous knowledge.

[0188] Based on the above embodiment, this application also provides an optional implementation method, which is to obtain expert information required by the project by setting up a data collection and processing module, a model training and optimization module, a talent portrait module, and a talent matching model. The specific explanation is as follows.

[0189] (1) Data acquisition and processing module:

[0190] 1.1 Data Collection. Connect to academic and patent databases, and automatically collect journal literature and patent data through API interfaces and crawler technology. This includes information such as the paper's title, author, keywords, abstract, publication date, publishing institution, contact information, journal, citations, and citations, as well as patent information such as the title, inventor, abstract, claims, and classification number. This data is also collected from enterprise databases and stored in a distributed database for subsequent processing and analysis.

[0191] 1.2 Data Cleaning and Preprocessing. Clean the collected data to remove duplicate, erroneous, and irrelevant information. Perform preprocessing operations such as word segmentation, noise removal, and stemming on the text data to improve data quality and the accuracy of subsequent analysis.

[0192] 1.3 Data Fusion and Standardization. Integrate paper and patent data to establish a unified data model and standards. Standardize terminology, talent data, and institutional information to achieve data disambiguation and normalization. Match and integrate author information, institutional information, contact information, email addresses, academic background, and work experience from papers and patents to form a standardized expert and team data model, providing a foundation for subsequent analysis and mining.

[0193] Dynamic academic data can be obtained through the processing from 1.1 to 1.3 above.

[0194] (2) Model training and optimization module:

[0195] 2.1 Pre-trained Model Training and Fine-tuning. The collected domain data is fine-tuned based on different domains to improve the representation of domain terminology. Using BGE as the base model, fine-tuning is performed on domain corpus to improve the semantic representation of professional terminology.

[0196] 2.2 Topic Model Training. We used the FASTopic topic model, a fine-tuned BGE pre-trained model, and the fast Sinkhorn algorithm to achieve optimal transfer between documents, topics, and word embeddings. We generated topic distributions for documents and constructed a topic-word co-occurrence matrix to capture the semantic associations between topic words, thereby mapping the patent data to corresponding research topics.

[0197] 2.3 Construction of the Product Talent Map. We used hierarchical clustering and the DBSCAN algorithm to cluster topics and keywords, forming multiple technical field clusters. Leveraging the contextual understanding and reasoning capabilities of the large model, we analyzed the hierarchical relationships between keywords within each cluster and constructed a hierarchical relationship between technical tags (industry field - industry direction - research direction - topic - topic keywords), thus forming a product talent and technology map.

[0198] (3) Talent portrait module:

[0199] 3.1 Calculation of Talent Contribution. Each candidate expert's contribution is quantified. The weight of a single paper = Journal Impact Factor × Author Ranking Coefficient Wi × Citation Attenuation Factor. The weight of a single patent = Technology Scarcity (IPC Classification Coverage) × Status Coefficient (e.g., 1.0 for authorized patents, 0.5 for publicly unauthorized patents) × Citation Attenuation Factor.

[0200] Using multi-dimensional data mapping, based on the generated industry domain map, we dynamically associate paper authors and patent inventors with the technical theme hierarchy. We also count the number of papers and patents and thematic distribution coefficients of talents in each thematic research cluster. Based on the calculated contribution of individual papers / patents, we calculate the talent's relevance across various thematic fields. We also analyze the expert's academic influence in each research field (i.e., the field contribution information for multiple candidate experts), which is used to enrich the talent profile in 3.2.

[0201] 3.2 Generation of Talent Resumes and Talent Portraits

[0202] The first target model trained with public resumes and reference academic information is used to process the collected dynamic academic information. For example, a talent resume (i.e. the above-mentioned initial resume) is generated based on the relevant information of the candidate expert (educational background, work experience, academic research results, contribution to the industry, etc.).

[0203] The first target model can output an instance of the initial resume in the following example form:

[0204] Personal information: Name: Omitted;

[0205] Research areas: machine learning, deep learning, feature selection, image processing (SAR / remote sensing), network architecture optimization, and unsupervised learning;

[0206] Technical fields: software defect prediction, multimodal data fusion, target detection, change detection, and model lightweighting.

[0207] Professional skills:

[0208] 1. Machine Learning and Deep Learning:

[0209] Proficient in feature selection algorithms (sparse representation, non-negative matrix factorization, manifold learning);

[0210] Proficient in using CNN, RNN, autoencoder, contrastive learning, graph neural network (GNN) and other models;

[0211] Develop multi-objective optimization algorithms (such as the decomposition-based MOMAD algorithm) and dynamic propagation models.

[0212] 2. Image Processing and Computer Vision:

[0213] Satellite / remote sensing image processing: object detection (SDANet), semantic segmentation (CMNet), change detection (StackedFisher autoencoder);

[0214] SAR image classification (DSNet, Contourlet-CNN) and 3D reconstruction (2D to 3D motion-adaptive depth estimation).

[0215] 3. Algorithm optimization and engineering implementation:

[0216] Network architecture search (EF-ENAS), model lightweighting (depthwise separable convolution, low-rank sparse representation);

[0217] High-performance computing: HDFS distributed storage optimization, multi-source data fusion (DMULN).

[0218] 4. Tools and Languages:

[0219] Programming languages: Python, MATLAB.

[0220] Frameworks: TensorFlow, PyTorch, Keras.

[0221] Data processing: NumPy, Pandas, OpenCV.

[0222] Research fields and application directions:

[0223] 1. Software Defect Prediction:

[0224] A two-stage data preprocessing method (RKEE) was proposed, combining feature selection and noise filtering to significantly improve the AUC index;

[0225] Application scenarios: Quality assurance of large software projects such as NASA and Eclipse.

[0226] 2. Image processing and pattern recognition:

[0227] SAR / remote sensing imagery: Development of lightweight networks (DSNet) and semantic embedding density adaptive detectors (SDANet);

[0228] Medical / industrial images: Contrastive Learning-based Anomaly Detection (MCOD), Weakly Supervised Semantic Segmentation (SAGN).

[0229] 3. Network architecture and optimization:

[0230] Design mixed precision quantization network (SSPS) and evolutionary neural network architecture search (EF-ENAS);

[0231] The dynamic immune node model (NICT) is used to control information propagation in complex networks.

[0232] 4. Unsupervised and Semi-supervised Learning:

[0233] Feature selection algorithms (UFSRL, NNSAFS), cluster alignment framework (Cluster Alignment);

[0234] Contrastive learning (MCOD), pseudo label generation (SLREO).

[0235] Representative scientific research achievements:

[0236] 1. Algorithm innovation

[0237] RKEE framework: Combines rough set theory and KNN noise filtering to solve the class imbalance problem of software datasets (significantly improves AUC);

[0238] DSNet: A lightweight network based on depthwise separable convolution, with the number of parameters reduced to 1 / 9 of traditional CNN, suitable for small-sample PolSAR classification;

[0239] SDANet: A semantically embedded density-adaptive detector for high-accuracy moving vehicle detection in satellite videos with reduced background clutter.

[0240] 2. Theoretical Contributions:

[0241] The generalized random matrix product theory is proposed to extend the traditional Hajnal inequality and apply it to the consistency analysis of multi-agent systems.

[0242] Develop a multi-objective sparse subspace learning framework based on Pareto front to improve the flexibility of feature selection for classification tasks.

[0243] 3. Engineering application:

[0244] HDFS optimization: Replica placement strategy based on multi-objective optimization improves storage and network load balancing;

[0245] Dynamic Propagation Model: Used for information suppression in social networks. Experimental verification shows that the immunity effect is better than similar methods in 8 real networks.

[0246] It should be noted that the above-mentioned specific content, specific technical means, technical names, and English abbreviations are merely schematic representations of the initial resume and do not limit the technical scope.

[0247] The above initial resume can be further restructured, and the predetermined query template can be as follows:

[0248] text="Analyze based on the information below."+="Name of this person:"+str(talent['authorName'])+"\n"

[0249] +"This person's education experience is as follows:"+str(talent['education_str'])+"\n"

[0250] +"This person's work experience is as follows:"+str(talent['employment_str'])+"\n"

[0251] +"The papers and patents published by this person are as follows:"+str(talent['article'])+"\n"

[0252] +"This person's field contribution is as follows: +str(talent['area'])+"\n"

[0253] {"role": "system", "content": "You are a senior HR in the field of life sciences and health / intelligent IoT. The output results are in Chinese. Generate a resume based on the following information and analyze the person's talent profile."},

[0254] {"role": "user", "content": text}

[0255] The "text" key is used to store text content and will be assigned and concatenated in the following code. "authorName" represents the author's name. In the code, talent['authorName'] retrieves the value corresponding to the key named authorName, which is the talent's name. "education_str" represents an educational background string. talent['education_str'] retrieves the value corresponding to the key named education_str, which is the talent's educational background information, presented as a string. "employment_str" represents an employment experience string. talent['employment_str'] retrieves the value corresponding to the key named employment_str, which is the talent's employment experience information, presented as a string. talent['article'] retrieves the value corresponding to the key named 'article'_str, which is used to retrieve the talent's papers and patents. talent['area'] retrieves the value corresponding to the key named 'area', which is used to retrieve the talent's contributions to a field.

[0256] "Role" means "role", "system" means the identity attribute of the system, and "content" is used to specify the specific tasks, descriptions, requirements, etc. corresponding to the role. For example, if the role is system, it specifies the tasks that should be performed by the system role.

[0257] The above query template can be used to obtain a reconstructed resume for the above resume example. Combined with the talent contribution obtained in 3.1, the quantitative values ​​are semantically converted to express understandable semantic expressions. For example, more than five years of in-depth experience in the same field can be expressed as "expert" or "senior", and project experience in multiple fields can be integrated into "comprehensive talent". For talent profile analysis, it can be expressed as follows:

[0258] Core advantages:

[0259] 1. Cross-domain technology integration capabilities:

[0260] From software engineering (defect prediction) to remote sensing image processing, to network architecture design, it demonstrates multidisciplinary innovation capabilities;

[0261] He is good at converting theories (such as rough sets and manifold learning) into engineering solutions (such as SDANet and DSNet).

[0262] 2. Complex problem modeling capabilities:

[0263] When processing high-dimensional, small-sample, and noisy data (such as SAR images), improve model robustness through methods such as sparse representation and contrastive learning;

[0264] In unsupervised scenarios, pseudo-label generation (SLREO) and cluster alignment strategies are designed to solve the problem of scarce labeled data.

[0265] 3. Algorithm performance optimization experts:

[0266] Model lightweighting: Deep Separable Convolution (DSNet) and Low-Rank Sparse Decomposition (NMF-LRSR) significantly reduce computational costs;

[0267] Real-time optimization: EF-ENAS uses evolutionary search to quickly determine mixed-precision quantization strategies to meet hardware resource constraints.

[0268] Suitable positions:

[0269] Industry: Senior algorithm engineer (autonomous driving / remote sensing), AI platform architect, cloud computing optimization expert.

[0270] Academia: Machine learning researcher, leader of a computer vision laboratory.

[0271] Professional competitiveness:

[0272] Academic influence: Published multiple papers in top journals such as IEEE TIP, IEEE TGRS, and Neural Networks, with high H-index potential.

[0273] Technology implementation capabilities: Many achievements (such as HDFS optimization and SAR classification) can be directly applied to industrial scenarios and have commercial potential.

[0274] Summary: This candidate is a compound talent with both theoretical depth and engineering breadth, especially outstanding in processing high-dimensional data, small sample learning, and model lightweighting. He is suitable for leading technical breakthroughs or cross-domain R&D teams.

[0275] 3.3 Industry talent map.

[0276] Utilize GraphRAG / Noderag technology to process and analyze the resumes generated by 3.2, construct heterogeneous graphs, and integrate and correlate multi-dimensional data such as personal information, educational background, work experience, skills, research direction, project experience, etc., to achieve more efficient talent recommendation and analysis.

[0277] GraphRAG (Graph Retrieval-Augmented Generation) is an advanced method that combines graph databases with retrieval-augmented generation (RAG) techniques. It leverages graph-structured data (such as graphs) to enhance the retrieval and generation capabilities of language models. Noderag, a key component of GraphRAG, focuses on node ranking mechanisms within graphs, optimizing node retrieval and ranking through the use of graph neural networks (GNNs) or other graph analysis techniques (such as community detection and node importance ranking).

[0278] The following are the processing solutions and implementation steps based on GraphRAG technology:

[0279] 3.3.1 Constructing Nodes and Edges of Heterogeneous Graphs

[0280] 3.3.1.1 Node type: 1) Personal information node: name, contact information, educational background (school, major, degree), work experience (company, position, time).

[0281] 2) Skill nodes: technical skills (such as Python, TensorFlow, image processing); soft skills (such as teamwork and communication skills).

[0282] 3) Research direction node: specific research direction (such as machine learning, deep learning, federated learning); research field (such as computer vision, biomedical engineering).

[0283] 4) Project experience nodes: project name, project description, technology stack used, and project results.

[0284] 5) Research achievement nodes: paper publication (journals, conferences, impact factors), patent applications, software open source (GitHub projects).

[0285] 6) Talent profiling nodes: core strengths (such as cross-domain technology integration capabilities and complex problem modeling capabilities); suitable positions (such as senior algorithm engineers and data scientists); professional competitiveness (such as academic influence and technology implementation capabilities).

[0286] 3.3.1.2 Edge Types: 1) Association Edge: Indicates the association relationship between different nodes, for example: the association between personal information nodes and skill nodes, the association between project experience nodes and skill nodes, and the association between scientific research results nodes and research direction nodes.

[0287] 2) Recommendation edge: used to represent suitable positions and career competitiveness in talent profiling analysis, for example: the recommendation relationship between the suitable position node and the specific position; the relationship between the career competitiveness node and potential development suggestions.

[0288] 3.3.2 Construction and Optimization of Graph Structure

[0289] 3.3.2.1 Graph Decomposition: Break down the information in your resume into fine-grained nodes. For example, you can further decompose "machine learning and deep learning" skills into specific skill nodes such as "feature selection algorithm," "CNN," and "RNN." Break down project experience into nodes such as project name, project description, and technology stack used.

[0290] 3.3.2.2 Graph Enhancement: Further enrich the graph structure by identifying key nodes and generating attribute summaries. For example, add an attribute summary for the "DSNet" node to explain its lightweight characteristics and application scenarios. Add an attribute summary for the "Federated Learning Protocol FEVERLESS" node to explain its privacy protection characteristics.

[0291] 3.3.2.3 Graph Enrichment: Embed personal information, skills, research interests, and other data into the graph and optimize the relationships between nodes using graph algorithms (such as PageRank and community detection). For example, use community detection algorithms to identify potential connections between skills and project experience. Use PageRank algorithms to identify key skills and project experience.

[0292] (IV) Talent matching model:

[0293] 4.1 Vector Model Enhancement. Collect text data from papers and patents, as well as relevant talent data. Use fine-tuned large-scale models such as BGE to pre-train embeddings on this data, generating vector representations for each paper, patent, and talent. Establish a dynamic update mechanism to regularly collect new patent and paper data and update the embedding model to capture emerging technologies.

[0294] 4.2 Intent Parsing and Query Expansion. Using the second target model to drive parsing, project requirement information is input and query corpus and expanded terms are output. These expanded terms can be processed in response to user clicks to obtain the final prompt term information.

[0295] 4.3 Using the prompt word information to generate screening criteria: Based on the reconstructed resumes obtained in section 3.2 above, a subset of candidate experts corresponding to the target profile is obtained. A second subset of candidate experts is also obtained based on the graph analysis results obtained in section 3.3 above, using the prompt word information for correlation analysis. The overlapping portion of these two subsets of candidate experts is selected as the final target expert.

[0296] 4.4 Talent Assessment Report Module. Use a large model to conduct in-depth semantic analysis of the papers and patents of recalled talents to identify their core technological contributions in specific fields. Through citation analysis and influence, further evaluate the influence and innovation of these technologies, and ultimately extract the expert's top 3 technical tags. Systematically search the expert's academic and patent records to screen out the papers and core patents with the highest impact factors. Analyze the academic value, practical application potential, and driving effect on the industry of these achievements to form the expert's academic and technological innovation trajectory. Combine the expert's technical advantages and achievements with the needs of the enterprise for matching analysis. Systematically evaluate the fit between the expert and the enterprise's technology gap, and rank and score them.

[0297] For example: 1. Target expert A:

[0298] Technological Advantages: Target Expert A has world-leading experience in optimizing gene editing tools and holds four high-value patents. Over the past three years, its research has shifted to AI-assisted off-target effect prediction, demonstrating its deep understanding and innovative capabilities in gene editing technology.

[0299] Milestone achievement: In 2022, expert A published a paper in an academic journal (impact factor 68.9), proposing a new variant. This achievement has been cited more than 500 times, demonstrating its outstanding influence in the field of gene editing.

[0300] Collaboration suggestion: Expert A's technical background is highly consistent with Pharmaceutical Company X's rare disease development needs. It is recommended that Expert A jointly develop a gene therapy pipeline for genetic diseases to fully utilize their expertise in gene editing and AI prediction.

[0301] Ranking score: 85 out of 100. Technical fit: 40 out of 50, impact of research: 35 out of 35, and collaboration potential: 10 out of 15. Expert A's technical strengths are highly aligned with the company's needs, his academic achievements are highly influential, and he has significant potential for collaboration.

[0302] 2. Target Expert B:

[0303] Technical Advantages: Target Expert B specializes in the application of CAR-T cell therapy in solid tumors, with a particular focus on tumor microenvironment modification. Its team has extensive experience in clinical translational research and combination therapy design, demonstrating strong clinical application capabilities.

[0304] Milestone achievement: In 2021, expert B published a paper in an academic journal (impact factor 66.8), detailing the new mechanism of CAR-T cell therapy in the treatment of solid tumors. The paper has been cited more than 300 times and is an important reference in this field.

[0305] Cooperation suggestion: Expert B's technical advantages in the field of solid tumor treatment are highly consistent with the company's needs. It is recommended to jointly develop solid tumor treatment solutions, especially in the aspects of tumor microenvironment modification and combination therapy design, to accelerate clinical application and promotion.

[0306] Ranking score: 80 out of 100. Technical fit: 38 out of 50, impact: 32 out of 35, and collaboration potential: 10 out of 15. Expert B has significant advantages in the application of CAR-T cell therapy for solid tumors, with strong academic impact and high collaboration potential.

[0307] 3. Target Expert C:

[0308] Technological Advantages: Target Expert C has extensive technical expertise in the development of PD-1 / PD-L1 inhibitors and immune checkpoint blockade. Its research team achieved significant breakthroughs in Phase III clinical trials, demonstrating strong R&D and clinical advancement capabilities.

[0309] Milestone achievement: In 2020, target expert C published a paper in an academic journal (impact factor 50.3), detailing the new progress of PD-1 / PD-L1 inhibitors in tumor treatment. The paper has been cited more than 200 times and is an important reference in this field.

[0310] Cooperation suggestion: Expert C's technical background is highly consistent with the company's needs in immunotherapy drug development. It is recommended to jointly develop immunotherapy drugs, especially in the optimization and clinical application of PD-1 / PD-L1 inhibitors, to accelerate product launch and market layout.

[0311] Ranking score: 75 out of 100. Technical fit: 35 out of 50, impact: 28 out of 35, and collaboration potential: 12 out of 15. Expert C has significant advantages in the development of PD-1 / PD-L1 inhibitors, with strong academic impact and high collaboration potential.

[0312] Compared with the prior art, this optional implementation has the following beneficial effects:

[0313] 1. Multi-dimensional data integration, comprehensive analysis of multi-dimensional data such as paper citations, keyword co-occurrence, author cooperation network, journal influence, etc., to provide a more comprehensive talent and team assessment, overcome the limitations of single-dimensional analysis of traditional methods, and improve the comprehensiveness and accuracy of the assessment. A topic model is used to construct a subject word co-occurrence matrix to capture the semantic associations between subject words, thereby mapping paper patent data to corresponding research topics. Hierarchical clustering + DBSCAN algorithm is used to cluster topics and subject words to form multiple technology field clusters. Based on the contextual understanding and reasoning capabilities of the large model, the hierarchical relationship between keywords in each cluster is analyzed, and the hierarchical relationship of technical labels is constructed to form a production and technology map.

[0314] 2. Dynamic tracking and updating: Introducing a dynamic analysis module to provide real-time updates on the latest research results of talent and teams, ensuring the timeliness and accuracy of evaluation results. This addresses the issue of delayed data updates in existing technologies and better reflects the latest developments of talent. Based on industry domain maps, this module dynamically associates paper authors and patent inventors with technical themes, calculates the relevance of talent across various thematic areas, analyzes the academic influence of experts in various research fields, and creates a talent profile.

[0315] 3. Talent profiles intuitively present key information about candidate experts, such as their educational background, work experience, and academic achievements. Recruiters or project leaders can directly access this information and understand each candidate's strengths and characteristics, without having to guess at the complex internal calculations and decision-making processes of black-box models. In black-box models, recommendation results often lack clear explanations, leaving users with only the final list of recommendations without understanding why certain candidates were selected. This provides strong interpretability when screening resumes based on project requirements, clearly explaining why certain experts were selected or excluded. Unlike black-box models, where the decision-making basis is often elusive and can lead to distrust of the results.

[0316] 4. A large-scale talent assessment report utilizes a large-scale model to conduct in-depth semantic analysis of talent's papers and patents, identifying their core technological contributions in specific fields. Citation analysis and impact analysis are used to further assess the impact and innovation of these technologies, ultimately extracting the expert's top three technology tags. A systematic search of the expert's academic and patent records identifies papers and core patents with the highest impact factors. These achievements are analyzed for their academic value, practical application potential, and impact on the industry, forming the expert's academic and technological innovation trajectory. This report combines the expert's technical strengths and achievements with enterprise needs, systematically assessing the fit between the expert and the enterprise's technological gaps, and ranking and scoring them.

[0317] 5. Cross-domain adaptability: The system can adapt to the unique needs of different fields, improving the versatility and flexibility of talent and team discovery. This makes the system not limited to a specific field, but can be widely applied in multiple industries and technology fields, enhancing the practicality and adaptability of the system.

[0318] The above beneficial effects make this optional implementation method superior to existing industrial talent discovery technologies in terms of technological advancement, evaluation accuracy, and wide application, providing more powerful data support for promoting regional scientific and technological innovation and talent introduction.

[0319] Figure 4 A schematic block diagram of an information processing method according to an embodiment of the present application is shown. Figure 4 As shown, corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application further provides an information processing device, including:

[0320] A field contribution determination module 401 is configured to determine field contribution information corresponding to a plurality of candidate experts based on dynamic academic data, wherein the dynamic academic data includes a citation count decay factor, which is used to indicate that the number of citations of an academic achievement decays over time, and the field contribution information includes the contribution of the corresponding candidate expert in a plurality of fields;

[0321] A resume generation module 402 is configured to generate initial resumes corresponding to a plurality of candidate experts based on dynamic academic data using a first target model, wherein the first target model is trained based on a training set including public resumes and reference academic information, and there is a correspondence between the reference academic information and the public resumes;

[0322] The resume reconstruction module 403 is used to screen and reconstruct the initial resume according to a predetermined query template to obtain reconstructed resumes corresponding to multiple candidate experts;

[0323] Talent portrait module 404 is used to perform semantic association based on the field contribution information and reconstructed resumes of the multiple candidate experts to obtain talent portraits corresponding to the multiple candidate experts;

[0324] Demand decomposition module 405 is used to process the project demand information of the target project using a second target model to obtain a query corpus; the second target model is trained based on historical demand information and historical expert information of reference projects;

[0325] The retrieval module 406 is used to screen the reconstructed resumes corresponding to the plurality of candidate experts based on the query corpus, and determine the target expert required for the target project from the plurality of candidate experts.

[0326] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0327] Figure 5 FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present application. Figure 5 As shown, the electronic device includes a memory 501 and a processor 502. The memory 501 stores a computer program that can be executed on the processor 502. When the processor 502 executes the computer program, the method of the above embodiment is implemented. The number of memory 501 and processor 502 can be one or more. In a specific implementation, the electronic device may also include a communication interface 503 for communicating with external devices and performing data exchange.

[0328] In a specific implementation, if the memory 501, processor 502, and communication interface 503 are implemented independently, the memory 501, processor 502, and communication interface 503 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0329] Optionally, in a specific implementation, if the memory 501 , the processor 502 , and the communication interface 503 are integrated on a chip, the memory 501 , the processor 502 , and the communication interface 503 may communicate with each other through an internal interface.

[0330] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the information processing method provided in the embodiment of the present application.

[0331] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the information processing method provided in the embodiment of the present application.

[0332] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the information processing method provided in the embodiment of the present application.

[0333] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the information processing method provided in the embodiment of the application.

[0334] It should be understood that the processor described above may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0335] Furthermore, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache memory. By way of example and not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).

[0336] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0337] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0338] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0339] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.

[0340] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from and execute instructions on an instruction execution system, apparatus or device), or used in conjunction with such instruction execution systems, apparatuses or devices.

[0341] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0342] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0343] The above are merely exemplary embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An information processing method, characterized in that: include: Determine, based on dynamic academic data, field contribution information corresponding to each of the multiple candidate experts, wherein the dynamic academic data includes a citation count decay factor, which is used to indicate that the number of citations of academic achievements decays over time, and the field contribution information includes the contribution of the corresponding candidate experts in multiple fields; Based on the dynamic academic data, a first target model is used for processing to generate initial resumes corresponding to the plurality of candidate experts, wherein the first target model is trained based on a training set including public resumes and reference academic information, and there is a correspondence between the reference academic information and the public resumes; According to a predetermined query template, the initial resume is screened and reconstructed to obtain reconstructed resumes corresponding to the multiple candidate experts; wherein the query template is set to include multiple dimensional screening conditions, and the initial resume is cross-screened and reconstructed according to the multiple dimensional screening conditions to remove irrelevant information in the initial resume; Performing semantic association based on the field contribution information and reconstructed resumes corresponding to the multiple candidate experts, obtaining talent portraits corresponding to the multiple candidate experts; Based on the project requirement information of the target project, a second target model is used to process the query corpus; the second target model is trained based on the historical requirement information and historical expert information of the reference project; Based on the query corpus, the talent portraits corresponding to the multiple candidate experts are screened, and the target expert required for the target project is determined from the multiple candidate experts.

2. The method according to claim 1, characterized in that The method further comprises: Obtaining a first number of citations of the academic achievement within a predetermined time range, and a second number of citations of the academic achievement outside the predetermined time range; Attenuation-adjust the preset first weight in each predetermined attenuation period according to a predetermined attenuation ratio to obtain a second weight outside the predetermined time range; The citation count attenuation factor is obtained by weighting based on the first weight, the second weight, the first citation count, and the second citation count.

3. The method according to claim 1, characterized in that The step of screening the talent profiles corresponding to the plurality of candidate experts based on the query corpus and determining the target expert required for the target project from the plurality of candidate experts includes: Based on the query corpus, generating prompt word information, the prompt word information including at least one of an expanded word, a translated name, a synonym, and an abbreviation for the query corpus; Generate screening conditions according to the prompt word information; Determining, from the talent portraits corresponding to the plurality of candidate experts, a talent portrait that meets the screening condition as a target portrait; Performing a correlation analysis based on the prompt word information to determine a target node from multiple types of nodes in a predetermined graph, wherein the multiple types of nodes include personal information nodes, skill nodes, project experience nodes, academic achievement nodes, and talent profile nodes; The target expert is determined based on the candidate experts corresponding to the target portrait and the candidate experts corresponding to the target node.

4. The method according to claim 3, characterized in that The method further includes: extracting data from the structured domain contribution information to obtain first entity data and first relationship data; generating the predetermined graph based on the first entity data and the first relationship data; or, Data extraction is performed based on the reconstructed resumes corresponding to the multiple candidate experts to obtain second entity data and second relationship data; based on the second entity data and the second relationship data, the predetermined graph is generated, and the reconstructed resume is in the form of unstructured data.

5. The method according to claim 1, wherein After the information included in the initial resume is arranged in chronological order and the dynamic academic data is processed using the first target model to generate initial resumes corresponding to the plurality of candidate experts, the method further includes: When it is detected that the dynamic academic data has incremental data, the incremental data is input into the first target model for processing to obtain updated data; The updated data is updated to the initial resume according to a predetermined time sequence.

6. The method according to any one of claims 1 to 5, characterized in that The determination of the field contribution information corresponding to each of the multiple candidate experts based on dynamic academic data includes: determining the subject to which each academic achievement included in the dynamic academic data belongs; Determining the field information corresponding to each academic achievement according to the field mapping relationship and the subject to which each academic achievement belongs, wherein the field mapping relationship represents the corresponding relationship between the multiple fields and the subject respectively; Based on the field information corresponding to each academic achievement, the contribution statistics of the academic achievements of the multiple candidate experts in different fields are performed to obtain the field contribution information corresponding to the multiple candidate experts respectively.

7. The method according to claim 6, characterized in that Determining the subject to which each academic achievement included in the dynamic academic data belongs includes: Performing feature embedding processing on the document of each academic achievement to obtain document embedding features, topic embedding features, and word embedding features; the document embedding features are used to represent the semantic features of the document of each academic achievement, the topic embedding features are used to represent the semantic features of each topic, and the word embedding features are used to represent the semantic features of each word included in the document of each academic achievement; Determining a first mapping relationship between documents and topics based on the distribution of documents of each academic achievement on different topics; Determining a second mapping relationship between the words and the topics based on the distribution of each word on different topics; Based on the first mapping relationship and the second mapping relationship, the subject to which each academic achievement belongs is determined.

8. The method according to claim 7, characterized in that Before performing feature embedding processing on the document based on each academic achievement, the method further includes: Obtaining domain corpus data corresponding to the multiple domains respectively; Performing semantic representation adjustments on the initial models corresponding to the multiple fields respectively on the domain corpus data of the corresponding fields to obtain adjusted initial models corresponding to the multiple fields respectively; The feature embedding process is performed using the adjusted initial models corresponding to the multiple fields.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 8 when executing the computer program. 10 . A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to claim 1 is implemented.

Citation Information

Patent Citations

  • Expert recommendation method and device based on artificial intelligence, terminal and medium

    CN110688405A

  • Expert portrait-oriented information tracking method and device

    CN118779507A

Cited By

  • Intelligent matching method and device, computer equipment and storage medium

    CN121144364A

  • Intelligent matching method and device, computer device and storage medium

    CN121144364B