Scientific research project recommendation method and device based on AI large language model
By comprehensively considering the research boundaries, related fields, and overlapping conflicts of scientific research projects through an AI large language model, and calculating research space indicators, the problem of insufficient accuracy in scientific research project recommendations in existing technologies is solved, and efficient and accurate scientific research project recommendations are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU KEAO INFORMATION TECH CO LTD
- Filing Date
- 2026-02-04
- Publication Date
- 2026-04-28
AI Technical Summary
Existing research project recommendation methods mostly rely on matching based on single-dimensional basic information, which cannot efficiently and accurately recommend high-quality and targeted research projects to researchers, resulting in redundant and inaccurate recommendation results.
Using an AI-based large language model, the research project's research boundary, the public space of related fields, the available space of the research entry point, and the scope of field overlap and conflict between adjacent research projects are obtained. The total research space and the difference in occupancy beyond the boundary are calculated. Taking into account the research potential and overlap and conflict of the projects, projects are recommended in descending order of research space indicators.
It enables efficient and accurate recommendations of high-quality research projects that have both research potential and low conflict risk to researchers, avoiding results that deviate from actual value due to recommendations based on a single dimension, and improving the accuracy and practicality of the recommendations.
Smart Images

Figure CN121660638B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of scientific research project management technology, and more specifically, to a method and apparatus for recommending scientific research projects based on an AI large language model. Background Technology
[0002] In the fields of scientific research project management and support technology, existing technologies already exist that recommend research projects to researchers to provide research assistance. The core objective of this recommendation is to provide researchers with guidance on research directions, helping them quickly select projects that align with their research areas and reducing the time cost of preliminary research. Given the current industry landscape of continuously increasing project numbers and increasingly diverse research directions, such project recommendation applications have gradually become an important component of the scientific research support system.
[0003] Existing research project recommendation methods mostly rely on matching based on single-dimensional basic information, failing to fully consider the core characteristic parameters of research projects and lacking efficient quantitative evaluation and data processing mechanisms. When faced with large-scale research project data, they cannot quickly complete project screening and prioritization, which not only reduces the efficiency of the recommendation process but also easily leads to problems such as redundant recommendation results and insufficient accuracy, making it difficult to meet researchers' actual needs for high-quality and targeted project recommendations. Summary of the Invention
[0004] The purpose of this application is to provide a method and apparatus for recommending scientific research projects based on an AI large language model, which solves the technical problem of recommending scientific research projects to researchers in an efficient and accurate manner, and achieves the technical effect of recommending scientific research projects to researchers in an efficient and accurate manner.
[0005] In a first aspect, embodiments of this application provide a research project recommendation method based on an AI large language model. The method includes: obtaining multiple research projects corresponding to the current research direction through the AI large language model; obtaining the research boundaries, public space of related fields, available space of research entry points, and overlapping conflict range of adjacent research projects for the multiple research projects; obtaining the maximum allowed overlap range of the research projects; determining the total research space of the multiple research projects based on the research boundaries, public space of related fields, and available space of research entry points; determining the difference between the overlapping conflict range of adjacent research projects and the maximum allowed overlap range as the boundary occupancy difference; determining the difference between the total research space of the multiple research projects and the boundary occupancy difference as the research space index of the multiple research projects; and recommending multiple research projects to the user in descending order of research space index.
[0006] In one possible implementation, the total research space of multiple research projects is determined based on their research boundaries, the common space of related fields, and the available space of research entry points. This includes: using an AI large language model to determine the research boundary vector of the research project's research boundary, the related field vector corresponding to the common space of related fields, the available space vector corresponding to the boundary set of the available space of the research entry point, and the field overlap conflict vector corresponding to the field overlap conflict range of adjacent research projects; determining the research boundary L2 norm corresponding to the research boundary vector, the related field L2 norm corresponding to the related field vector, and the available space L2 norm corresponding to the available space vector; and determining the sum of the research boundary L2 norm, the related field L2 norm, and the available space L2 norm of the research project as the total research space of the research project.
[0007] In another possible implementation, the difference between the domain overlap conflict range and the maximum allowable overlap range of adjacent research projects is determined as the over-limit occupancy difference. This includes: using an AI large language model to determine the domain overlap conflict vector corresponding to the domain overlap conflict range of adjacent research projects and the maximum allowable overlap vector corresponding to the maximum allowable overlap range; determining the domain overlap conflict L2 norm corresponding to the domain overlap conflict vector and the maximum allowable overlap L2 norm corresponding to the maximum allowable overlap vector; and determining the difference between the domain overlap conflict L2 norm and the maximum allowable overlap L2 norm of the research project as the over-limit occupancy difference of the research project.
[0008] In another possible implementation, the method further includes: obtaining the resource path coefficient and average resource acquisition time of the research project; determining the product of the resource path coefficient and average resource acquisition time of the research project as the comprehensive resource efficiency value of the research project; determining the comprehensive research space index of the research project based on the research space index and the comprehensive resource efficiency value of the research project; and recommending multiple research projects to the user in descending order of the comprehensive research space index.
[0009] In another possible implementation, the comprehensive research space index of the research project is determined based on the research space index and the comprehensive resource efficiency value of the research project. This includes: obtaining the weights of the research space index and the comprehensive resource efficiency weight; and determining the sum of the product of the research space index and the weight of the research space index, and the product of the comprehensive resource efficiency value and the comprehensive resource efficiency weight, as the comprehensive research space index of the research project.
[0010] In another possible implementation, the method further includes: obtaining the research scope and research trajectory of multiple historical research projects of researchers; scaling the research trajectory of multiple historical research projects within the research scope to obtain a unified dimension research trajectory of multiple historical research projects; performing overlap analysis on the unified dimension research trajectory of multiple historical research projects to determine the union of research trajectories; and determining the degree of overlap between the union of research trajectories and the research boundaries, common space of related fields, and available space of research entry points of multiple research projects, as a weight for research space indicators.
[0011] In another possible implementation, the method further includes: obtaining the research field to which the research project belongs, obtaining the average execution time of the research field to which the research project belongs; determining the mean of the average execution time of all research projects in the research field to which the research project belongs, as the field average execution time; and determining the ratio of the field average execution time to the field basic execution time, as the resource comprehensive efficiency weight of the research project.
[0012] In another possible implementation, the method further includes: acquiring research direction information of researchers and the research field to which the research project belongs; mapping the research direction information of researchers to a unified subject classification system to obtain structured research direction data; mapping the research field to which the research project belongs to a unified subject classification system to obtain structured project field data; determining the subject classification level matching degree based on the structured research direction data and the structured project field data; and multiplying the comprehensive research space index of the research project by the subject classification level matching degree to adjust the comprehensive research space index of the research project.
[0013] In another possible implementation, the subject classification level matching degree is determined based on the structured research direction data and the structured project domain data, including: determining the structured research direction vector corresponding to the structured research direction data; determining the structured project domain vector corresponding to the structured project domain data; and determining the cosine similarity between the structured research direction vector and the structured project domain vector as the subject classification level matching degree.
[0014] Secondly, embodiments of this application provide a research project recommendation device based on an AI large language model, including units for implementing the above-described method.
[0015] The beneficial effects of the embodiments in this application compared with the prior art are:
[0016] This application provides a research project recommendation method based on an AI large language model. The method includes: obtaining multiple research projects corresponding to the current research direction through the AI large language model; obtaining the research boundaries, common space of related fields, available space of research entry points, and overlapping conflict range of adjacent research projects for multiple research projects; obtaining the maximum allowed overlap range of research projects; determining the total research space of multiple research projects based on the research boundaries, common space of related fields, and available space of research entry points; determining the difference between the overlapping conflict range of adjacent research projects and the maximum allowed overlap range as the over-limit occupancy difference; determining the difference between the total research space of multiple research projects and the over-limit occupancy difference as the research space index of multiple research projects; and recommending multiple research projects to the user in descending order of research space index. This application comprehensively considers the research space of the project itself, the expansion space of related fields, the available space of research entry points, and the overlapping conflict exceeding the allowed range. It can comprehensively reflect the effective research space of the project, avoid the recommendation results deviating from the actual value of the project due to focusing on only a single dimension, and help users quickly screen out high-quality projects with both research potential and low conflict risk. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating the first research project recommendation method based on an AI large language model provided in this application embodiment;
[0019] Figure 2 A schematic diagram illustrating the workflow of the first research project recommendation method based on an AI large language model provided in this application embodiment;
[0020] Figure 3 A flowchart illustrating the second research project recommendation method based on an AI large language model provided in this application embodiment;
[0021] Figure 4 A flowchart illustrating the third research project recommendation method based on an AI large language model provided in this application embodiment;
[0022] Figure 5 A flowchart illustrating the fourth research project recommendation method based on an AI large language model provided in this application embodiment;
[0023] Figure 6 A schematic diagram illustrating the workflow of the fourth research project recommendation method based on an AI large language model provided in this application embodiment;
[0024] Figure 7 A flowchart illustrating the fifth research project recommendation method based on an AI large language model provided in this application embodiment;
[0025] Figure 8 A flowchart illustrating the sixth research project recommendation method based on an AI large language model provided in this application embodiment;
[0026] Figure 9 A schematic diagram illustrating the workflow of the sixth research project recommendation method based on an AI large language model provided in this application embodiment;
[0027] Figure 10 A flowchart illustrating the seventh research project recommendation method based on an AI large language model provided in this application embodiment;
[0028] Figure 11 A schematic diagram illustrating the workflow of the seventh research project recommendation method based on an AI large language model provided in this application embodiment;
[0029] Figure 12 A flowchart illustrating the eighth research project recommendation method based on an AI large language model provided in this application embodiment;
[0030] Figure 13 A flowchart illustrating the ninth method for recommending research projects based on an AI large language model, provided in this application embodiment;
[0031] Figure 14 This is a schematic diagram of the logical structure of a research project recommendation device based on an AI large language model, provided in an embodiment of this application. Detailed Implementation
[0032] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0033] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0034] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0035] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0036] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0037] Existing research project recommendation methods mostly rely on matching based on a single dimension of basic information, which makes it difficult to meet researchers' actual needs for high-quality, targeted project recommendations.
[0038] Based on the above reasons, this application provides a research project recommendation method based on an AI large language model. The method includes: obtaining multiple research projects corresponding to the current research direction through the AI large language model; obtaining the research boundaries, public space of related fields, available space of research entry points, and overlapping conflict range of adjacent research projects for multiple research projects; obtaining the maximum allowed overlap range of research projects; determining the total research space of multiple research projects based on the research boundaries, public space of related fields, and available space of research entry points; determining the difference between the overlapping conflict range of adjacent research projects and the maximum allowed overlap range as the over-limit occupancy difference; determining the difference between the total research space of multiple research projects and the over-limit occupancy difference as the research space index of multiple research projects; and recommending multiple research projects to the user in descending order of research space index. This application comprehensively considers the research space of the project itself, the expansion space of related fields, the available space of research entry points, and the overlapping conflict exceeding the allowed range. It can comprehensively reflect the effective research space of the project, avoid the recommendation results deviating from the actual value of the project due to focusing only on a single dimension, and help users quickly screen out high-quality projects with both research potential and low conflict risk.
[0039] In some scenarios, the research project recommendation method based on an AI large language model according to the embodiments of this application can be applied to research assistance, which can improve the recommendation effect of research projects in research assistance for users.
[0040] The following section provides a detailed explanation of a research project recommendation method based on an AI large language model, as provided in the embodiments of this application, using specific examples.
[0041] Figure 1 A flowchart illustrating the first research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 1 As shown in the embodiment of this application, a research project recommendation method based on an AI large language model is provided. The method includes steps S110 to S130, which are described in detail below.
[0042] S110. Utilize an AI-powered large language model to identify multiple research projects corresponding to the current research direction. Obtain the research boundaries of these projects, the common space of related fields, the available space for research entry points, and the overlapping and conflicting areas of adjacent research projects. Determine the maximum permissible overlap range of the research projects.
[0043] Figure 2 A schematic diagram illustrating the workflow of the first research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 2As shown, in this implementation, the semantic understanding capability of the AI large language model can be used to parse the research direction keywords (such as research field, technical direction, application scenario, etc.) input by the user, and match highly relevant projects from scientific research project databases (such as the National Natural Science Foundation of China database, enterprise R&D project database, academic paper preprint database, etc.). The AI large language model can identify the deeper connotation of the research direction and avoid project omissions due to keyword matching bias.
[0044] For example, when a user's current research direction is "the application of artificial intelligence in medical image diagnosis", the AI big language model will extract three core dimensions: "artificial intelligence", "medical image" and "diagnosis", and match multiple research projects such as "Research on lung cancer CT image classification based on Transformer", "Breast cancer screening method based on multimodal medical image fusion" and "Optimization of AI-assisted brain tumor MRI image segmentation algorithm".
[0045] In this implementation, by analyzing the application documents, research plans, or publicly available results of research projects, the research boundaries, common spaces of related fields, available spaces of research entry points, and overlapping and conflicting areas of adjacent research projects can be extracted. The research boundary is the clearly defined research scope of the project (e.g., the boundary of "Research on Lung Cancer CT Image Classification Based on Transformer" is "Application of Transformer Model in Lung Cancer CT Image Classification", excluding other cancer types or modalities such as ultrasound and MRI). The common space of related fields is the intersection of the project with other related fields (e.g., the common space between this project and "Application of Nanomaterials in Medical Imaging Contrast" is "Feature Enhancement Methods for Medical Images"). The available space of research entry points is the innovative direction that has not been fully explored in the project (e.g., the available space of this project is "Optimization of Transformer Attention Mechanism in Small Sample Lung Cancer Images"). The overlapping and conflicting areas of adjacent research projects is the proportion of overlap between the research content of the current project and other projects in the same field (e.g., the overlap between this project and "Research on Lung Cancer CT Image Classification Based on CNN" is "Feature Extraction Methods for Lung Cancer CT Images", accounting for 30%).
[0046] In this implementation, the maximum allowable overlap range is a threshold determined based on research practices, project funding agency requirements, or industry standards, used to limit excessive duplication of research between projects.
[0047] For example, the maximum allowable overlap for basic research projects can be set at 25% (meaning the overlap of research content between two similar projects shall not exceed 25%); the maximum allowable overlap for corporate R&D projects is usually more stringent (e.g., 15%) to avoid wasting internal resources.
[0048] S120. Based on the research boundaries of multiple research projects, the common space of related fields, and the available space of research entry points, determine the total research space of multiple research projects. Determine the difference between the overlapping conflict range of adjacent research projects and the maximum allowable overlap range, as the boundary occupancy difference.
[0049] In this implementation, the total research space is a comprehensive quantification of the project's "research scope, cross-domain expansion potential, and innovation space". The calculation formula is: Total research space = range value of research boundary + expansion value of public space in related fields + unexplored value of available space at the research entry point.
[0050] For example, the research boundary of "Lung Cancer CT Image Classification Research Based on Transformer" covers 10 specific research points (such as the number of layers in the Transformer model, the number of attention heads, and learning rate optimization), and the related domain public space expands to 5 research points (such as weakly supervised learning in medical imaging, cross-modal transfer learning, and noise reduction methods for clinical labels). The available space for research entry points includes 3 unexplored directions (such as the model generalization ability in small sample scenarios, the lesion localization accuracy of the attention mechanism, and the adaptability of multi-center image data). The total research space is 10+5+3=18.
[0051] In this implementation, the over-limit occupancy difference is a quantification of "overlapping degree exceeding the allowable range", reflecting the degree to which overlapping conflicts crowd out the research space of the project. The calculation formula is: over-limit occupancy difference = domain overlap conflict range of adjacent projects - maximum allowable overlap range.
[0052] For example, the maximum allowed overlap range for "Transformer-based lung cancer CT image classification study" is 25%, and the overlap conflict range with adjacent projects is 30%. The difference in the over-limit occupancy is 30%-25%=5% (if the overlap range does not exceed the threshold, the difference is 0).
[0053] S130. Determine the sum of the research space of multiple research projects and the difference in the over-limit occupancy value, as the research space index of multiple research projects. Recommend multiple research projects to the user in descending order of research space index.
[0054] In this implementation, the research space index is the final quantitative result of the project's "effective research space". By deducting the over-limit occupancy difference from the total research space, the balance between "potential value and conflict loss" is achieved. The calculation formula is: Research space index = Total research space - Over-limit occupancy difference.
[0055] For example, the total research space of "Research on Lung Cancer CT Image Classification Based on Transformer" is 18, and the over-limit occupancy difference is 5, so its research space index is 18-5=13; if the over-limit occupancy difference of a certain project is 0 (not exceeding the maximum allowable range), then the research space index is equal to the total research space.
[0056] In this implementation, all matched research projects can be sorted according to the research space index. The larger the index, the more effective research space the project has and the lower the risk of conflict, and the more likely it is to be recommended to the user.
[0057] For example, if the research space indicators for three research projects are: Project A (13), Project B (11, total research space 15, excess space difference 4), and Project C (9, total research space 12, excess space difference 3), then the recommended order is Project A → Project B → Project C. Users can directly filter projects with both research potential and low conflict risk based on the sorting results.
[0058] This approach takes into account both the potential research space of a project and the negative impact of overlap and conflict, ensuring that the recommended projects have sufficient research space without excessive overlap and conflict. This effectively improves the accuracy and practicality of research project recommendations, helping users quickly select high-quality projects that have both research potential and low conflict risk.
[0059] This approach comprehensively considers the research space of the project itself, the expansion space of related fields, the available space of research entry points, and overlapping conflicts that exceed the allowable range. It can fully reflect the effective research space of the project and avoid the recommendation results deviating from the actual value of the project due to focusing on only a single dimension. The quantitative processing of overlapping conflicts transforms the negative impact of overlapping conflicts into a specific calculable value, which can accurately measure the adverse effects of overlapping conflicts on the project. This avoids recommending research projects with excessive overlap and ensures that the recommended projects have reasonable research boundaries and development space.
[0060] Figure 3 A flowchart illustrating the second research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 3 As shown, in some implementations, in S120 above, the total research space of multiple research projects is determined based on the research boundaries of multiple research projects, the public space of related fields, and the available space of research entry points, including S121 to S122. S121 to S122 will be explained in detail below.
[0061] S121. Using the AI large language model, determine the research boundary vector of the research project, the related domain vector corresponding to the public space of the related domain, the available space vector corresponding to the boundary set of the available space of the research entry point, and the domain overlap and conflict vector corresponding to the domain overlap and conflict range of adjacent research projects.
[0062] In this implementation, the semantic encoding capabilities of AI large language models can be used to transform the abstract research boundaries, related domain public spaces, available space for research entry points, and overlapping and conflicting ranges of adjacent projects in scientific research projects into structured high-dimensional vectors.
[0063] In this implementation, the research boundary of a research project (such as "Research on lung cancer CT image diagnosis method based on deep learning, which only focuses on the classification of benign and malignant lung nodules and does not involve other cancer types") is input into the AI large language model. The model will first parse the core semantic elements ("deep learning", "lung cancer CT image", "lung nodule classification", "excluding other cancers"), and then, based on the pre-trained semantic library in the fields of medical imaging and machine learning, map these elements into research boundary vectors in high-dimensional space (such as real number vectors with a dimension of 768).
[0064] In this implementation, for the common space of related domains (such as "the common research scope of 'disease classification based on image features' in the intersection of deep learning and medical imaging diagnosis"), the model will extract the core features of the common space ("intersection domain", "image features", "disease classification") and generate related domain vectors.
[0065] In this implementation, for the boundary set of the available space of research entry points (such as "in lung cancer CT image diagnosis research, available entry points include 'model training under small sample data' and 'multimodal image fusion feature extraction'"), the model will integrate the semantic associations of each boundary to generate an available space vector.
[0066] In this implementation, for the overlapping and conflicting domains of adjacent research projects (such as "the overlapping part of this project and adjacent projects in 'deep learning model applied to lung nodule classification'"), the model will identify the overlapping semantic range and encode it as a domain overlap and conflict vector.
[0067] For example, if the research boundary of a research project is "a study on Chinese news text classification based on Transformer, focusing only on political news and excluding entertainment or sports," after inputting this boundary into an AI large language model, the model will parse out elements such as "Transformer," "Chinese news," "political classification," and "excluding other domains," and then generate a research boundary vector with a dimension of 768 based on a pre-trained semantic database of the natural language processing domain. If the common space of related domains is "the common range of 'pre-trained model applied to text classification' in the intersection of Transformer and Chinese text processing," the model will extract features such as "intersection domain," "pre-trained model," and "text classification" to generate related domain vectors. If the boundary set of available research entry points is "in Chinese news classification research, available entry points include 'long-tail category sample optimization' and 'cross-domain transfer learning'," the model will integrate the semantic associations of these boundaries to generate a usable space vector. If the scope of overlapping conflicts between adjacent projects is "the overlap between this project and adjacent projects in 'Transformer applied to political news classification'," the model will encode it as a domain overlap conflict vector.
[0068] S122. Determine the L2 norm of the research boundary vector, the L2 norm of the related domain vector, and the L2 norm of the available space vector. The sum of the L2 norms of the research boundary, related domain, and available space for the research project is taken as the total research space of the research project.
[0069] In this implementation, the "semantic range size" of semantic vectors in each dimension can be converted into a quantifiable value by calculating the L2 norm of the vector (i.e., the "length" of the vector in the high-dimensional space), and then the summation is used to obtain the total study space.
[0070] For example, if the research boundary vector of a research project is a vector of dimension 768 with a sum of squares of 100, then the L2 norm of the research boundary is 10; the sum of squares of the related domain vectors is 80, and the L2 norm is approximately 8.94; the sum of squares of the available space vectors is 60, and the L2 norm is approximately 7.75; the sum of the three (10 + 8.94 + 7.75 ≈ 26.69) is the total research space of the project.
[0071] For example, if Project A has a research boundary L2 norm of 12, an associated domain L2 norm of 9, and a usable space L2 norm of 7, its total research space is 28; Project B has a research boundary L2 norm of 10, an associated domain L2 norm of 10, and a usable space L2 norm of 8, also totaling 28. This indicates that the two projects are comparable in terms of the comprehensive research space in terms of "breadth of research boundary, breadth of associated domain, and usable scope of entry point," and the differences can be further distinguished through subsequent steps.
[0072] This implementation method uses an AI large language model to transform the research boundary of a research project into a research boundary vector, the public space of related fields into a related field vector, and the boundary set of the available space of the research entry point into a available space vector. Then, it calculates the research boundary L2 norm corresponding to the research boundary vector, the related field L2 norm corresponding to the related field vector, and the available space L2 norm corresponding to the available space vector. Finally, it adds these three L2 norms to obtain the total research space of the research project. This transforms the originally abstract research space-related indicators into quantifiable vector norms, avoiding the uncertainty of traditional qualitative or fuzzy quantitative methods. This makes the calculation of the total research space more objective and accurate, providing a more reliable foundation for the determination of subsequent research space indicators.
[0073] This implementation extends the capabilities of AI large language models from information collection to structured information encoding. It ensures that subsequent L2 norm calculations and the determination of the research space summation are based on scientific semantics, avoiding calculation errors caused by improper information representation and improving the accuracy of the research space summation and its relevance to the actual situation of scientific research projects.
[0074] Figure 4 A flowchart illustrating the third research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 4 As shown, in some implementations, in the above-mentioned S120, the difference between the overlapping conflict range of adjacent scientific research projects and the maximum allowable overlapping range is determined as the boundary occupancy difference, including S123 to S124. S123 to S124 will be explained in detail below.
[0075] S123. Using an AI large language model, determine the domain overlap conflict vector corresponding to the domain overlap conflict range of adjacent research projects and the maximum allowable overlap vector corresponding to the maximum allowable overlap range.
[0076] In this implementation, the textual descriptions of the overlapping conflict range and the maximum allowable overlap range of adjacent scientific research projects can be input into the AI large language model. Based on its understanding of the semantics of the scientific research field, the model extracts key dimensions (such as technology type, application scenario, research task, etc.) in the range and quantifies the features of each dimension into vector elements, thereby generating the corresponding field overlapping conflict vector and the maximum allowable overlap vector. By utilizing the semantic encoding capability of the AI large language model, the abstract range description is transformed into a structured vector, ensuring that the vector accurately reflects the substantive content of the range.
[0077] For example, if adjacent research projects are "Application of machine learning in medical image diagnosis" and "Research on deep learning in medical image segmentation", and the scope of their domain overlap and conflict is "Technical application of machine learning / deep learning in medical image processing", the AI big language model will extract three key dimensions: "technology type, application scenario, and task type", quantify the matching degree of each dimension into a value between 0 and 1, and generate a domain overlap and conflict vector such as [0.8, 0.9, 0.7]. If the maximum allowable overlap is "medical image application of different tasks under the same technology type", the model will also extract the above dimensions and quantify to generate a maximum allowable overlap vector such as [0.6, 0.7, 0.5].
[0078] S124. Determine the L2 norm of the domain overlap conflict vector and the L2 norm of the maximum allowed overlap vector. Determine the difference between the L2 norm of the domain overlap conflict and the L2 norm of the maximum allowed overlap for each research project, using this as the overstepping occupancy difference for that research project.
[0079] In this implementation, the L2 norm can be calculated for the generated domain overlap conflict vector and the maximum allowed overlap vector. The L2 norm is the square root of the sum of the squares of the elements of the vector, which is used to quantify the "size" of the range represented by the vector. After calculating the two L2 norms, the difference between the two can be taken as the cross-border occupancy difference. This difference can intuitively reflect the degree to which the actual overlap range exceeds the allowed range.
[0080] For example, if the domain overlap conflict vector is [0.8, 0.9, 0.7], its L2 norm is 1.39; if the maximum allowed overlap vector is [0.6, 0.7, 0.5], its L2 norm is 1.05; then the difference in boundary crossing is 1.39-1.05=0.34, which indicates that the actual overlap range exceeds the allowed range by 0.34.
[0081] This implementation method quantifies the abstract range into a computable value using vectors and the L2 norm, preserving the spatial characteristics of the range and avoiding information loss in traditional non-vector methods. This makes the quantification of domain overlap conflicts and maximum allowed overlap more accurate, improving the accuracy of the calculation results for the boundary occupancy difference. This uniformity ensures the consistency of calculations for different indicators, reduces errors caused by different calculation methods, and thus keeps the calculation logic of the boundary occupancy difference and the sum of the research space consistent, improving the standardization of the calculation of indicators in the entire research space and avoiding indicator deviations caused by differences in calculation methods.
[0082] This implementation method generates domain overlap and conflict vectors and maximum allowed overlap vectors using an AI large language model. The AI large language model can understand the domain semantics and scope boundaries of scientific research projects, accurately convert the scope of textual descriptions into vectors, and use the semantic understanding capabilities of the AI large language model to process professional content in the scientific research field. This ensures that the vectors accurately reflect the substantive content of the scope, improves the semantic relevance and accuracy of the vectors, provides a reliable basis for calculating the difference in boundary crossing, and avoids calculation errors caused by inaccurate vectors.
[0083] Figure 5 A flowchart illustrating the fourth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 5 As shown, in some implementations, the above method also includes S210 to S220, which will be described in detail below.
[0084] S210. Obtain the resource path coefficient and average resource acquisition time of the research project. Determine the product of the resource path coefficient and average resource acquisition time of the research project as the overall resource efficiency value of the research project.
[0085] Figure 6 A schematic diagram illustrating the workflow of the fourth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 6 As shown, in this implementation, the resource path coefficient and average resource acquisition time of key resources required for scientific research projects can be collected. The resource path coefficient is used to measure the convenience of the resource acquisition path (such as supply chain stability, process complexity, supplier reliability, etc.), and the larger the value, the smoother the path. The average resource acquisition time is used to measure the time cost of acquiring key resources (such as the average time from initiating an application to the arrival of resources), and the smaller the value, the faster the acquisition speed. These parameters can be determined by surveying the supply chain of project resources, historical acquisition data, etc., and provide a basis for subsequent calculation of the comprehensive resource efficiency value.
[0086] For example, for the "New Energy Battery Material R&D" project, lithium ore raw materials need to be obtained: if there are 3 stable supply chains (supplier on-time delivery rate ≥ 95%), and the procurement process only requires 3 steps: "submit requirements - sign contract - deliver goods", the resource path coefficient can be set to 0.85; if the average time to obtain lithium ore for the past 5 similar projects is 15 days, then the average time to obtain resources is 15 days.
[0087] For example, for the "AI-based medical image diagnosis" project, which requires acquiring medical data: if it can directly connect to the electronic medical record systems of three top-tier hospitals, and the data authorization process only requires two steps, "submitting for ethical approval - system integration", the resource path coefficient can be set to 0.9; if the average time for acquiring medical data for similar projects in the past is 10 days, then the average time for acquiring resources is 10.
[0088] In this implementation, the resource efficiency value is synthesized by multiplying the resource path coefficient by the average resource acquisition time, thus combining "convenience" and "time cost." The higher the resource path coefficient (more convenient the path) and the shorter the average resource acquisition time (lower time cost), the smaller the product, indicating higher resource acquisition efficiency. This calculation method integrates two dimensions of resource efficiency factors into comparable values, facilitating subsequent integration with spatial indicators in research.
[0089] For example, the resource path coefficient of the "New Energy Battery Material R&D" project is 0.85, and the average time to acquire resources is 15 days. The product of the two is 0.85 × 15 = 12.75, which is the comprehensive resource efficiency value of the project. The resource path coefficient of the "Artificial Intelligence Medical Image Diagnosis" project is 0.9, and the average time to acquire resources is 10 days. The product of the two is 0.9 × 10 = 9. The latter value is smaller, indicating that the resource acquisition efficiency of the medical imaging project is higher.
[0090] S220. Based on the research space indicators and resource comprehensive efficiency values of the research projects, determine the comprehensive research space indicators for each project. Recommend multiple research projects to the user in descending order of their comprehensive research space indicators.
[0091] In this implementation method, the comprehensive research space index is a combination of the research space index (reflecting the research potential and conflict risk of the project) and the comprehensive resource efficiency value (reflecting the efficiency of resource acquisition). The comprehensive research space index measures both "research value" and "execution feasibility". The larger the index, the more research space the project has, the more efficient it can acquire resources, and the lower the execution risk.
[0092] For example, Project A has a research space index of 90 (high research potential, few conflicts) and a resource comprehensive efficiency value of 9 (high resource acquisition efficiency), so its research space comprehensive index can be 90 ÷ 9 = 10; Project B has a research space index of 85 (slightly lower research potential) and a resource comprehensive efficiency value of 8 (even higher resource acquisition efficiency), so its research space comprehensive index can be 85 ÷ 8 = 10.625. In this case, Project B has a larger comprehensive index, indicating that its combination of "research potential and resource efficiency" is better.
[0093] In this implementation, the comprehensive index of research space is used as the basis for the final recommendation ranking. Projects are presented to users in descending order of index. This approach unifies "research potential" and "resource efficiency" into a single dimension, avoiding the recommendation of "potential-only" projects that have "large research space but are difficult to obtain resources". This ensures that the recommendation results are both academically valuable and practically feasible.
[0094] For example, there are three projects to be recommended: Project X: Research space index 90, resource comprehensive efficiency value 9 → comprehensive index 10; Project Y: Research space index 85, resource comprehensive efficiency value 8 → comprehensive index 10.625; Project Z: Research space index 95, resource comprehensive efficiency value 12 → comprehensive index ≈ 7.92. Sorted from highest to lowest comprehensive index, the recommended order is Project Y → Project X → Project Z.
[0095] This implementation integrates resource utilization efficiency into the recommendation logic, ensuring that the recommendation results consider not only the size and conflict of the research space but also the actual efficiency of resource acquisition. This avoids recommending projects that seem to have a large research space but are difficult to acquire resources efficiently, thus improving the practicality of the recommendations. Compared to using only research space indicators, the comprehensive indicators are more comprehensive, and the recommended projects have sufficient research potential and reasonable domain boundaries while also making efficient use of resources, thereby improving the scientific rigor of the recommendations.
[0096] Figure 7 A flowchart illustrating the fifth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 7 As shown, in some implementations, in S210 above, the comprehensive research space index of the research project is determined based on the research space index and resource comprehensive efficiency value of the research project, including S211 to S212. S211 to S212 will be explained in detail below.
[0097] S211. Obtain the weights of research space indicators and the weights of comprehensive resource efficiency.
[0098] In this implementation, the weights of research space indicators and resource comprehensive efficiency can be obtained. These weights can be determined based on the application scenarios recommended for scientific research projects or the needs and preferences of users. For example, when the recommended target is an academic team focusing on basic research, the weight of research space indicators can be set to a higher value to highlight the research potential of the project; when the recommended target is a corporate R&D department that emphasizes implementation efficiency, the weight of resource comprehensive efficiency can be appropriately increased to focus on the convenience of resource acquisition.
[0099] In this implementation, the weight values that meet the needs of most users can be adjusted by analyzing the proportion of users' attention to research space and resource efficiency in historical recommendation data, thus ensuring the rationality and practicality of the weight allocation.
[0100] For example, the weights of research space indicators and resource comprehensive efficiency can be set according to different needs and scenarios. For instance, in the recommendation of basic research projects for universities, the weight of research space indicators is set to 0.7 and the weight of resource comprehensive efficiency is set to 0.3, reflecting the emphasis on research potential. In the recommendation of applied research and development projects for enterprises, the weight of research space indicators is set to 0.4 and the weight of resource comprehensive efficiency is set to 0.6, highlighting the importance of resource acquisition efficiency.
[0101] S212. The sum of the product of the research spatial index and the weight of the research spatial index, and the product of the comprehensive resource efficiency value and the weight of the comprehensive resource efficiency value, is determined as the comprehensive research spatial index of the research project.
[0102] In this implementation, when determining the comprehensive research space index, the research space index of the research project is multiplied by its corresponding weight to obtain the weighted contribution value of the research space dimension; then, the comprehensive resource efficiency value is multiplied by its corresponding weight to obtain the weighted contribution value of the resource efficiency dimension; finally, the two weighted contribution values are added together to obtain the comprehensive research space index that comprehensively reflects the research potential and resource acquisition efficiency of the project. This weighted summation method clarifies the relative importance of the two dimensions through weights, avoids the one-sided influence of a single dimension, and makes the comprehensive index more in line with the needs and priorities in actual applications.
[0103] For example, suppose a research project has a research space index of 85 and a weight of 0.6; a resource comprehensive efficiency value of 75 and a resource comprehensive efficiency weight of 0.4. Then the weighted contribution value of the research space index is 85 × 0.6 = 51, and the weighted contribution value of the resource comprehensive efficiency value is 75 × 0.4 = 30. The sum of the two, 51 + 30 = 81, is the research space comprehensive index of the project.
[0104] This implementation method clarifies the relative importance of the two indicators through weight allocation, integrating the originally independent research space indicator and resource comprehensive efficiency value into a comprehensive indicator. This allows for a more accurate quantification of the combined impact of the two indicators on research projects, avoiding the one-sidedness of a single indicator and making the comprehensive indicator more aligned with the weight preferences in actual needs.
[0105] This implementation method combines the different levels of importance that users place on the two factors in different scenarios to achieve a balanced consideration of multiple factors. It avoids over-biasing on one factor, which would result in recommended projects having either a large research space but difficult resource acquisition, or easily accessible resources but limited research space, thereby improving the applicability of the recommendation results to different needs.
[0106] Figure 8 A flowchart illustrating the sixth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 8As shown, in some implementations, the above method also includes S310 to S320, which will be described in detail below.
[0107] S310. Obtain the research scope and trajectory of multiple historical research projects of the researcher. Scale the research trajectories of multiple historical research projects within their research scope to obtain a unified dimension of research trajectory for each project. Perform overlap analysis on the unified dimension research trajectories of multiple historical research projects to determine the union of the research trajectories.
[0108] Figure 9 A schematic diagram illustrating the workflow of the sixth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 9 As shown, this implementation method can collect the research scope and trajectory of multiple historical research projects conducted by researchers. The research scope of each historical project refers to the boundary of the specific research field involved in the project; the research trajectory refers to the specific progress path of the project within its research scope from initial conception to final completion. This information provides a foundation for subsequent analysis of researchers' research strengths and coverage.
[0109] It should be noted that the research scope and trajectory of a researcher's multiple historical research projects can be illustrated with examples from specific scenarios. For instance, a researcher in the field of computer science may have historical research projects including "Optimization of Image Classification Algorithms Based on Deep Learning" and "Research on Model Compression Technology for Edge Computing." The research scope of "Optimization of Image Classification Algorithms Based on Deep Learning" covers the application and optimization methods of deep learning algorithms in image classification, and its research trajectory progresses from structural optimization of classic CNN models (such as ResNet) to the introduction of attention mechanisms, and then to the design of lightweight models. The research scope of "Research on Model Compression Technology for Edge Computing" covers model compression methods and deployment optimization on edge devices, and its research trajectory progresses from improvements to pruning algorithms to the integration of quantization techniques, and then to hardware-algorithm co-optimization.
[0110] In this implementation, the research trajectory can be scaled and adjusted within the research scope of each historical research project, converting the research trajectories of historical projects with different research scopes into a representation with the same dimension. This is done to solve the problem of inconsistent trajectory dimensions caused by differences in the research scope of historical projects, and to facilitate the subsequent overlapping analysis of the research trajectories of multiple historical projects.
[0111] For example, a researcher has two historical projects: Project A's research scope is "sentiment analysis in natural language processing" (covering 10 research sub-directions), and its research trajectory covers 6 of these sub-directions; Project B's research scope is "attention mechanism optimization in machine translation" (covering 5 research sub-directions), and its research trajectory covers 3 of these sub-directions. The research trajectory of Project A is scaled by increasing the dimension of the research direction upwards, thus scaling the 6 covered sub-directions into 3 units of a unified dimension; the research trajectory of Project B remains unchanged (already corresponding to 3 units of 5 sub-directions), ultimately resulting in a unified dimension research trajectory for both projects.
[0112] For example, when converting the research trajectories of historical projects of different research scopes into a representation of the same dimension, a large language model can be used to understand and transform the research trajectories of historical projects of different research scopes into a representation of the same dimension.
[0113] In this implementation, the research trajectories of multiple historical research projects with a unified dimension can be compared to find the overlapping parts and unique parts of different trajectories. Then, the coverage of all trajectories is merged to obtain a set containing the coverage area of all historical research trajectories, that is, the union of research trajectories. The union can accurately reflect the coverage of the historical research scope of researchers.
[0114] For example, in two unified-dimensional historical research trajectories of a researcher, trajectory 1 covers dimension units 1-3 (corresponding to the scaled trajectory of project A), and trajectory 2 covers dimension units 2-4 (corresponding to the scaled trajectory of project B). When performing an overlap analysis on the two, the overlapping part is dimension units 2-3, and the unique part is 1 of trajectory 1 and 4 of trajectory 2. Therefore, the union of the research trajectories is dimension units 1-4.
[0115] S320. The overlap of the union of research trajectories and the research boundaries of multiple research projects, the public space of related fields, and the available space of research entry points is determined as the weight of the research space indicator.
[0116] In this implementation, the degree of overlap (i.e., coincidence) between the union of the researchers' historical research trajectories and the three parts of the current research project's research boundaries, the common space of related fields, and the available space of the research entry point can be calculated. This degree of overlap is used as the weight of the research space index. This approach can reflect the researchers' familiarity with the current project's research space and make the weight more in line with their research experience.
[0117] For example, the union of a researcher's historical research trajectory is "algorithm optimization and edge deployment of deep learning in computer vision," the research boundary of a current research project is "image segmentation algorithm optimization based on Transformer," the common space of related fields is "application of Transformer model in computer vision," and the available space for research entry point is "structural design of lightweight Transformer model." Calculating the overlap between the union and these three parts, we get 80% overall. Therefore, the weight of the research space index for this project is 0.8.
[0118] For example, when determining the overlap between the union of research trajectories and the research boundaries of multiple research projects, the common space of related fields, and the available space of research entry points, the union of research trajectories and the research boundaries of multiple research projects, the common space of related fields, and the available space of research entry points can first be converted into vector form using a large language model. Then, the similarity between the vector corresponding to the union of research trajectories and the vector corresponding to the overlap between the research boundaries of multiple research projects, the common space of related fields, and the available space of research entry points can be further determined.
[0119] This approach determines the weights of spatial indicators based on the researchers' historical research background, replacing general weights. This makes the comprehensive indicators of spatial research more closely aligned with the researchers' research experience, thereby improving the matching degree between recommended research projects and the researchers' research background.
[0120] This implementation method scales the research trajectories of multiple historical research projects within their respective research scopes to obtain research trajectories with a unified dimension. This solves the problem of inconsistent trajectory dimensions caused by different research scopes of different historical projects. Then, through overlap analysis, the union of historical research trajectories is obtained, which accurately reflects the research scope coverage of researchers. The degree of overlap between this union and the research boundary of the research project, the common space of related fields, and the available space of the research entry point is calculated as the weight of the research space index. This ensures that the weight can reflect the researcher's familiarity with the current project's research space, making the comprehensive research space index better reflect the researcher's adaptability to the project and improving the applicability of recommended projects.
[0121] Figure 10 A flowchart illustrating the seventh research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 10 As shown, in some implementations, the above method also includes S330 to S340, which will be described in detail below.
[0122] S330. Obtain the research field to which the research project belongs, and obtain the average execution time of the research field to which the research project belongs. Determine the average execution time of all research projects in the research field to which the research project belongs, as the average total execution time of the field.
[0123] Figure 11 A schematic diagram illustrating the workflow of the seventh research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 11 As shown, in this implementation, the research field to which the research project belongs can be obtained first. This is the basis for calculating the comprehensive resource efficiency weight in combination with the overall characteristics of the field. At the same time, the average execution time of the research field can be obtained. This average execution time reflects the general time cost level of the research project execution in the field, providing basic data support for the subsequent quantification of the overall time consumption characteristics of the field.
[0124] For example, the research field to which a research project belongs can be determined based on the core research content and subject classification of the project. For instance, the research field of "semantic analysis of medical text based on large language model" belongs to "interdisciplinary field of natural language processing and medical information", and the research field of "stability study of novel solar cells based on perovskite" belongs to "photovoltaic materials and devices".
[0125] For example, the average execution time of a research project in a particular research field can be obtained by statistically analyzing the execution cycles of completed projects in that field. For instance, in the "interdisciplinary field of natural language processing and medical information," the execution time of 25 projects completed in the past four years ranged from 14 months to 26 months, with an average execution time of 20 months. In the "photovoltaic materials and devices" field, the average execution time of 30 projects completed in the past three years was 18 months.
[0126] In this implementation, the average execution time of all research projects within the research field to which the research project belongs can be determined as the field average time. The field average time integrates the general time consumption of projects in various sub-directions or types within the field, and quantifies the overall time cost scale of the field.
[0127] For example, if the "interdisciplinary field of natural language processing and medical information" includes three sub-directions of research projects: "structured parsing of medical records", "generation of medical guide texts" and "recognition of clinical dialogue intent", and the average execution time for each direction is 19 months, 20 months and 21 months respectively, then adding the three together gives (19+20+21) / 3=20 months, which is the average total time for this field.
[0128] S340. Determine the ratio of the average time spent in the field to the basic time spent in the field, and use it as the weight for the overall resource efficiency of scientific research projects.
[0129] In this implementation, the ratio of the average time spent in the field to the basic time spent in the field can be calculated. This ratio is used as the resource comprehensive efficiency weight of the scientific research project. This weight integrates the overall time spent characteristics of the field and is used to adjust the contribution of the resource comprehensive efficiency value in the comprehensive index of the research space.
[0130] It should be noted that the domain-based time is the baseline time for project execution within that research domain. It is usually set based on the domain's conventional research processes, resource allocation standards, and project complexity thresholds. For example, the domain-based time for the "interdisciplinary field of natural language processing and medical information" can be set at 20 months (based on the conventional process time for data annotation, model training, and clinical validation in this field); the domain-based time for the "photovoltaic materials and devices field" can be set at 15 months (based on the conventional cycle of material synthesis, device fabrication, and performance testing).
[0131] This implementation method ensures that the comprehensive index of research space not only considers the comprehensive resource efficiency value of the research project itself, but also incorporates the overall time consumption characteristics of the research field. This makes the comprehensive index more comprehensively reflect the resource efficiency and field background of the project, avoids the bias caused by relying solely on the project's own data, and improves the accuracy of the recommendation results.
[0132] This implementation method, which determines weights based on actual time consumption data in the research field, avoids the bias of subjective assignment, making the comprehensive resource efficiency weights more consistent with the time consumption characteristics of the field. This improves the accuracy of subsequent calculations of comprehensive research space indicators, and makes recommended research projects more in line with the actual resource efficiency of the field.
[0133] Figure 12 A flowchart illustrating the eighth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 12 As shown, in some implementations, the above method also includes S410 to S420, which will be described in detail below.
[0134] S410. Obtain information on the research directions of researchers and the research fields to which the research projects belong. Map the researchers' research direction information to a unified subject classification system to obtain structured research direction data. Map the research fields to which the research projects belong to a unified subject classification system to obtain structured project field data.
[0135] In this implementation method, information on the research direction of researchers can be collected, such as the topics of their published papers, the research directions of the projects they participate in, and the research fields of the projects they apply for. This information can reflect the core research directions of researchers. At the same time, the research fields to which the research projects belong can be collected, such as the field classification specified in the project application, the research field description in the project completion report, and the technical fields corresponding to the project results. This information can clearly define the field affiliation of the research projects. This raw information will serve as the basic data for subsequent mapping to a unified classification system.
[0136] In this implementation, the hierarchical structure of a unified subject classification system (such as a tree-like classification from first-level disciplines to second-, third-, and fourth-level disciplines) can be used to transform researchers' research direction information into structured data. Specifically, the core keywords or themes in the researchers' research directions can be extracted first, and then these keywords can be matched one by one with the categories in the unified subject classification system to finally form structured research direction data containing multi-level subject categories. This mapping process can transform unstructured textual descriptions into structurally consistent and directly comparable classification data.
[0137] It should be noted that the unified subject classification system here can be an industry-standard system (such as the national standard "Subject Classification and Codes" GB / T 13745-2009), or it can be an internal standardized classification system developed by research institutions based on their own research fields. Its core function is to eliminate the differences between research directions information from different sources and with different expressions, and to achieve the structuring and standardization of data.
[0138] For example, if a researcher's research direction is "image semantic segmentation based on deep learning", the core keywords such as "deep learning" and "image semantic segmentation" are extracted first. Then, the unified subject classification system is queried to find the corresponding first-level discipline "computer science and technology", second-level discipline "computer application technology", third-level discipline "computer vision", and fourth-level discipline "image semantic segmentation". Finally, these hierarchical categories are combined to form structured research direction data of "computer science and technology - computer application technology - computer vision - image semantic segmentation", which clearly shows the specific position of the research direction in the unified classification system.
[0139] In this implementation, a method consistent with the processing of researchers' research direction information can be adopted to map the research field of research projects to a unified subject classification system, forming structured project field data. The core content (such as technical direction, application scenario, etc.) in the research field of research projects can be extracted first, and then matched with the categories in the unified subject classification system. Finally, a multi-level subject category combination consistent with the structured research direction data structure is formed. This consistent structural design can provide a unified comparison basis for subsequent subject classification matching analysis.
[0140] For example, if the research field of a research project is "the application of blockchain technology in supply chain finance", the core keywords such as "blockchain technology" and "supply chain finance" are extracted first. Then, the unified subject classification system is queried to find the first-level discipline "Management", the second-level discipline "Business Administration", the third-level discipline "Supply Chain Management", and the fourth-level discipline "Supply Chain Finance" corresponding to the main category, as well as the first-level discipline "Computer Science and Technology", the second-level discipline "Computer Application Technology", and the third-level discipline "Blockchain Technology" corresponding to the cross-category. Finally, based on the main research direction of the project, the main category is determined as "Management-Business Administration-Supply Chain Management-Supply Chain Finance" and the cross-category is determined as "Computer Science and Technology-Computer Application Technology-Blockchain Technology", forming structured project field data that includes the main category and the cross-category, comprehensively reflecting the field coverage of the project.
[0141] S420. Based on the structured research direction data and structured project area data, determine the subject classification level matching degree. Multiply the research space comprehensive index of the research project by the subject classification level matching degree to adjust the research space comprehensive index of the research project.
[0142] In this implementation, the matching degree between structured research direction data and structured project domain data can be quantitatively calculated based on the hierarchical depth of the unified subject classification system. For example, if the unified subject classification system is divided into four levels—level one, level two, level three, and level four—each level's matching status corresponds to a different weight (generally, the finer the level, the higher the weight). By weighted summing the matching results of each level, the subject classification hierarchical matching degree can be obtained. This hierarchical matching method can accurately reflect the deep consistency between the two in subject classification.
[0143] It should be noted that the calculation logic of the subject classification level matching degree is based on the principle that "the finer the classification level, the higher the matching difficulty and the stronger the correlation".
[0144] For example, the matching weight of a fourth-level subject may be higher than that of a third-level subject, third-level higher than second-level, and second-level higher than first-level. This way, the calculated matching degree can more accurately reflect the deep relationship between the two in terms of subject classification.
[0145] For example, if the structured research direction data is "Computer Science and Technology - Computer Application Technology - Computer Vision - Image Semantic Segmentation" and the structured project domain data is "Computer Science and Technology - Computer Application Technology - Computer Vision - Video Content Analysis", the two are completely matched in the first, second, and third level disciplines, and partially matched in the fourth level discipline (both belong to the sub-domain of computer vision). Assuming that the weight of each level is the same, and the partial matching degree of the fourth level discipline is set to 0.8, then the overall discipline classification level matching degree is (1+1+1+0.8) / 4=0.95, which can better reflect the discipline relevance of the two.
[0146] In this implementation method, the original comprehensive research space index of the research project (which has taken into account factors such as research space, resource efficiency, and weight) can be multiplied with the matching degree of the discipline classification level to obtain the adjusted comprehensive research space index.
[0147] For example, if a research project's original comprehensive index for research space is 80 and its subject classification level matching degree is 0.95, then the adjusted index will be 80 × 0.95 = 76; if another project's original index is 80 and its matching degree is 0.5, then the adjusted index will be 40. The adjusted index directly reflects the impact of the matching degree between the researcher's subject background and the project field on the project. The higher the matching degree, the higher the adjusted index, and the higher the project's recommendation priority.
[0148] It should be noted that the core principle of this adjustment method is to integrate the matching degree between the researcher's disciplinary expertise and the project field into the comprehensive index, so as to avoid recommending projects that are "mismatched in discipline but have good objective attributes". This ensures that the recommended projects not only have good research space and resource efficiency, but also highly match the researcher's disciplinary background, thereby improving the success rate of project execution and the work efficiency of researchers.
[0149] For example, suppose there are two research projects: Project A originally had a comprehensive research space index of 90 and a subject classification level matching degree of 0.9; Project B originally had an index of 85 and a matching degree of 1.0. After adjustment, Project A's comprehensive index is 90 × 0.9 = 81, and Project B's is 85 × 1.0 = 85. At this point, Project B has a higher recommendation priority. Although its original index is slightly lower than Project A's, its subject classification matching degree is higher, better aligning with the researcher's subject expertise, and its implementation is smoother.
[0150] This implementation eliminates the problem of inconsistent classification standards between researchers' research directions and project fields through a unified classification system, providing a structured data foundation for direct comparison between the two. This avoids errors in subsequent matching calculations due to differences in classification systems and provides a consistent data source for accurately quantifying the correlation between the two. The calculation can accurately quantify the correlation between researchers' research directions and project fields in terms of disciplinary structure. Compared with the previous approach that only considered the overlap of research scope, it delves deeper into the underlying logic of disciplinary classification, effectively improving the accuracy of correlation assessment.
[0151] Through this implementation method, the adjusted comprehensive index of research space is more aligned with the disciplinary expertise of researchers. The recommended research projects not only possess good research space and resource efficiency attributes, but also highly match the researchers' research directions in terms of disciplinary classification, significantly improving the accuracy and practicality of the recommendation results.
[0152] Figure 13 A flowchart illustrating the ninth research project recommendation method based on an AI large language model provided in this application embodiment is shown below. Figure 13 As shown, in some implementations, in S420 above, the subject classification level matching degree is determined based on the structured research direction data and the structured project field data, including S421 to S422. S421 to S422 will be explained in detail below.
[0153] S421. Determine the structured research direction vector corresponding to the structured research direction data. Determine the structured project domain vector corresponding to the structured project domain data.
[0154] In this implementation, a large language model can be used to convert structured research direction data mapped to a unified subject classification system into structured research direction vectors using a consistent vector encoding method. At the same time, a large language model can also be used to convert structured project domain data mapped to the same system into structured project domain vectors with the same dimensions. Here, the vector encoding needs to correspond to the level or category of the unified subject classification system. For example, each dimension represents a specific classification node in the system, ensuring that the vector can accurately carry the semantic information of the structured data.
[0155] S422. Determine the cosine similarity between the structured research direction vector and the structured project domain vector, as the matching degree of the subject classification level.
[0156] In this implementation, the cosine similarity between the structured research direction vector and the structured project domain vector can be calculated and used as the matching degree at the subject classification level. The cosine similarity measures the consistency of direction by calculating the cosine value of the angle between the two vectors. The smaller the angle and the closer the cosine value is to 1, the higher the degree of matching between the two at the subject classification level; the larger the angle and the closer the cosine value is to -1, the lower the degree of matching. This method transforms the abstract subject matching relationship into a quantifiable numerical value, avoiding the ambiguity of subjective judgment.
[0157] This implementation transforms structured research directions and project domain data into features in a mathematical space through vector representation. Cosine similarity is used to measure the similarity between the two directions, and the similarity of directions directly corresponds to the degree of matching of the subject classification level. This avoids the bias of subjective judgment and effectively improves the accuracy of the matching degree of the subject classification level. The result of adjusting the comprehensive index of the research space with this matching degree is more in line with the matching needs of researchers' actual research directions and project domains.
[0158] This implementation method ensures that the calculation standard for the matching degree of different researchers and different projects at different subject classification levels is consistent. It will not cause the result deviation due to differences in data format or classification system. The consistent calculation standard makes the matching degree result more valuable for reference. This allows the comprehensive index of research space adjusted by the matching degree to more fairly and accurately reflect the degree of matching between researchers' research direction and project field, thereby improving the rationality of the recommendation result.
[0159] This application also provides a research project recommendation device based on an AI large language model, including a unit for implementing the method described above.
[0160] Figure 14 A schematic diagram of the logical structure of a research project recommendation device based on an AI large language model, provided for embodiments of this application, is shown below. Figure 14 As shown, the apparatus 1 of this embodiment includes a processing unit 11, a storage unit 12, and a transceiver unit 13. The processing unit 11 is used to process data, the storage unit 12 is used to store data, and the transceiver unit 13 is used to send and receive data. The processing unit 11, the storage unit 12, and the transceiver unit 13 cooperate with each other to implement the above-described method. The beneficial effects of the embodiments of this application have been described in the above-described method and will not be repeated here.
[0161] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0162] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0164] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0165] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0167] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0168] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A research project recommendation method based on an AI large language model, characterized in that, The method includes: This method utilizes an AI-powered large language model to identify multiple research projects corresponding to the current research direction. It also identifies the research boundaries, common spaces within related fields, available space for research entry points, and overlapping / conflicting areas between adjacent research projects. Finally, it determines the maximum permissible overlap range for each research project. Specifically, by analyzing project proposals, research plans, or publicly available results, the method extracts these elements. The research boundary is the explicitly defined research scope of the project; the common space within related fields is the intersection of the project with other relevant fields; the available space for research entry points represents unexplored innovative directions within the project; the overlapping / conflicting areas between adjacent research projects represent the proportion of overlap between the current project and other projects in the same field; and the maximum permissible overlap range is a threshold determined based on research conventions, funding agency requirements, or industry standards to limit excessive duplication of research between projects. Based on the research boundaries of multiple research projects, the public space of related fields, and the available space of research entry points, determine the total research space of multiple research projects; determine the difference between the overlapping conflict range of adjacent research projects and the maximum allowable overlap range, as the boundary occupancy difference; The total research space of multiple research projects and the difference in the over-limit occupancy value are determined as the research space index of multiple research projects; multiple research projects are recommended to users in descending order of research space index.
2. The method according to claim 1, characterized in that, Based on the research boundaries of multiple research projects, the common space of related fields, and the available space of research entry points, the total research space of multiple research projects is determined, including: By using the AI large language model, we can determine the research boundary vector of the research project, the related domain vector corresponding to the public space of the related domain, the available space vector corresponding to the boundary set of the available space of the research entry point, and the domain overlap and conflict vector corresponding to the domain overlap and conflict range of adjacent research projects. Determine the L2 norm of the research boundary vector, the L2 norm of the related domain vector, and the L2 norm of the available space vector; determine the sum of the L2 norm of the research boundary, the L2 norm of the related domain, and the L2 norm of the available space for the research project, which is taken as the total research space of the research project.
3. The method according to claim 2, characterized in that, The difference between the overlapping conflict range and the maximum allowable overlap range of multiple adjacent research projects is determined as the boundary occupancy difference, including: Using AI large language model, determine the domain overlap conflict vector corresponding to the domain overlap conflict range of adjacent scientific research projects, and the maximum allowable overlap vector corresponding to the maximum allowable overlap range. Determine the L2 norm of the domain overlap conflict vector and the L2 norm of the maximum allowable overlap vector; determine the difference between the L2 norm of the domain overlap conflict and the L2 norm of the maximum allowable overlap of the research project, and use it as the over-limit occupancy difference of the research project.
4. The method according to claim 3, characterized in that, The method further includes: Obtain the resource path coefficient and average resource acquisition time of the research project; determine the product of the resource path coefficient and average resource acquisition time of the research project as the comprehensive resource efficiency value of the research project. Based on the research space indicators and resource comprehensive efficiency values of research projects, the comprehensive research space indicators of research projects are determined; and multiple research projects are recommended to users in descending order of comprehensive research space indicators.
5. The method according to claim 4, characterized in that, Based on the research space indicators and resource comprehensive efficiency values of research projects, the comprehensive research space indicators of research projects are determined, including: Obtain the weights of research space indicators and the weights of comprehensive resource efficiency; The sum of the product of the research spatial index and the weight of the research spatial index, and the product of the comprehensive resource efficiency value and the weight of the comprehensive resource efficiency value, is determined as the comprehensive research spatial index of the research project.
6. The method according to claim 5, characterized in that, The method further includes: The research scope and trajectory of multiple historical research projects of researchers are obtained; the research trajectory of multiple historical research projects is scaled within the research scope to obtain the research trajectory of multiple historical research projects in a unified dimension; the overlap analysis of the research trajectory of multiple historical research projects in a unified dimension is performed to determine the union of the research trajectories; The degree of overlap between the union of research trajectories and the research boundaries of multiple research projects, the public space of related fields, and the available space of research entry points is determined and used as the weight of research space indicators.
7. The method according to claim 6, characterized in that, The method further includes: Find the research field to which the research project belongs, and obtain the average execution time of the research field to which the research project belongs; determine the average execution time of all research projects in the research field to which the research project belongs, as the average execution time of the field. The ratio of the average time spent in the field to the basic time spent in the field is determined and used as the weight for the overall resource efficiency of scientific research projects.
8. The method according to claim 7, characterized in that, The method further includes: Obtain information on researchers' research directions and the research fields to which research projects belong; map researchers' research direction information to a unified subject classification system to obtain structured research direction data; map the research fields to which research projects belong to a unified subject classification system to obtain structured project field data. Based on structured research direction data and structured project field data, the matching degree of subject classification level is determined; the comprehensive index of research space of research projects is adjusted by multiplying the matching degree of subject classification level by the comprehensive index of research space of research projects.
9. The method according to claim 8, characterized in that, Based on structured research direction data and structured project area data, the matching degree of subject classification hierarchy is determined, including: Determine the structured research direction vector corresponding to the structured research direction data; determine the structured project domain vector corresponding to the structured project domain data; The cosine similarity between the structured research direction vector and the structured project domain vector is determined as the matching degree of the subject classification level.
10. A research project recommendation device based on an AI large language model, characterized in that, Includes units for implementing the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent matching and recommendation method and device based on user and content similarity
CN118760801A
KR20210143434A