A scientific research achievement whole-cycle management system and method based on multi-modal intelligent analysis
The multimodal intelligent analysis-based full-cycle management system for scientific research achievements solves the problems of inefficient processes, fragmented knowledge, and low utilization rates in the management of research achievements in scientific research institutions. It realizes intelligent management and efficient transformation of scientific research achievements throughout the entire process, thereby enhancing the competitiveness and utilization rate of research achievements in scientific research institutions.
Patent Information
- Application Number
- CN202511163739.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Research institutions face problems in managing research results, such as inefficient process control, unclear data on research results, lack of knowledge transformation, and insufficient knowledge utilization. This leads to repetitive work, low efficiency, and waste of resources. Furthermore, existing systems cannot achieve intelligent management throughout the entire lifecycle of research results.
The system adopts a full-lifecycle management system for scientific research achievements based on multimodal intelligent analysis. Through business process automation, knowledge-enhanced retrieval, and intelligent push, it realizes intelligent management of the entire process of scientific research achievements from generation, review, submission, archiving to reuse. This includes automated multi-departmental collaborative approval processes, semantic parsing and intelligent transformation of multimodal scientific research achievements, and intelligent push and service closed-loop optimization of scientific research achievements.
It improved process efficiency, eliminated repetitive work, transformed scattered results into a searchable knowledge base, enabled precise knowledge delivery, enhanced user experience and satisfaction, and promoted the comprehensive application and innovation of scientific research results.
Smart Images

Figure CN120723980B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of collaborative management and knowledge enhancement of scientific research achievements of scientific research and testing units, and in particular to a scientific research achievement whole-cycle management system and method based on multi-modal intelligent analysis. BACKGROUND
[0002] The statements in this section merely provide background information related to the present disclosure and can not constitute the prior art.
[0003] Scientific research achievements (such as reports, standards, papers, patents, systems, etc.) are the core assets of scientific research and testing units, and have important value and significance in the academic and technical influence, resource acquisition and sustainable development, technology transformation and economic benefits, institutional reputation and talent attraction, strategic support and social responsibility of scientific research and testing units. However, for a long time, scientific research units have generally faced three major pain points in the management and utilization of scientific research achievements:
[0004] First, the business control process is inefficient. The number, type, and status of scientific research achievements are not clear, information is scattered, and heterogeneous data integration is difficult, leading to repeated registration and reporting, repeated labor of scientific researchers (such as the need to submit the same materials for patent application and expense reimbursement processes), repeated checking by government management personnel, low efficiency, and errors.
[0005] Second, the knowledge transformation of achievements is lacking. Scientific research achievement data is managed in a scattered manner, archiving is not standardized, there is a lack of a unified knowledge management platform, and there is a lack of intelligent search and analysis tools, which cannot effectively realize the transformation and accumulation of single and scattered scientific research achievements into systematic knowledge assets that can be inherited and reused by the unit, and it is difficult to support the sharing and application of field professional knowledge.
[0006] Third, the knowledge fusion application is insufficient. A large amount of scientific research knowledge such as reports, standards, papers, and systems accumulated by scientific research units is stored in static files and continues to grow. This knowledge is necessary for scientific researchers to carry out work in an orderly manner, on the one hand, scientific researchers need to spend time and effort to find and organize it, on the other hand, it is not associated with existing scientific research business-related systems and processes, and is disconnected from the work environment, resulting in low knowledge utilization rate and resource waste.
[0007] The above problems involve the main links of the generation, archiving and reuse process of scientific research achievements, and interact and restrict each other: inefficient process leads scientific researchers to consume a lot of energy in repeated reporting and manual coordination instead of core scientific activities; knowledge fragmentation hinders the accumulation and inheritance of field professional knowledge, making it difficult for historical achievements to be transformed into reusable intelligent assets; application disconnection makes the knowledge base become a static archive library and cannot be deeply integrated with business scenarios such as test design and project declaration, forming a “knowledge island”. These problems and their intertwined effects not only restrict the efficiency of scientific research and the transformation rate of achievements, but also weaken the competitiveness of scientific research units in complex scientific research tasks (such as tests and multidisciplinary collaborative innovation), and an intelligent management solution for scientific research achievements is urgently needed to realize the fundamental change from “dispersed storage” to “intelligent service”.
[0008] According to the existing relevant literature query, there are few public literatures on the intelligent management of the whole process of scientific research achievements from generation, review, submission, archiving to reuse in scientific research units at home and abroad.
[0009] From these public reports, the existing research literatures mainly have the limitations of single data management integration (most systems only focus on the management of single-class achievements such as patents and papers, and lack the unified integration of multi-modal data such as test reports and project documents), weak cross-system collaboration (approval processes are isolated from business systems such as financial reimbursement and security review, and data synchronization relies on manual operation), low knowledge reuse rate (the matching degree between search results and user needs is poor, and the citation rate of historical achievements in existing systems is low), insufficient scene adaptability (mostly using passive service mode, relying on users to initiate queries, and unable to predict scientific research needs, such as actively pushing risk prompts in the project start-up stage), etc. cannot meet the business needs of scientific research and test units in knowledge management, job training, knowledge application, and operation guidance, and there is no public report on a scientific research achievement whole life cycle management system suitable for scientific research and test units. SUMMARY
[0010] The purpose of the present application is to provide a scientific research achievement whole cycle management system and method based on multi-modal intelligent analysis, aiming at the problems of low process control efficiency, unclear achievement data, lack of knowledge transformation, and broken knowledge circulation in the management of scientific research achievements in scientific research and test units.
[0011] Specifically, the present application aims to solve the following technical problems:
[0012] First, the problem of repeated labor and manual errors caused by inefficient business processes;
[0013] Second, the problem of low transformation efficiency caused by dispersed management of scientific research achievements knowledge;
[0014] Third, the utilization rate problem caused by the disconnection between static knowledge storage and business scenarios.
[0015] Specifically, the present application proposes a scientific research achievement full-cycle management system and method based on multi-modal intelligent analysis to solve the above problems, which realizes the intelligent management of the whole process of scientific research achievements from generation, review, submission, archiving to reuse through business process automation, knowledge enhanced search and intelligent push, and effectively promotes the comprehensive application of scientific research achievements of scientific research test units, and better stimulates the innovation vitality of scientific research personnel.
[0016] The technical scheme of the present application is as follows:
[0017] A scientific research achievement full-cycle management system based on multi-modal intelligent analysis, comprising:
[0018] A business management and control module for realizing condition-triggered multi-department collaborative approval process automation through a weighted state transition equation and a dynamic role weight matrix;
[0019] A knowledge enhancement module based on search enhancement generation technology for performing semantic analysis and intelligent conversion on multi-modal scientific research achievements to generate structured knowledge assets;
[0020] An accurate service module for realizing intelligent push and service closed-loop optimization of scientific research achievements through a collaborative filtering algorithm and a dynamic weight adaptation mechanism.
[0021] Further, the business management and control module comprises:
[0022] A process modeling unit configured to adopt a business process model and symbolically define process nodes, roles and rules, and describe process state transitions through a weighted state transition equation and a dynamic role weight matrix;
[0023] A role modeling unit configured to generate a dynamic role weight matrix based on the number of inter-department collaborations and approval efficiency;
[0024] An automatic triggering unit configured to calculate a triggering probability through the product of role weight and state transition condition, and trigger automatic flow when the calculation result is greater than or equal to a preset threshold, otherwise, manual processing.
[0025] Further, the weighted state transition equation is as follows:
[0026]
[0027] wherein, is the current state, is the first condition rule for the process to enter the next state from the current state, is the weight of the condition rule, is the weight of the role. The role weight of each condition To trigger the event, This is the urgency coefficient;
[0028] The dynamic role weight matrix is as follows:
[0029]
[0030] in, For the department With Department The collaboration weight between them Given the rarity of interdepartmental collaboration, For the department With Department The relative frequency of collaboration This refers to the number of approval processes completed per unit of time.
[0031] Furthermore, the trigger probability is expressed as:
[0032]
[0033] in, In the current state and roles Under the condition, transition to the next state. The probability is calculated as follows: if the result is 1, it means that the transfer will definitely occur and the process can be triggered automatically; if the result is 0, it means that the transfer will not occur; if the result is an intermediate value, it means that the transfer may occur, but there are probability limits or role weight adjustments, which require manual operation. From state Spontaneous transition to state The base probability, This indicates that the transfer path is allowed. This indicates that the transfer path is prohibited; Indicates role For the transfer path The influence weight.
[0034] Furthermore, the knowledge enhancement module includes:
[0035] The multimodal processing unit is configured to perform structured parsing of scientific research results in Excel, PDF, and Word formats. Excel parsing maps cell data to structured objects, PDF parsing extracts logical structure through optical character recognition and layout analysis, and Word parsing extracts hierarchical elements of paragraphs and tables.
[0036] The hybrid retrieval unit is configured to generate document vectors using a semantic coding model and perform linear weighted retrieval by combining keyword relevance scores.
[0037] Furthermore, the hybrid retrieval unit performs the following operations:
[0038] The document content is semantically encoded using a dual encoder model to generate dense vectors representing semantic information.
[0039] Keyword weight scores are calculated by statistically analyzing word frequency and inverse document frequency.
[0040] The semantic search results and keyword search results are fused together using a weighting coefficient. The calculation formula is: semantic weighting coefficient × semantic similarity score + (1 - semantic weighting coefficient) × keyword relevance score.
[0041] Furthermore, the knowledge enhancement module also includes:
[0042] The fusion unit is configured to calculate the semantic similarity between the retrieved fragment and the user query, combine the temperature parameter to allocate fusion weights, and achieve information recombination through an attention mechanism;
[0043] The generation unit is configured to generate structured text based on a pre-trained language model and perform grammatical correction and logical consistency optimization.
[0044] Furthermore, the precision service module includes:
[0045] The dynamic synchronization unit is configured to filter the candidate document set based on a semantic vector similarity threshold.
[0046] The collaborative filtering unit is configured to calculate user similarity and construct a neighbor set based on user behavior data;
[0047] The scenario adaptation unit is configured to perform weighted filtering of the data pool based on the type of business process.
[0048] Furthermore, the precision service module also includes:
[0049] The push calculation unit is configured to take into account semantic similarity, user similarity, scenario weight, and initial weight coefficients selected from user information corresponding to business processes. Generate a TOP-N push sequence;
[0050] The real-time feedback unit is configured to dynamically adjust the weight coefficient based on the user click-through rate. .
[0051] This invention also proposes a method for full-cycle management of scientific research achievements based on multimodal intelligent analysis. The aforementioned full-cycle management system for scientific research achievements based on multimodal intelligent analysis includes:
[0052] The dynamic jump and automatic triggering of the multi-department collaborative approval process are realized through the collaborative calculation of the weighted state transition equation and the dynamic role weight matrix.
[0053] The multi-modal scientific research achievements are subjected to optical character recognition, semantic coding and mixed retrieval, and a structured knowledge base is generated.
[0054] The collaborative filtering algorithm is optimized based on user behavior feedback data, and the dynamic adaptation of knowledge pushing and business scenarios is realized.
[0055] Compared with the existing technology, the beneficial effects of the present application are:
[0056] 1. Process efficiency is improved: the system eliminates the repetitive labor of researchers in the process of achievement declaration and expense reimbursement through intelligent process automation and knowledge reuse.
[0057] 2. Systematization of knowledge assets: scattered scientific research achievements are converted into searchable and related knowledge base, such as mapping wind tunnel test parameters, paper conclusions and patent technology points into unified entity relationships; through intelligent analysis of user behavior (such as frequently queried literature and frequently cited test databases), common needs in the field are identified and automatically archived as "knowledge template library", and implicit knowledge is made explicit.
[0058] 3. User experience and satisfaction are improved: accurate knowledge pushing is realized, for example, when the user starts the "acoustic wind tunnel test design" task, the system automatically pushes related historical test anomaly data reports, similar scene patents and other related materials; new employees can quickly retrieve the best practice documents and historical problem solving solutions in the field through the system, reducing the learning curve. BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 The data flow chart of the knowledge enhancement module;
[0060] Figure 2 The flow chart of the scene data pool screening. DETAILED DESCRIPTION
[0061] It should be noted that the relational terms herein, such as first and second and the like, are used solely to distinguish one from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0062] The features and characteristics of the present application will be further described with reference to the following examples.
[0063] Example One
[0064] A scientific research achievement whole-cycle management system based on multi-modal intelligent analysis, comprising:
[0065] A business management and control module, configured to realize conditional trigger type multi-department collaborative approval process automation through a weighted state transition equation and a dynamic role weight matrix; it should be noted that the function description of the business management and control module is: realizing multi-department collaborative approval based on a workflow engine, and supporting conditional trigger type process automation;
[0066] A knowledge enhancement module, configured to perform semantic analysis and intelligent conversion on multi-modal scientific research achievements based on a retrieval enhancement generation technology, and generate structured knowledge assets; it should be noted that the function description of the knowledge enhancement module is: taking semantic analysis and intelligent conversion of scientific research achievements as a core goal, converting scattered, unstructured, multi-modal scientific research achievements (such as reports, papers, patents, standards, etc.) into reusable, associable structured knowledge assets through a RAG retrieval enhancement generation (retrieval-generation-fusion) technology, and supporting dynamic generation of text content meeting user needs, realizing structured storage and intelligent generation of scientific research achievements;
[0067] A precise service module, configured to realize intelligent push and service closed-loop optimization of scientific research achievements through a collaborative filtering algorithm and a dynamic weight adaptation mechanism; it should be noted that the function description of the precise service module is: establishing seamless connection with the knowledge enhancement module, taking the document vector library formed by the above-mentioned BGE model as the semantic basis of the push system, and matching user scenarios to dynamically adjust weights, performing push calculation, triggering bidirectional optimization according to user behaviors (such as click rate, etc.), and forming a collaborative evolution mechanism of knowledge update→scene adaptation→push calculation→parameter optimization→closed-loop feedback.
[0068] In the embodiment, specifically, the business management and control module comprises:
[0069] A flow modeling unit configured to define flow nodes, roles and rules by BPMN (Business Process Model and Notation), and describe flow state transition by a weighted state transition equation and a dynamic role weight matrix;
[0070] A role modeling unit configured to generate a dynamic role weight matrix based on the number of inter-department collaborations and approval efficiency;
[0071] An automatic triggering unit configured to calculate a triggering probability by the product of role weight and state transition condition, trigger automatic flow when the calculation result is greater than or equal to a preset threshold, and otherwise, transfer to manual processing.
[0072] In the embodiment, specifically, the weighted state transition equation is as follows:
[0073]
[0074] Wherein:
[0075] is the current state;
[0076] is the first condition rule (such as “only when the achievement passes the audit, can it enter the approval stage”) for the flow to enter the next state from the current state; is the th condition rule;
[0077] is the role weight of the th condition (such as file integrity weight , approval level weight );
[0078] is a triggering event (such as “automatic flow to the next node” or “triggering financial audit”);
[0079] is an urgency coefficient (such as urgent task , ordinary task );
[0080] The dynamic role weight matrix based on user behavior is as follows:
[0081]
[0082] Wherein:
[0083] is the department and the department the collaboration weight between departments;
[0084] the rarity of the collaboration relationship between departments;
[0085] the department collaborating with; the relative frequency of collaboration;
[0086] the number of processes completed in a unit of time.
[0087] In this embodiment, the rarity of the collaboration relationship between departments is calculated by the following formula:
[0088]
[0089] wherein:
[0090] is the total number of collaborations between all departments;
[0091] is the number of collaborations between the department and the department collaborating with;
[0092] the constant 1 is to avoid the denominator being 0.
[0093] In this embodiment, the relative frequency of collaboration between the department and the department collaborating with is calculated by the following formula:
[0094]
[0095] wherein:
[0096] is the total number of departments.
[0097] In this embodiment, the number of processes completed in a unit of time is calculated by the following formula:
[0098]
[0099] wherein:
[0100] is the number of collaboration approvals completed by the department and the department in a unit of time.
[0101] In this embodiment, the trigger probability is specifically expressed as:
[0102]
[0103] in:
[0104] In the current state and roles Under the condition, transition to the next state. The probability is calculated as follows: if the result is 1, it means that the transfer will definitely occur and the process can be triggered automatically; if the result is 0, it means that the transfer will not occur; if the result is an intermediate value, it means that the transfer may occur, but there are probability limits or role weight adjustments, which require manual operation.
[0105] From state Spontaneous transition to state The base probability, This indicates that the transfer path is allowed. This indicates that the transfer path is prohibited;
[0106] Indicates role For the transfer path The influence weight.
[0107] In this embodiment, it should also be noted that after each task is completed, the efficiency is adjusted according to the process. (e.g., completion time t) to dynamically adjust weights:
[0108]
[0109] in:
[0110] For the first The updated weights for each condition;
[0111] For learning efficiency (e.g., 0.1);
[0112] It refers to the efficiency of the current task process, which may be a metric such as completion time or error rate;
[0113] This represents the historical expected value, or average value, of process efficiency.
[0114] The variance representing process efficiency is used for standardization. The changes.
[0115] In this embodiment, as Figure 1 As shown, the knowledge enhancement module includes:
[0116] a multi-modal processing unit configured to perform structured parsing on scientific research achievements in Excel, PDF, and Word formats, wherein the Excel parsing implementation unit maps cell data to structured objects, the PDF parsing extracts logical structures through optical character recognition and layout analysis, and the Word parsing extracts hierarchical elements of paragraphs and tables;
[0117] a hybrid retrieval unit configured to generate document vectors using a semantic encoding model and perform linearly weighted retrieval in combination with keyword relevance scores.
[0118] In this embodiment, it should be noted that the processing target of the multi-modal processing unit is to uniformly represent the input achievements (text, charts, formulas, etc.) in multiple modalities, which is as follows:
[0119] a. Excel data parsing
[0120] Excel data is parsed using openpyxl to extract the title, options, and answers in the cells, map each row of data to a structured object containing complete title information, and finally format the structured data into text paragraphs;
[0121] b. PDF data parsing
[0122] When processing PDF data, OCR (Optical Character Recognition) processing is first required to convert the text in the image into editable text. Layout analysis is to identify the text layout in the PDF file, including text lines, paragraphs, and page structure. This helps to understand the logical structure of the text and more accurately extract information.
[0123] c. Word data parsing
[0124] The paragraphs, tables, images, etc. in the Word document are parsed using python-docx. The text format is cleaned and converted (e.g. removing redundant symbols, unifying encoding, etc.). The parsed data is saved in the target format (e.g. JSON, text file, or database record).
[0125] In this embodiment, specifically, the hybrid retrieval unit performs the following operations:
[0126] The document content is semantically encoded by a double-encoder model to generate dense vectors representing semantic information;
[0127] The keyword weight score is calculated by counting the term frequency and inverse document frequency;
[0128] The semantic retrieval result and the keyword retrieval result are fused by a weight coefficient, and the calculation formula is: semantic weight coefficient x semantic similarity score + (1-semantic weight coefficient) x keyword relevance score.
[0129] In the present embodiment, it should be noted that the target of the hybrid retrieval unit is to perform hybrid retrieval based on BGE semantic retrieval and BM25 keyword retrieval, to efficiently retrieve semantic information related to user demand from massive heterogeneous scientific research achievements, as follows:
[0130] a. BGE model text representation
[0131] The BGE model is used for semantic encoding of the document content to generate a dense vector representation:
[0132]
[0133] Wherein:
[0134] is the input scientific research achievement original document (such as scientific research report, patent abstract, paper, etc.);
[0135] is a double-encoder model for mapping text to a dense semantic vector, such as a BERT-based architecture;
[0136] is the vector dimension of the model output (such as 768, 1024, etc.), used to control the fineness of semantic representation. Each dimension of this vector corresponds to a semantic direction captured by the model, used to describe the semantic information, theme, entity or syntax structure of the document, etc.
[0137] b. BM25 keyword extraction
[0138] The BM25 algorithm is used for keyword extraction, and the keyword weight is calculated by statistics of document term frequency TF and inverse document frequency IDF:
[0139]
[0140] Wherein:
[0141] represents the relevance score of the word t (which is a word in the query) in the document d to the query q;
[0142] measures the importance of word t in the corpus;
[0143] represents the number of times the word t appears in the document;
[0144] Parameter Controlling the saturation of term frequency, TF;
[0145] Parameter Adjusting the intensity of document length normalization;
[0146] For document length;
[0147] For corpus average document length.
[0148] c. Semantic retrieval cooperates with keyword retrieval
[0149] Through linear weighting, balance the pros and cons of semantic retrieval (capture deep semantic association, but may not be sensitive enough to keywords) and keyword retrieval (quickly match explicit keywords, but cannot understand implicit semantics), realize the cooperation of BGE semantic retrieval and BM25 keyword retrieval:
[0150]
[0151] Among them:
[0152] Semantic weight (usually set to 0.7), balance the contribution of semantic retrieval and keyword retrieval;
[0153] According to the comprehensive score Sort the documents, and return the top-K document fragments with high semantic similarity and matching keywords in priority.
[0154] In this embodiment, specifically, the knowledge enhancement module further includes:
[0155] The fusion unit is configured to calculate the semantic similarity between the retrieval fragments and the user query, combine the temperature parameter to perform fusion weight distribution, and realize information reorganization through attention mechanism; it should be noted that the fusion target of the fusion unit is to align the retrieved fragmented information with the user's original input context semantically, and generate a contextualized and structured new input for the generation model;
[0156] The generation unit is configured to generate structured text based on a pre-trained language model, and perform syntax correction and logical consistency optimization; it should be noted that the target of the generation unit is to generate text content (such as test reports, abstracts, etc.) that meets scientific research specifications and is semantically coherent based on the fused input.
[0157] In this embodiment, it should be noted that the fusion unit is implemented through the following specific algorithm:
[0158] a. Context similarity calculation
[0159] For each retrieval fragment The matching degree with the user query q is calculated, and the vector generated by the BGE model is used to calculate the retrieval fragment and the semantic similarity of the query q:
[0160]
[0161] wherein:
[0162] denotes the retrieval fragment and the semantic similarity of the query q;
[0163] denotes the retrieval fragment and the dense semantic vector of the query q; denotes the dense semantic vector of the query q;
[0164] denotes the length of the modified retrieval fragment is the length penalty coefficient (such as = -0.2), to avoid long text dominating the results.
[0165] b. Fusion weight distribution
[0166] According to the similarity and information relevance, the fusion weight is distributed:
[0167]
[0168] wherein:
[0169] is the fusion weight, which is dynamically distributed based on the similarity ;
[0170] is the temperature parameter, which controls the concentration of the weight distribution, when it is larger (such as 1.0), the general fusion is emphasized, and the weights of all fragments are close; when it is smaller (such as 0.5), the precise match is emphasized, and the weight of the high-similarity fragment is significantly higher than that of the low-similarity fragment;
[0171] denotes the retrieval fragment and the semantic similarity of the query q;
[0172] c. Information recombination rule
[0173] According to , the fragment information is weighted and integrated, and the key semantics of the original input is preserved:
[0174]
[0175] Where:
[0176] Represents semantic fusion based on attention mechanism (such as Cross-Attention of Transformer);
[0177] Can be a fragment Text content, vector representation, or structured information containing metadata such as location, timestamp, etc.
[0178] In this embodiment, it should be noted that the generation unit specifically has the following functions:
[0179] Text planning: according to the fused input, make the structure and content outline of the generated text;
[0180] Language generation: use a pre-trained text generation model such as GPT to generate text according to the planned structure; During generation, the model will consider grammar, semantics, and contextual coherence; In the process of generation, through multi-head self-attention mechanism Capture long-distance dependencies and correct ambiguities:
[0181]
[0182] Where:
[0183] Represents the target word generated by the model at the t-th step, which is linearly transformed after weighting the value vector and output;
[0184] Represents the semantic query vector of the current target word Generated from the linear transformation of the input text;
[0185] Represents the key vector of the input text, used to calculate the relevance with the query vector;
[0186] Represents the vector dimension of the query vector And the key vector Used for normalization calculation;
[0187] Represents the value vector of the input text, carrying the semantic information of the original text, which is weighted by attention weight to generate output.
[0188] Post-processing: The generated text may need to be post-processed, including correcting grammatical errors, eliminating duplicate content, ensuring transaction consistency, etc.
[0189] Output: The final generated text is output to the user.
[0190] In the embodiment, specifically, the accurate service module comprises:
[0191] The dynamic synchronization unit is configured to filter the candidate document set based on a semantic vector similarity threshold;
[0192] The collaborative filtering unit is configured to calculate user similarity and construct a neighbor set through user behavior data;
[0193] The scene adaptation unit is configured to perform weighted filtering on the data pool according to a business link type;
[0194] The push calculation unit is configured to select an initial weight coefficient by comprehensively considering semantic similarity, user similarity, scene weight, and initial weight coefficient of user information corresponding to a business link to generate a TOP-N push sequence;
[0195] The real-time feedback unit is configured to dynamically adjust the weight coefficient according to a user click rate .
[0196] In the embodiment, it should be noted that the dynamic synchronization unit is mainly used to realize dynamic knowledge weight synchronization, and specifically as follows:
[0197] The BGE model based on the knowledge enhancement module generates semantic vectors of the scientific research achievement document and the user query, and uses a similarity calculation result to filter a preliminary candidate document set, so as to ensure that the push calculation unit always uses the latest semantic vector generated by the knowledge enhancement module;
[0198]
[0199] Among them:
[0200] represents the cosine similarity of the semantic vector generated by the BGE model encoding;
[0201] represents the semantic vector generated by the BGE model encoding of the input document d;
[0202] represents the semantic vector generated by the BGE model encoding of the user query q.
[0203] The filtering threshold is the candidate document set , and the similarity is sorted in size:
[0204]
[0205] Among them:
[0206] represents the input document.
[0207] In the present embodiment, it is to be noted that the collaborative filtering unit is mainly used to implement push based on collaborative filtering, and the implementation is as follows:
[0208] The similarity between users is evaluated by the scores of different users on the items, and the recommendation is made according to the similarity between users. It is known that user u has inquired about a scientific research project knowledge achievement, and user u is highly similar to user v, so it can be predicted that user v also wants to inquire about the scientific research project knowledge achievement, and the scientific research project related knowledge achievement can be pushed to user v.
[0209] User , The content similarity is calculated as follows:
[0210]
[0211] Wherein:
[0212] The content similarity of user ,
[0213] The set of the project knowledge inquired by user u is represented as:
[0214] The set of the project knowledge inquired by user v is represented as:
[0215] Through the above algorithm, the similarity of the search content of the current user with other users can be calculated, and the n users with the highest similarity with the current user are listed as the neighbor set. The collaborative filtering algorithm recommends the existing knowledge model of the users in the neighbor set to the current user, i.e. the potential knowledge model of the user;
[0216] Based on the calculated similarity, the demand probability of the current user u for the unsearched knowledge is predicted by using the weighted average method, and the prediction table formula is:
[0217]
[0218] Wherein:
[0219] The set of the demand probability of the current user u for the unsearched knowledge predicted by using the weighted average method is represented as:
[0220] The set of users similar to the current user is represented as:
[0221] The search frequency of similar user v on knowledge i is represented as:
[0222] In the embodiment, it should be noted that the scene adaptation unit is mainly used for weight adaptation of scene data pool filtering, and the specific process is as follows:
[0223] After the initial push knowledge is calculated, before the push, in order to improve the push efficiency and accuracy, the data pool needs to be filtered according to different business links. Each user corresponding to each business scene has a set of initial data pool weight system S representing the user's tendency to each data pool data, so as to realize more effective personalized knowledge push to the user. As shown in the following formula: Figure 2
[0224] Taking the project management system as an example, the project management system provides a help interface and a push knowledge display to the user as the main input and output. The probability array P obtained after the knowledge array is filtered through the data pool can be represented as:
[0225]
[0226] Among them:
[0227] Respectively represent the weight coefficient of user u in data pool 1, data pool 2 and data pool n;
[0228] Respectively represent the knowledge corresponding to data pool 1, data pool 2 and data pool n.
[0229] In the embodiment, it should be noted that the push calculation unit is mainly used for TOP-N push and dynamic weight correction, and the specific process is as follows:
[0230] The candidate knowledge set is calculated through the algorithm weight coefficient, and finally the initial weight coefficient is selected according to the user information corresponding to the business link to fuse and calculate the candidate knowledge set to obtain the probability of each knowledge push, so as to obtain the push order. According to the respective push knowledge order, the corresponding ranking can be obtained, and then the final push value is calculated according to the weight to obtain the recommended ranking. The specific formula is as follows:
[0231]
[0232] In the formula, represents the weight size of the ranking value based on the collaborative filtering algorithm; represents the weight size of the ranking value based on the similarity of the candidate document set; represents the recommended ranking value based on the collaborative filtering algorithm; represents the recommended ranking value based on the similarity of the candidate document set; represents the recommended ranking value of the mixed ranking calculated by the weight. Finally, the probability set is sorted in descending order, and the top The first N probabilities correspond to knowledge pushing.
[0233] In this embodiment, it should be noted that the real-time feedback unit mainly functions to feed back user behavior to the knowledge enhancement module, update the semantic vector, and specifically as follows: dynamic weight correction is performed by collecting user feedback to continuously optimize the pushing effect and improve the system adaptive ability.
[0234]
[0235] wherein: is the content similarity weight coefficient before iteration;
[0236] is the updated weight coefficient after user feedback correction;
[0237] is the learning coefficient, which controls the step size of weight correction;
[0238] is the gradient of the cross-entropy loss function Loss with respect to the weight coefficient .
[0239]
[0240] The cross-entropy loss function is used to quantify the error between the pushing result and the user's real behavior (such as clicking or not clicking).
[0241] wherein: is the binary label of whether the user clicks the document d (1 = click, 0 = no click);
[0242] The system predicts the user's click probability, and the score is mapped to the interval [0, 1] by the sigmoid function.
[0243] The embodiment also proposes a scientific research achievement whole-cycle management method based on multi-modal intelligent analysis, based on the above-mentioned scientific research achievement whole-cycle management system based on multi-modal intelligent analysis, comprising:
[0244] The dynamic role weight matrix is calculated to realize dynamic jump and automatic triggering of the multi-department collaborative approval process;
[0245] Optical character recognition, semantic encoding and hybrid retrieval are performed on the multi-modal scientific research achievements to generate a structured knowledge base;
[0246] The collaborative filtering algorithm is optimized based on user behavior feedback data to realize dynamic adaptation of knowledge pushing and business scenarios.
[0247] Embodiment Two
[0248] Example two further illustrates the scientific research achievement full-cycle management system and method based on multi-modal intelligent analysis proposed in example one, covering the whole process from generation, archiving to reuse, as follows:
[0249] 1. Patent application stage
[0250] Business control module:
[0251] After the user submits the patent draft, the system automatically triggers the approval process:
[0252] Call the rule engine to dynamically generate approval nodes (such as technical review, financial audit, etc.) according to the patent type (such as invention, utility model, etc.) and budget status.
[0253] Push tasks to the corresponding department (such as the technology department, finance department) through the BPMN engine, automatically extract keywords (BM25 algorithm) in the patent abstract to generate classification tags.
[0254] After approval, the system automatically generates a reimbursement form and associates the patent number to avoid duplicate submissions.
[0255] Knowledge enhancement module:
[0256] OCR scans patent documents, extracts metadata such as titles and abstracts, uses BGE models to generate semantic vectors, and stores them in the knowledge base.
[0257] Intelligently fill in repetitive content (such as claim templates) to reduce manual writing time.
[0258] 2. Achievement archiving stage
[0259] Knowledge enhancement module:
[0260] Associate archived patent documents with historical data and automatically generate technical point summaries through the RAG generation stage to form structured knowledge items.
[0261] Establish a mapping relationship between patents and wind tunnel test data (such as patent number → test number) to support cross-system retrieval.
[0262] 3. Knowledge reuse stage
[0263] Precision service module:
[0264] When the user logs into the system, the system calls the hybrid push algorithm:
[0265] Content matching: based on the user's current query (such as "acoustic wind tunnel design"), calculate the semantic similarity of patents in the knowledge base .
[0266] Collaborative filtering: analyze user historical behavior (e.g. browse "aerodynamic acoustic" related patents), recommend knowledge items in similar domain.
[0267] Weighted ranking: combine scores , push TOP-5 related patents and technical reports.
[0268] According to user feedback such as clicking on "a patent detail page", reading time exceeding 2 minutes, dynamically adjust (alpha) value (such as increasing the weight of collaborative filtering).
[0269] 4. Full life cycle closed loop
[0270] The system continuously collects user behavior data and optimizes the push algorithm.
[0271] Periodically reconstruct the historical achievements into knowledge graph to enhance the relevance of domain professional knowledge.
[0272] The above-described embodiments only express the specific implementation of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the protection scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the technical scheme concept of the present application, a number of modifications and improvements can be made, which all belong to the protection scope of the present application.
[0273] This background section is provided to generally present the context of the application, the work of the current named inventors, the work described in this background section to the extent that it is described, and the aspects described in this section at the time of filing, neither expressly nor implicitly, are recognized as prior art of the present application.
Claims
1. A scientific research achievement whole-cycle management system based on multi-modal intelligent analysis, characterized in that, Comprise: Business management and control module, for realizing condition-triggered multi-department collaborative approval process automation through weighted state transition equation and dynamic role weight matrix; Knowledge enhancement module, based on retrieval enhancement generation technology, for performing semantic analysis and intelligent conversion on multi-modal scientific research achievements, and generating structured knowledge assets; Precise service module, through collaborative filtering algorithm and dynamic weight adaptive mechanism, forming a collaborative evolution mechanism of knowledge update→scene adaptation→push calculation→parameter optimization→closed-loop feedback; The business management and control module comprises: Process modeling unit, configured to adopt business process model and symbol definition process node, role and rule, and describe process state transition through weighted state transition equation and dynamic role weight matrix; Role modeling unit, configured to generate dynamic role weight matrix based on the number of inter-department collaboration and approval efficiency; Automatic triggering unit, configured to calculate the triggering probability through the product of role weight and state transition condition, and trigger automatic flow when the calculation result is greater than or equal to the preset threshold, otherwise, manual processing; The weighted state transition equation is as follows: wherein, is the current state, is the flow from the current state into the next state the th condition rule, is the role weight of the th condition, is the trigger event, is the urgency coefficient; The dynamic role weight matrix is as follows: wherein, is the department with which the department collaborates, is the rarity of the inter-departmental collaboration, is the department with which the department collaborates, is the number of processes completed per unit of time; The triggering probability is expressed as: wherein, is the probability of transition to the next state and role under the current state condition, if the result is 1, it means that the transition must occur, allowing the process to be automatically triggered, if the result is 0, it means that the transition is impossible to occur, for intermediate values it means that the transition can occur, but there is a probability limit or role weight adjustment, which needs manual operation; is the basic probability of spontaneous transition from state to state , indicates that the transition path is allowed, indicates that the transition path is prohibited; indicates the influence weight of the role on the transition path ; The knowledge enhancement module comprises: Multi-modal processing unit, configured to perform structured analysis on scientific research achievements in Excel, PDF and Word formats, wherein the Excel analysis implementation unit realizes mapping of cell data to structured objects, the PDF analysis extracts logical structure through optical character recognition and layout analysis, and the Word analysis extracts hierarchical elements of paragraphs and tables; Hybrid retrieval unit, configured to generate document vectors using semantic encoding model, and perform linear weighted retrieval combining keyword relevance score; The hybrid retrieval unit performs the following operations: Semantically encode the document content through a double-encoder model to generate a dense vector representing semantic information; Calculate keyword weight score by counting word frequency and inverse document frequency; Fuse the semantic retrieval result and the keyword retrieval result through a weight coefficient, and the calculation formula is: semantic weight coefficient × semantic similarity score + (1-semantic weight coefficient) × keyword relevance score; The knowledge enhancement module further comprises: Fusion unit, configured to calculate the semantic similarity between the retrieval segment and the user query, distribute the fusion weight combining the temperature parameter, and realize information reorganization through attention mechanism; Generation unit, configured to generate structured text based on pre-trained language model, and perform syntax correction and logical consistency optimization; The precise service module comprises: Dynamic synchronization unit, configured to filter the candidate document set based on semantic vector similarity threshold; Collaborative filtering unit, configured to calculate user similarity and construct neighbor set through user behavior data; Scene adaptation unit, configured to weight filter the data pool according to the business link type; The precise service module further comprises: The push computing unit is configured to comprehensively select an initial weight coefficient of user information corresponding to a business link according to semantic similarity, user similarity, scene weight, and the initial weight coefficient generate a TOP-N push sequence A real-time feedback unit configured to dynamically adjust the weight coefficient according to the user click rate .
2. A scientific research achievement whole cycle management method based on multi-modal intelligent analysis, characterized in that, A scientific research achievement whole-cycle management system based on multi-modal intelligent analysis according to claim 1, comprising: Through the collaborative calculation of weighted state transition equation and dynamic role weight matrix, realize the dynamic jump and automatic trigger of multi-department collaborative approval process; Optical character recognition, semantic encoding and hybrid retrieval of multi-modal scientific research achievements to generate structured knowledge base; Optimize collaborative filtering algorithm based on user behavior feedback data to achieve dynamic adaptation of knowledge push and business scenarios.
Citation Information
Patent Citations
Intelligent agent role switching method and system based on multi-modal perception and related components
CN120197139A
Media information analysis and recommendation platform
US20100235313A1