Multi-source heterogeneous data intelligent matching system and method applied to department and wound supply chain online platform
By building an intelligent matching system for multi-source heterogeneous data, the problem of data structure and format differences on the scientific and technological innovation supply chain platform has been solved, multi-dimensional quality assessment and personalized matching have been achieved, and the matching accuracy and efficiency of scientific and technological innovation resources have been improved.
Patent Information
- Application Number
- CN202510760241.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-05
AI Technical Summary
Existing data matching technologies for the science and technology supply chain have difficulty handling the structural and format differences of multi-source heterogeneous data, lack personalized matching mechanisms, and are difficult to optimize matching strategies based on user feedback, resulting in matching results being out of touch with actual needs.
Build a multi-source heterogeneous data intelligent matching system that adapts to the scientific and technological innovation supply chain platform, including a data preprocessing module, a cross-modal semantic understanding module, an intelligent matching engine module and a feedback optimization module, and adopt multi-dimensional quality assessment, cross-modal information fusion and personalized adjustment mechanism to achieve data standardization, intent recognition and continuous optimization.
It improves the accuracy and efficiency of scientific and technological resource matching, solves the problems of semantic understanding deviation and poor user experience, and achieves precise docking and efficient flow.
Smart Images

Figure CN120597889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data matching technology, and in particular to a multi-source heterogeneous data intelligent matching system and method applied to an online platform for a scientific and technological innovation supply chain. Background Art
[0002] With the vigorous development of the science and technology innovation industry, the transformation of scientific and technological achievements and the synergy of the industrial chain have become key drivers of innovative economic growth. As a vital bridge connecting scientific research institutions, high-tech enterprises, and industry demanders, online platforms for science and technology innovation supply chains carry massive amounts of heterogeneous data from multiple sources, including scientific research results, patented technologies, enterprise needs, equipment resources, and other diverse information. Currently, mainstream science and technology innovation supply chain platforms generally use keyword matching, simple semantic similarity calculations, or traditional recommendation algorithms for data matching and resource recommendations. These platforms achieve cross-domain information association and matching through the construction of knowledge graphs, vector databases, and semantic retrieval systems.
[0003] However, existing data matching technologies for the S&T supply chain face numerous challenges. First, data in the S&T field is highly specialized and domain-specific. Data from different sources differ significantly in structure, format, and expression, resulting in uneven data quality and difficulty in unified processing. Second, traditional matching technologies struggle to accurately understand user query intent, especially when complex cross-domain technical requirements are involved, often leading to semantic understanding biases. Furthermore, existing systems lack effective personalized matching mechanisms, making it difficult to dynamically adjust matching strategies based on user historical behavior and preferences, resulting in a disconnect between matching results and actual needs.
[0004] Patent publication number CN118193694A discloses a large-scale voice question-answering system for a multi-source heterogeneous local knowledge base. This system processes multi-source heterogeneous data through semantic integrity segmentation, employs an adaptive knowledge base matching method for reasoning, and integrates virtual digital human technology to achieve voice interaction. Its matching mechanism is primarily based on simple vector similarity calculations, using the L1 norm to calculate inter-text similarity, lacking a multi-dimensional similarity assessment mechanism. Furthermore, its similarity threshold adjustment utilizes incremental updates based on a Gaussian function, resulting in a coarse update granularity and difficulty in achieving refined matching adjustments. Furthermore, while the system provides vector embedding and matching capabilities, it cannot distinguish the reliability and applicability of data from different sources. Finally, the system primarily focuses on a one-way query matching process, making it difficult to continuously optimize matching strategies based on user feedback. This limits the system's adaptability and matching accuracy over long-term use. Summary of the Invention
[0005] In light of this, the present invention provides a multi-source heterogeneous data intelligent matching system and method for an online platform for scientific and technological innovation supply chains. The system constructs a unified semantic representation system adapted to the characteristics of multi-source heterogeneous data on the platform, designs a fine-grained demand intent recognition algorithm, establishes a multi-dimensional resource quality assessment mechanism, constructs a cross-modal information fusion algorithm, designs an adaptive matching strategy adjustment mechanism, and establishes a continuous optimization mechanism based on user feedback. This system aims to improve the matching accuracy and efficiency of scientific and technological innovation resources, and achieve the precise connection and efficient flow of scientific and technological innovation elements such as scientific and technological achievements, enterprise needs, and expert resources.
[0006] The technical solution of the present invention is achieved as follows:
[0007] In one aspect, the present invention provides a multi-source heterogeneous data intelligent matching system applied to a scientific and technological supply chain online platform, comprising:
[0008] The data preprocessing module is used to standardize the multi-source heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform and generate data quality scores;
[0009] The cross-modal semantic understanding module is used to receive user queries, perform unified semantic feature representation on user queries and standardized scientific and technological resource data, and identify user query intent;
[0010] The intelligent matching engine module is used to calculate the similarity between the user's query intent and its semantic feature representation and the semantic feature representation of the scientific and technological resource data, and generate a matching score based on the data quality score and output a matching list;
[0011] The feedback optimization module is used to collect user feedback data on the matching list, convert the feedback data into real-time feedback features and feed them back to the intelligent matching engine module, recalculate the matching score and update the matching list until the preset stop condition is met.
[0012] Preferably, the data preprocessing module includes:
[0013] Data structure parser, used to perform structured analysis on multi-source heterogeneous scientific and technological resource data, identify and extract key fields and metadata;
[0014] Content normalization processor, used to perform format standardization, text cleaning and normalization on the parsed data;
[0015] Entity linking and disambiguation tools, which link and disambiguate professional terms, institution names, and technical concepts in the data;
[0016] A multi-dimensional quality assessor is used to perform a multi-dimensional assessment of data quality based on data completeness, consistency, timeliness, and authority, and generate a data quality score.
[0017] Preferably, the multi-dimensional quality assessor uses the following formula to calculate the data quality score:
[0018]
[0019] Among them, Q multi (r) is the comprehensive quality score of data r, K is the number of quality assessment dimensions, q k (r) is the original quality score of data r in the kth dimension, w k is the importance weight of the k-th dimension, σ(·) is the Sigmoid activation function, α k is the adjustment parameter, Ω k (type(d)) is the weight adjustment factor based on the data type.
[0020] Preferably, the cross-modal semantic understanding module includes:
[0021] A semantic understanding model for the science and technology innovation field, which is used to perform deep semantic encoding on user queries and standardized science and technology innovation resource data, and generate text semantic features;
[0022] A cross-modal information fusion model is used to deeply fuse text semantic features with structured information to form a cross-modal feature representation. For scientific and technological resource data, its text semantic features and structured information are integrated; for user queries, its text semantic features and query context information are integrated.
[0023] The user intent recognition model is used to analyze user queries and context information and identify the user's query intent type.
[0024] Preferably, the cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation of scientific and technological resource data:
[0025] F unified (r)=LayerNorm(W text F text (r)+W meta F meta (r)+W struct F struct (r)+C cross (r))
[0026] Among them, r represents scientific and technological innovation resource data, F unified (r) is the cross-modal feature representation of r, F text (r) is the text semantic feature of r, F meta (r) is the metadata feature of r, F struct (r) is the structural feature of r, W text 、W meta 、W structis the corresponding weight matrix, LayerNorm(·) is the layer normalization function, C cross (r) is the cross-modal interaction term;
[0027] The cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation of user queries:
[0028] F unified (q) = LayerNorm(W′ text F text (q)+W′ context F context (q)+C′ cross (q))
[0029] Among them, q represents the user query, F unified (q) is the cross-modal feature representation of q, F text (q) is the text semantic feature of q, F context (q) is the query context feature of q, W′ text and W′ context is the corresponding weight matrix, C′ cross (q) is the feature interaction term within the query.
[0030] Preferably, the intelligent matching engine module includes:
[0031] Intent-aware matching weight configurator, used to adjust the weight distribution of multi-dimensional similarity calculation based on the identified user query intent;
[0032] A multi-dimensional similarity calculator is used to calculate the multi-dimensional similarity between user queries and scientific and technological resource data based on semantic similarity, domain relevance, technical level similarity, and application scenario matching, combined with weight distribution;
[0033] A personalized regulator, used to generate a personalized adjustment factor based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics;
[0034] Matching score calculator, which integrates multi-dimensional similarity, data quality score and personalized adjustment factors to generate a matching score;
[0035] Candidate result ranking optimizer, used to generate matching lists based on matching scores.
[0036] Preferably, the calculation formula for the matching score is:
[0037]
[0038] Among them, q represents user query, r represents scientific and technological resource data, and M intent(q,r) is the matching score between q and r, D is the number of similarity calculation dimensions, intent(q) is the user query intent type of q, φ d (intent(q)) is the d-th dimension similarity weight based on query intent, F unified (r) is the cross-modal feature representation of r, F unified (q) is the cross-modal feature representation of q, Sim d (F unified (r),F unified (q)) is the similarity between q and r in the dth dimension, Q multi (r) is the data quality score of r, and Ψ(u,r) is the personalized adjustment factor of user u for resource r.
[0039] Preferably, the feedback optimization module includes:
[0040] User feedback collector, used to collect user feedback data on matching results, including explicit feedback and implicit feedback;
[0041] A real-time feedback feature converter, used to convert the collected feedback data into standardized real-time feedback features;
[0042] Feedback transmitter, used to feed back real-time feedback features to the intelligent matching engine module, triggering the recalculation of matching scores and the update of matching lists;
[0043] The stopping condition detector is used to monitor whether the preset stopping condition is met and determine whether the feedback optimization process is terminated.
[0044] Preferably, the feedback optimization module further includes an online learner, which regularly updates the model parameters using an online learning mechanism based on user feedback and newly added data.
[0045] On the other hand, the present invention also provides a multi-source heterogeneous data intelligent matching method applied to the online platform of the scientific and technological supply chain, the method is used to implement any of the above-mentioned systems, and the method includes:
[0046] S1. Perform structured analysis, content normalization, entity linking and disambiguation processing, and multi-dimensional quality assessment on the multi-source and heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform to generate standardized scientific and technological resource data and its quality score;
[0047] S2. Receive user queries, use cross-modal semantic understanding methods to create unified semantic feature representations for the user queries and standardized scientific and technological resource data, and identify the user's query intent.
[0048] S3. Based on the user's query intent and its semantic feature representation, a multi-dimensional similarity calculation is performed with the semantic feature representation of the scientific and technological resource data, and a matching score is generated by combining the data quality score and a personalized adjustment factor. The personalized adjustment factor is calculated based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics.
[0049] S4. Generate a matching list based on the matching score and push the matching list to the user;
[0050] S5, monitoring whether the preset stop condition is met, if so, ending the matching process, if not, executing step S6;
[0051] S6. Collect user feedback data on the matching list, including explicit feedback and implicit feedback;
[0052] S7: Convert the collected feedback data into standardized real-time feedback features and feed them back to step S3.
[0053] The present invention has the following beneficial effects compared to the prior art:
[0054] (1) The multi-source heterogeneous data intelligent matching system provided by the present invention realizes the standardized processing, unified semantic representation, intention-aware matching and feedback optimization of multi-source heterogeneous data on the scientific and technological innovation supply chain platform through four collaborative functional modules, effectively solving the problems of semantic understanding deviation, insufficient matching accuracy and poor user experience in data matching in the scientific and technological innovation field, and improving the accuracy and efficiency of scientific and technological innovation resource matching;
[0055] (2) The data preprocessing module of the present invention adopts a multi-dimensional quality assessment mechanism, which avoids the negative impact of low-quality data on matching results through comprehensive scoring of dimensions such as data integrity, consistency, timeliness and authority;
[0056] (3) The cross-modal semantic understanding module of the present invention is based on a pre-trained Transformer architecture and a multi-head attention mechanism, which converts data from different modalities into a unified semantic feature representation and accurately grasps the query purpose through a user intent recognition model, thus overcoming the problem that traditional matching systems lack understanding of complex professional terms and cross-domain concepts.
[0057] (4) The intelligent matching engine module of the present invention realizes the refined regulation and personalized customization of matching strategies through dynamic adjustment of intention-aware weights and multi-dimensional similarity calculation, combined with personalized adjustment factors, so that the matching results are more in line with the actual needs of different users in different scenarios;
[0058] (5) The feedback optimization module of the present invention constructs a complete closed-loop feedback mechanism. By collecting user explicit feedback and implicit feedback and converting them into standardized real-time feedback features, it continuously optimizes the matching strategy, giving the system adaptive learning capabilities and continuously improving matching accuracy as users interact. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0060] Figure 1 A schematic diagram of the system of the present invention;
[0061] Figure 2 It is a technical flow chart of the present invention;
[0062] Figure 3 Flow chart of the method of the present invention. DETAILED DESCRIPTION
[0063] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0064] like Figure 1 and Figure 2 As shown, the present invention provides a multi-source heterogeneous data intelligent matching system applied to the online platform of scientific and technological supply chain, comprising:
[0065] The data preprocessing module is used to standardize the multi-source heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform and generate data quality scores;
[0066] The cross-modal semantic understanding module is used to receive user queries, perform unified semantic feature representation on user queries and standardized scientific and technological resource data, and identify user query intent;
[0067] The intelligent matching engine module is used to calculate the similarity between the user's query intent and its semantic feature representation and the semantic feature representation of the scientific and technological resource data, and generate a matching score based on the data quality score and output a matching list;
[0068] The feedback optimization module is used to collect user feedback data on the matching list, convert the feedback data into real-time feedback features and feed them back to the intelligent matching engine module, recalculate the matching score and update the matching list until the preset stop condition is met.
[0069] Specifically, in one embodiment of the present invention, the data preprocessing module includes:
[0070] Data structure parser, used to perform structured analysis on multi-source heterogeneous scientific and technological resource data, identify and extract key fields and metadata;
[0071] Content normalization processor, used to perform format standardization, text cleaning and normalization on the parsed data;
[0072] Entity linking and disambiguation tools, which link and disambiguate professional terms, institution names, and technical concepts in the data;
[0073] A multi-dimensional quality assessor is used to perform a multi-dimensional assessment of data quality based on data completeness, consistency, timeliness, and authority, and generate a data quality score.
[0074] The overall process of data preprocessing is as follows: first, the data structure parser performs structured parsing on the input heterogeneous data, identifies and extracts key fields and metadata; then the content normalization processor performs format standardization, text cleaning and normalization on the parsed data; then the entity linking and disambiguation processor links and disambiguates the professional terms, institution names and technical concepts in the data; finally, the multi-dimensional quality assessor evaluates the data quality based on dimensions such as data integrity, consistency, timeliness and authority, and generates a comprehensive quality score.
[0075] In this embodiment, the data structure parser is used to automatically identify and extract key information fields of different types of scientific and technological data. For patent data, the system parses and extracts key information such as technical field classification, core technical points, and application scenarios; for paper data, it extracts research methods, technical innovations, and other content; for expert data, it identifies information such as professional fields and representative achievements; for corporate needs, it understands elements such as demand types and technical indicators. The data structure parser first identifies the format of the input data to determine whether it is structured data, semi-structured data, or unstructured data; then applies the corresponding parsing strategy based on the data type. For structured data, it extracts fields directly based on its structure; for semi-structured data, it uses regular expressions, DOM parsing, or PDF parsing tools to extract key information; for unstructured data, it uses natural language processing technology to identify key entities and relationships. Finally, the parsing results are converted into a unified internal data representation format, which contains information such as original data, extracted key fields, metadata, and data sources.
[0076] The content normalization processor standardizes the extracted key information fields, including date format unification, organization name standardization, and technical terminology mapping. This includes: first, text cleaning to remove interfering elements such as HTML tags, special characters, extra spaces, and line breaks; then, punctuation and formatting unification, converting full-width characters to half-width characters and standardizing the formats of dates, times, and numbers; then, organization names are standardized to resolve homonymous or heteronymous homonymous issues; and finally, technical terminology mapping, mapping the same technical concepts in different expressions to standard terms in a predefined vocabulary.
[0077] The entity linking and disambiguation component is used to link entities in key information fields to a unified entity library and eliminate ambiguity. This component links entities such as authors and institutions mentioned in the text to the platform's unified entity library to eliminate ambiguity. The entity linking and disambiguation process mainly includes three steps: first, the entity mentions in the text are identified through named entity recognition (NER) technology, including names of people, institution names, technical terms, etc.; second, a list of candidate entities is generated for the identified entity mentions, that is, possible matching entities are retrieved from the unified entity library; finally, the entity disambiguation algorithm is used to select the most matching entity from the candidate entities for linking.
[0078] The multi-dimensional quality evaluator is used to evaluate data quality based on timeliness, authority, completeness, and accuracy. The evaluator establishes a comprehensive evaluation system that includes multiple dimensions. The multi-dimensional quality evaluator uses the following formula to calculate the data quality score:
[0079]
[0080] Among them, Q multi (r) is the comprehensive quality score of data r, K is the number of quality assessment dimensions, q k (r) is the original quality score of data r in the kth dimension (after normalization, such as mapping to the [0,1] interval), w k is the importance weight of the k-th dimension, σ(·) is the Sigmoid activation function, α k is the adjustment parameter, Ω k (type(r)) is the weight adjustment factor based on the data type. multi (r), it is necessary to perform normalization and map it to the [0,1] interval to facilitate subsequent calculations.
[0081] The original quality score q of each dimension k (r) The calculation method is as follows:
[0082] Timeliness evaluation: Based on the data release time age(r) (such as patent application date, paper publication date), the exponential decay function is used for calculation, for example where λ time It is a time decay coefficient specific to the timeliness dimension. Authority assessment: a comprehensive assessment, such as the number of citations of the patent, the size of the patent family, and the authorization status; the impact factor and citation frequency of the journal in which the paper is published; the academic title, project experience, and awards of the expert; the certification level of the publisher of the enterprise demand, etc. This information is weighted or predicted by the model to obtain the original score. Completeness assessment: evaluates the fill rate of key fields and the level of information detail. For example, whether the word count of the patent abstract and claims meets the standards, and whether the demand description is clear and specific. Accuracy assessment: evaluation is carried out through cross-validation with other data sources, consistency between text content and declared fields, etc.
[0083] The data type weight adjustment factor is defined as:
[0084]
[0085] Among them, ω patent,k 、ω paper,k 、ω expert,k 、ω demand,k The kth quality dimension is defined as the type-specific basic weights of patents, papers, experts, and demand data, respectively. These weights are set by domain experts or optimized through historical data analysis.
[0086] Specifically, in one embodiment of the present invention, the cross-modal semantic understanding module includes:
[0087] A semantic understanding model for the science and technology innovation field, which is used to perform deep semantic encoding on user queries and standardized science and technology innovation resource data, and generate text semantic features;
[0088] A cross-modal information fusion model is used to deeply fuse text semantic features with structured information to form a cross-modal feature representation. For scientific and technological resource data, its text semantic features and structured information are integrated; for user queries, its text semantic features and query context information are integrated.
[0089] The user intent recognition model is used to analyze user queries and context information and identify the user's query intent type.
[0090] In this embodiment, the semantic understanding model for the science and technology field is based on a pre-trained Transformer architecture (such as BERT or RoBERTa), and is fine-tuned for domain adaptation on professional corpus in the science and technology field to enhance the ability to understand professional terms and expressions in science and technology. The main implementation steps of this model include:
[0091] First, build a special vocabulary database and concept map in the field of science and technology innovation, which includes technical terms, research institutions, subject classifications, technology classifications and other field knowledge.
[0092] The pre-trained Transformer model was then fine-tuned for domain adaptability. This fine-tuning process utilized masked language model tasks and next sentence prediction tasks to adapt to the language characteristics and expression conventions of scientific and technological innovation texts. For patent documents, the model was trained to understand the expression of claims and technical specifications. For academic papers, the model's understanding of research methods and experimental results was enhanced. For expert data, the model's grasp of research directions and results descriptions was strengthened. For enterprise needs, the model's understanding of technical requirements and performance indicators was improved.
[0093] Finally, the fine-tuned model performs deep semantic encoding on the text content of user query text and scientific and technological resource data, outputting high-dimensional semantic feature vectors. For long texts (such as patent specifications and full-text academic papers), a segmented encoding and hierarchical fusion strategy is adopted. The text is first encoded by paragraph, and then the semantic representations of each paragraph are integrated through an attention mechanism to generate a holistic text semantic feature.
[0094] Since scientific and technological resource data and user queries differ in structure and characteristics, the cross-modal information fusion model designs different fusion strategies for the two.
[0095] For scientific and technological resource data, the cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation:
[0096] F unified (r)=LayerNorm(W text F text (r)+W meta F meta (r)+W struct F struct (r)+C cross (r))
[0097] Among them, r represents scientific and technological innovation resource data, F unified (r) is the cross-modal feature representation of r, F text (r) is the text semantic feature of r (generated by the semantic understanding model of science and technology), G meta (r) is the metadata feature of r (such as author, institution, publication time, etc.), F struct (r) is the structural feature of r (such as technology classification, citation relationship, etc.), W text 、W meta 、W struct is the corresponding weight matrix, LayerNorm(·) is the layer normalization function, C cross(r) is the cross-modal interaction term.
[0098] Specifically, F meta (r) Extraction through a combination of feature engineering and embedding learning: For categorical metadata, embedding tables are used to convert them into low-dimensional dense vectors; for numerical metadata, they are directly used after normalization; for textual metadata, they are first converted into vectors through embedding tables and then pooled to obtain fixed-dimensional representations. struct (r) Mainly includes the location information of scientific and technological innovation resources in the technology classification system, citation relationship network characteristics, etc.: the technology classification system characteristics are obtained through multi-level classification coding, first converting each level into an embedding vector, and then integrating the representations of each level through the attention mechanism; the citation relationship network characteristics are extracted through graph neural networks, and scientific and technological innovation resources are constructed as a citation relationship graph, and the structural information of resources in academic or technical networks is captured through the message passing mechanism.
[0099] Cross-modal interaction term C cross (r) The multi-head attention mechanism is used to calculate the relationship between different modal features and capture the semantic dependency and complementary information between modalities. The calculation formula is:
[0100]
[0101] Among them, W i→j is the interaction weight from mode i to mode j, ⊙ is the element-wise product, Attention(F i (r),F j (r)) is the cross-modal attention mechanism, which calculates the weighted combination of the attention weight and features of modality i to modality j.
[0102] For user queries, the cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation:
[0103] F unified (q) = LayerNorm(W′ text F text (q)+W′ context F context (q)+C′ cross (q))
[0104] Among them, q represents the user query, F unified (q) is the cross-modal feature representation of q, F text (q) is the text semantic feature of q, F context (q) is the query context feature of q (including user preferences, query history, filter conditions, etc.), W′ text and W′ context is the corresponding weight matrix, C′ cross(q) is the feature interaction term within the query, which is used to establish the association between the query text and the context. Its calculation formula is the same as C cross (r) Similar.
[0105] Specifically, F context (q) includes the following components: user preference features, which reflect the user's historical interests and behavior patterns; session state features, which describe the position and context of the current query in the entire interaction process; filter condition features, including the filter conditions explicitly set by the user; and device environment features, which describe the environmental information of the user's query.
[0106] In this embodiment, the user intent recognition model uses the following formula to calculate the probability distribution of the user's query intent:
[0107]
[0108] Among them, P(intent c |q) is the probability that query q belongs to the cth type of intent, F unified (q) is the unified feature representation of query q (generated by the cross-modal information fusion model), w c is the weight vector of the c-th type of intention, b c is the bias term, and C is the total number of intent categories. Intent categories include technical solution query intent, expert query intent, funding query intent, policy query intent, cooperation query intent, and equipment query intent.
[0109] Specifically, in one embodiment of the present invention, the intelligent matching engine module includes:
[0110] Intent-aware matching weight configurator, used to adjust the weight distribution of multi-dimensional similarity calculation based on the identified user query intent;
[0111] A multi-dimensional similarity calculator is used to calculate the multi-dimensional similarity between user queries and scientific and technological resource data based on semantic similarity, domain relevance, technical level similarity, and application scenario matching, combined with weight distribution;
[0112] A personalized regulator, used to generate a personalized adjustment factor based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics;
[0113] Matching score calculator, which integrates multi-dimensional similarity, data quality score and personalized adjustment factors to generate a matching score;
[0114] Candidate result ranking optimizer, used to generate matching lists based on matching scores.
[0115] In this embodiment, the intent-aware matching weight configurator assigns corresponding weights to different similarity dimensions based on the user query intent type identified by the cross-modal semantic understanding module.
[0116] For different query intent types, the configurator will adopt different weight allocation strategies: for technical solution query intent, more emphasis will be placed on semantic similarity and technical level similarity; for expert query intent, more emphasis will be placed on professional background matching and results influence; for funding query intent, more emphasis will be placed on project stage matching and funding scale adaptability; for policy query intent, more emphasis will be placed on policy applicability conditions and timeliness.
[0117] The intent-aware matching weight configurator dynamically calculates the weight of similarity in each dimension based on the user's query intent and the current context. The calculation formula is:
[0118]
[0119] Among them, φ d (intent(q)) is the d-th dimension similarity weight based on query intent, intent(q) is the intent type of user query q, e d is the embedding representation of the d-th dimension feature, c intent is the intent context vector, W intent 、W dim 、W ctx is the weight matrix, b is the bias vector, softmax(·) d It means applying the softmax function to the vector and obtaining the value of the d-th dimension.
[0120] In this embodiment, the multi-dimensional similarity calculator uses a combination of multiple algorithms to accurately calculate the similarity of different dimensions. The similarity dimensions include semantic similarity: measuring the similarity between user queries and scientific and technological resources at the semantic level, based on the unified feature representation calculation generated by the cross-modal semantic understanding module, capturing the matching degree between queries and resources in terms of concepts and content. Domain relevance: evaluating the correlation between user queries and scientific and technological resources in subject areas and technical classifications, considering the position of scientific and technological resources in the technical classification system and the degree of matching with the domain of user queries. Technical level similarity: comparing the matching degree between user queries and scientific and technological resources in terms of technical maturity, technical depth, etc., which is particularly suitable for matching patents and technical solutions. Application scenario matching: evaluating the degree of fit between the application requirements described in the user query and the applicable scenarios of scientific and technological resources, focusing on the actual application value and applicability of the technology.
[0121] For the similarity calculation of each dimension, a combination of multiple measurement methods is used:
[0122]
[0123] Among them, Sim d (F unified (r),F unified (q)) is the similarity between q and r in the dth dimension, cos(·,·) is the cosine similarity, is the Gaussian kernel similarity, Jaccard(K q ,K r ) is the Jaccard similarity of the keyword set, α d , β d , γ d is the combined weight, σ d is the bandwidth parameter of the Gaussian kernel.
[0124] In this embodiment, the personalized adjuster generates a customized adjustment factor for each user and query scenario by analyzing the user's historical interaction data, current query context, and real-time feedback information, thereby further improving the personalization of the matching results.
[0125] The calculation formula of the personalized adjustment factor is:
[0126]
[0127] Among them, Ψ(u,r) is the personalized adjustment factor of user u to resource r, h ih (u, r) is the user's historical behavior characteristics, which includes the user's past interaction information with resource r or its similar resources, that is, historical feedback data, including whether the user has browsed, clicked, or collected related resources. qic (q, r) is the current context feature, which describes the specific relationship between the current query q and session context and resource r, such as whether the resource meets the filtering conditions in the query and the occurrence of query keywords in the resource. is the real-time feedback item, w rt is the user feedback weight, f rt is the real-time feedback feature. ih and w qic is the weight vector, b user is the bias term. tanh(·) is the activation function.
[0128] In this embodiment, the matching score calculator calculates the comprehensive matching degree between the user query and the scientific and technological innovation resources.
[0129] The matching score is calculated as follows:
[0130]
[0131] Among them, M intent (q, r) is the matching score between q and r, and D is the number of similarity calculation dimensions.
[0132] In this embodiment, the candidate result ranking optimizer first ranks the candidate results according to the matching scores, and then applies a diversity optimization algorithm to ensure that the final list contains scientific and technological resources of different types and sources to avoid the results being too single.
[0133] Specifically, in one embodiment of the present invention, the feedback optimization module includes:
[0134] User feedback collector, used to collect user feedback data on matching results, including explicit feedback and implicit feedback;
[0135] A real-time feedback feature converter, used to convert the collected feedback data into standardized real-time feedback features;
[0136] Feedback transmitter, used to feed back real-time feedback features to the intelligent matching engine module, triggering the recalculation of matching scores and the update of matching lists;
[0137] The stopping condition detector is used to monitor whether the preset stopping condition is met and determine whether the feedback optimization process is terminated.
[0138] The feedback optimization module also includes an online learner, which uses an online learning mechanism to regularly update model parameters based on user feedback and new data.
[0139] In this embodiment, user behavior is divided into two categories: explicit feedback and implicit feedback. Explicit feedback includes user ratings of recommended results on a scale of 1-5, user actions such as actively saving results of interest, user sharing of recommended content, and user reporting of irrelevant content. Implicit feedback is inferred through users' natural interactive behavior, including user clicks on recommended results, user stay time on the recommended results page, the level of detail users review the recommended content, and user exit behavior (e.g., quickly leaving the recommended results).
[0140] The conversion strategies of the real-time feedback feature converter for explicit feedback data include:
[0141] Rating data conversion: linearly map the ratings from 1 to 5 to the interval [-1, 1]. The conversion formula is: Among them, f score The converted score feature value ranges from [-1, 1], where -1 indicates extreme dissatisfaction and 1 indicates extreme satisfaction.
[0142] Binary behavior conversion: For binary behaviors such as collecting, sharing, and reporting, predefined weight values are assigned: the collection action is given a weight of +0.3, indicating that the user has clear interest; the sharing action is given a weight of +0.25, indicating that the user thinks the content is valuable; the reporting action is given a weight of -0.5, indicating that the user is strongly dissatisfied.
[0143] The conversion strategies for implicit feedback data include: assigning a weight of +0.1 to click behavior, indicating the user's basic attention; stay time conversion: assigning weights according to the user's stay time on the content page: a long stay (more than 30 seconds) is assigned a weight of +0.2, indicating the user's deep interest; a medium stay (10-30 seconds) is assigned a weight of +0.1, indicating general interest; a short stay (5-10 seconds) is assigned a weight of 0, indicating a neutral evaluation; a quick exit (less than 5 seconds) is assigned a weight of -0.1, indicating possible dissatisfaction; browsing depth conversion: assigning weights according to the completeness of the user's browsing content: a complete browse (viewing the full text or all details) is assigned a weight of +0.15; a partial browse (viewing part of the content) is assigned a weight of +0.05.
[0144] Based on the above conversion strategy, the real-time feedback feature converter integrates all feedback behaviors in the current session into comprehensive real-time feedback features:
[0145]
[0146] in, Indicates the weights of different behaviors, Represents the characteristic value of a single behavior, N actions Indicates the total number of actions in the current session.
[0147] When a user submits a query, the system first initializes f rt = 0, indicating that there is no feedback information in the current session. The system then returns the initial matching list calculated based on historical data. When the user generates any feedback behavior, the system immediately updates f rt The personalized adjustment factor is recalculated based on the value of , and the matching score is updated. Finally, the matching list is dynamically rearranged and the optimized result is pushed to the user.
[0148] The feedback transmitter adopts an event-driven mechanism. When the real-time feedback feature f rt When a change occurs, the following processing flow is triggered immediately:
[0149] The updated real-time feedback feature f rt The personalized regulator component is passed to the intelligent matching engine module; the personalized regulator receives real-time feedback features and recalculates the personalized adjustment factor Ψ(u,r); the matching score calculator recalculates the matching scores of all candidate scientific and technological resources based on the updated personalized adjustment factor; the candidate result ranking optimizer regenerates the optimized matching list according to the new matching score; the system pushes the updated matching list to the user, replacing the original results.
[0150] To improve system performance, the feedback transmitter adopts an incremental update strategy, recalculating only the candidate results relevant to the current query. It also utilizes a caching mechanism to accelerate the computation process and ensure real-time responsiveness. Furthermore, a feedback update threshold is set. A full recalculation process is triggered only when the change in real-time feedback features exceeds a preset threshold, thus avoiding the computational overhead of frequent, small updates.
[0151] In this embodiment, the preset stopping conditions include reaching the maximum number of iterations: the system sets the maximum number of iterations for feedback optimization. After exceeding this number, further optimization will be stopped regardless of the matching results. User satisfaction is met: when the user expresses clear satisfaction with the matching results, it is considered that the optimization goal has been achieved and further optimization is stopped. Matching score stability: when the change in the matching scores of the top N results in the matching list after two consecutive optimizations is less than the preset threshold, the optimization is considered to have converged and further iterations are stopped. User termination interaction: when the user actively terminates the current query session, the current feedback optimization process is stopped. Time window limit: set the maximum optimization time window for a single query. After exceeding this time, further optimization is stopped to avoid long-term occupation of system resources.
[0152] The stop condition detector monitors the above conditions and, when any of the conditions is met, notifies the feedback transmitter to stop sending update requests to the intelligent matching engine, thus completing the feedback optimization process for the current query.
[0153] In this embodiment, the online learner uses the feedback data of the current users to make real-time adjustments. It also accumulates a certain amount of feedback data and regularly updates the parameters of various models of the system to achieve continuous evolution of system performance.
[0154] The online learner updates the model parameters using the following formula:
[0155]
[0156] Among them, Θ (t+1) is the updated model parameter, Θ (t) is the current model parameter, α is the learning rate, is the user feedback data stream collected at time t (including query q, recommendation result z and user feedback label f), Lranking(·,·) is the ranking loss function, M(Θ (t) ,q,z) is the matching score predicted by the intelligent matching model, and f is the true feedback label. Here, the ranking loss function can use RankNet Loss or ListMLE.
[0157] This formula is used to update all learnable parameters of each sub-model in the previous three modules.
[0158] like Figure 3As shown, the present invention also provides a multi-source heterogeneous data intelligent matching method applied to the scientific and technological supply chain online platform, characterized in that the method is used to implement the system as described in any one of the above items, and the method includes:
[0159] S1. Perform structured analysis, content normalization, entity linking and disambiguation processing, and multi-dimensional quality assessment on the multi-source and heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform to generate standardized scientific and technological resource data and its quality score;
[0160] S2. Receive user queries, use cross-modal semantic understanding methods to create unified semantic feature representations for the user queries and standardized scientific and technological resource data, and identify the user's query intent.
[0161] S3. Based on the user's query intent and its semantic feature representation, a multi-dimensional similarity calculation is performed with the semantic feature representation of the scientific and technological resource data, and a matching score is generated by combining the data quality score and a personalized adjustment factor. The personalized adjustment factor is calculated based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics.
[0162] S4. Generate a matching list based on the matching score and push the matching list to the user;
[0163] S5, monitoring whether the preset stop condition is met, if so, ending the matching process, if not, executing step S6;
[0164] S6. Collect user feedback data on the matching list, including explicit feedback and implicit feedback;
[0165] S7: Convert the collected feedback data into standardized real-time feedback features and feed them back to step S3.
[0166] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-source heterogeneous data intelligent matching system applied to the online platform of scientific and technological supply chain, characterized by: include: The data preprocessing module is used to standardize the multi-source heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform and generate data quality scores; The cross-modal semantic understanding module is used to receive user queries, perform unified semantic feature representation on user queries and standardized scientific and technological resource data, and identify user query intent; The intelligent matching engine module is used to calculate the similarity between the user's query intent and its semantic feature representation and the semantic feature representation of the scientific and technological resource data, and generate a matching score based on the data quality score and output a matching list; The feedback optimization module is used to collect user feedback data on the matching list, convert the feedback data into real-time feedback features and feed them back to the intelligent matching engine module, recalculate the matching score and update the matching list until the preset stop condition is met.
2. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 1 is characterized in that: The data preprocessing module includes: Data structure parser, used to perform structured analysis on multi-source heterogeneous scientific and technological resource data, identify and extract key fields and metadata; Content normalization processor, used to perform format standardization, text cleaning and normalization on the parsed data; Entity linking and disambiguation tools, which link and disambiguate professional terms, institution names, and technical concepts in the data; A multi-dimensional quality assessor is used to perform a multi-dimensional assessment of data quality based on data completeness, consistency, timeliness, and authority, and generate a data quality score.
3. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 2 is characterized in that: The multidimensional quality evaluator uses the following formula to calculate the data quality score: Among them, Q multi (r) is the comprehensive quality score of data r, K is the number of quality assessment dimensions, q k (r) is the original quality score of data r in the kth dimension, w k is the importance weight of the k-th dimension, σ(·) is the Sigmoid activation function, α k is the adjustment parameter, Ω k (type(d)) is the weight adjustment factor based on the data type.
4. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 1 is characterized in that: The cross-modal semantic understanding module includes: A semantic understanding model for the science and technology innovation field, which is used to perform deep semantic encoding on user queries and standardized science and technology innovation resource data, and generate text semantic features; A cross-modal information fusion model is used to deeply fuse text semantic features with structured information to form a cross-modal feature representation. For scientific and technological resource data, its text semantic features and structured information are integrated; for user queries, its text semantic features and query context information are integrated. The user intent recognition model is used to analyze user queries and context information and identify the user's query intent type.
5. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 4 is characterized in that: The cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation of scientific and technological resource data: F unified (r) =LayerNorm(W text F text (r)+W meta F meta (r)+W struct F struct (r)+C cross (r)) Among them, r represents scientific and technological innovation resource data, F unified (r) is the cross-modal feature representation of r, F text (r) is the text semantic feature of r, F meta (r) is the metadata feature of r, F struct (r) is the structural feature of r, W text 、W meta 、W struct is the corresponding weight matrix, LayerNorm(·) is the layer normalization function, C cross (r) is the cross-modal interaction term; The cross-modal information fusion model uses the following formula to calculate the cross-modal feature representation of user queries: F unified (q)=LayerNorm(W′ text F text (q)+W′ context F context (q)+C′ cross (q)) Among them, q represents the user query, F unified (q) is the cross-modal feature representation of q, F text (q) is the text semantic feature of q, F context (q) is the query context feature of q, W′ text and W′ context is the corresponding weight matrix, C′ cross (q) is the feature interaction term within the query.
6. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 1 is characterized in that: The intelligent matching engine module includes: Intent-aware matching weight configurator, used to adjust the weight distribution of multi-dimensional similarity calculation based on the identified user query intent; A multi-dimensional similarity calculator is used to calculate the multi-dimensional similarity between user queries and scientific and technological resource data based on semantic similarity, domain relevance, technical level similarity, and application scenario matching, combined with weight distribution; A personalized regulator, used to generate a personalized adjustment factor based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics; Matching score calculator, which integrates multi-dimensional similarity, data quality score and personalized adjustment factors to generate a matching score; Candidate result ranking optimizer, used to generate matching lists based on matching scores.
7. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 6 is characterized in that: The matching score is calculated as follows: Among them, q represents user query, r represents scientific and technological resource data, and M intent (q, r) is the matching score between q and r, D is the number of similarity calculation dimensions, intent(q) is the user query intent type of q, φ d (intent(q)) is the d-th dimension similarity weight based on query intent, F unified (r) is the cross-modal feature representation of r, F unified (q) is the cross-modal feature representation of q, Sim d (F unified (r),F unified (q)) is the similarity between q and r in the dth dimension, Q multi (r) is the data quality score of r, and Ψ(u,r) is the personalized adjustment factor of user u for resource r.
8. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 1 is characterized in that: The feedback optimization module includes: User feedback collector, used to collect user feedback data on matching results, including explicit feedback and implicit feedback; A real-time feedback feature converter, used to convert the collected feedback data into standardized real-time feedback features; Feedback transmitter, used to feed back real-time feedback features to the intelligent matching engine module, triggering the recalculation of matching scores and the update of matching lists; The stopping condition detector is used to monitor whether the preset stopping condition is met and determine whether the feedback optimization process is terminated.
9. The multi-source heterogeneous data intelligent matching system applied to the online platform of the science and technology supply chain according to claim 8 is characterized in that: The feedback optimization module also includes an online learner, which uses an online learning mechanism to regularly update model parameters based on user feedback and new data.
10. A multi-source heterogeneous data intelligent matching method applied to the online platform of scientific and technological supply chain, characterized by: The method is used to implement the system according to any one of claims 1 to 9, and the method includes: S1. Perform structured analysis, content normalization, entity linking and disambiguation processing, and multi-dimensional quality assessment on the multi-source and heterogeneous scientific and technological resource data on the scientific and technological innovation supply chain platform to generate standardized scientific and technological resource data and its quality score; S2. Receive user queries, use cross-modal semantic understanding methods to create unified semantic feature representations for the user queries and standardized scientific and technological resource data, and identify the user's query intent. S3. Based on the user's query intent and its semantic feature representation, a multi-dimensional similarity calculation is performed with the semantic feature representation of the scientific and technological resource data, and a matching score is generated by combining the data quality score and a personalized adjustment factor. The personalized adjustment factor is calculated based on the user's historical behavior characteristics, current context characteristics, and real-time feedback characteristics. S4. Generate a matching list based on the matching score and push the matching list to the user; S5, monitoring whether the preset stop condition is met, if so, ending the matching process, if not, executing step S6; S6. Collect user feedback data on the matching list, including explicit feedback and implicit feedback; S7: Convert the collected feedback data into standardized real-time feedback features and feed them back to step S3.
Citation Information
Patent Citations
Large model voice question-answering system for multi-source heterogeneous local knowledge base
CN118193694A
Intelligent data retrieval optimization method based on large model and deep learning
CN119377281A
Multi-dimensional police data intelligent search system based on NLP semantic analysis
CN119739906A
Government affair information intelligent retrieval and generation system applying RAG technology
CN119807261A
SQL (Structured Query Language) statement generation method based on large-model multi-stage iteration
CN120086244A
Cited By
Multi-modal data quality evaluation method based on deep learning
CN120804084A
Adaptive semantic-driven data set field matching method and system
CN121167326A