Information matching method based on multi-source heterogeneous data fusion
By constructing a resource knowledge graph and training a demand matching model, and by adjusting user adoption rate, push delay time, and demand deviation rate, the problem of insufficient accuracy of matching results in multi-source heterogeneous data fusion is solved, and more efficient information matching is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN INSTITUTE OF SCIENCE & TECHNOLOGY INFORMATION (HUNAN INSTITUTE OF SOFT SCIENCE)
- Filing Date
- 2026-03-09
- Publication Date
- 2026-06-05
AI Technical Summary
Existing information matching technologies lack effective repair basis in the fusion of multi-source heterogeneous data, resulting in unstable repair results for missing data, difficulty in reflecting the true relationship between data, and insufficient accuracy in generating matching results.
By collecting multi-source resource data and performing cleaning, noise reduction, information extraction, and entity matching, a resource knowledge graph is constructed, resource attribute features are extracted, a demand matching model is trained, and the matching results are adjusted based on user adoption rate, push delay time, and demand deviation rate to optimize the matching process.
This improved the accuracy and real-time performance of matching results, enhanced the system's responsiveness to dynamic needs, and increased user satisfaction and the usability of the matching results.
Smart Images

Figure CN122153485A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information matching technology, and in particular to an information matching method based on the fusion of multi-source heterogeneous data. Background Technology
[0002] In fields such as government-enterprise cooperation and scientific research innovation, information matching methods based on multi-source heterogeneous data fusion have become a core technology for integrating scattered information resources and improving service accuracy and response efficiency. Since the quality of matching user needs with available resources directly affects the efficiency of government processing, enterprise decision-making, and scientific research collaboration, the depth of data fusion and dynamic matching capabilities of this method are crucial. When a matching system relies solely on a single data source or static rules, the lack of information dimensions can lead to recommendations deviating from actual needs, or data updates may lag and fail to reflect the true state of resources. Furthermore, the lack of continuous learning of user behavior and preferences can limit the accuracy of personalized matching, and a rigid response mechanism cannot adapt to real-time changes in application scenarios, leading to risks such as decreased service satisfaction and lost opportunities. While existing information matching technologies have achieved basic keyword retrieval and simple recommendations, they often employ limited data sources or fixed matching models, failing to fully address core pain points such as difficulties in aligning multi-source heterogeneous data, dynamic recognition of user intent, and the inability to balance real-time response with long-term personalization. Moreover, their ability to mine implicit needs and cross-domain resource associations is insufficient, making it difficult to achieve highly accurate, adaptable, and efficient intelligent resource matching in complex application scenarios. Therefore, information matching methods based on multi-source heterogeneous data fusion need to overcome the limitations of traditional technologies and achieve the coordinated operation of extensive data collection, in-depth demand perception, dynamic model optimization, and real-time accurate push.
[0003] Chinese Patent Publication No. CN120296659A discloses a multi-source data fusion processing method and system. The method includes: 1. Data collection: collecting the latest data through data collection equipment; 2. Data classification and labeling: matching and classifying the data to generate updated data; 3. Data fusion and preprocessing: developing efficient data preprocessing methods, including data cleaning, normalization, and missing value handling; 4. Data visualization and dynamic correlation analysis: developing dynamic data visualization tools to intuitively display the multi-dimensional and multi-modal characteristics of the data, including visualization of real-time data streams; 5. Fusion data export: exporting data and processing it to generate fused data; 6. Data matching and missing data labeling: matching the updated data with the fused data, labeling the data differences, and processing them to generate repaired data; 7. Integrated platform and system function testing: integrating all developed modules into a unified system platform to ensure seamless connection and smooth data transmission. It is evident that the aforementioned multi-source data fusion processing method and system suffer from problems due to the lack of effective repair criteria based on data source consistency, resulting in unstable missing data repair results, difficulty in reflecting the true correlation between data, and insufficient accuracy in generating matching results. Summary of the Invention
[0004] To address this, the present invention provides an information matching method based on multi-source heterogeneous data fusion, which overcomes the problem in the prior art that the lack of effective repair basis based on data source consistency leads to unstable missing data repair results, difficulty in reflecting the true correlation between data, and insufficient accuracy in generating matching results.
[0005] To achieve the above objectives, the present invention provides an information matching method based on multi-source heterogeneous data fusion, comprising:
[0006] Collect multi-source resource data from several data sources, and sequentially clean, denoise, extract information, match entities, and fuse the multi-source resource data to construct a resource knowledge graph. Then, extract features from the resource knowledge graph to output resource attribute features.
[0007] The initial model is trained based on the resource attribute features and user needs to obtain a demand matching model, and a matching result is generated based on the demand matching model and historical user needs.
[0008] The matching results are optimized and adjusted based on real-time user needs to obtain optimized matching results, and the optimized matching results and their corresponding multi-source resource data are pushed to the user terminal.
[0009] Obtain the user acceptance rate of the matching results, and determine whether the accuracy of the generated matching results meets the requirements based on the user acceptance rate of the matching results;
[0010] If the accuracy of the generated matching result does not meet the requirements, the real-time performance of the matching result adjustment is determined by the push delay time of the multi-source resource data.
[0011] If the real-time performance of the matching result adjustment does not meet the requirements, then the discrete judgment threshold for the matching result adjustment is determined based on the user demand matching deviation rate.
[0012] Furthermore, the accuracy of the generated matching results is determined based on the user acceptance rate, including:
[0013] Compare the user acceptance rate of the matching results with the preset acceptance rate;
[0014] If the user acceptance rate of the matching result is greater than the preset acceptance rate, it is determined that the accuracy of the generated matching result meets the requirements.
[0015] Furthermore, if the user acceptance rate of the matching result is less than or equal to the preset acceptance rate, it is determined that the accuracy of the generated matching result does not meet the requirements.
[0016] Furthermore, if the accuracy of the generated matching results does not meet the requirements, the real-time performance of the matching result adjustment is determined based on the push delay time of multi-source resource data.
[0017] Furthermore, the real-time performance of matching result adjustments is determined based on the push delay duration of multi-source resource data, including:
[0018] The push delay time of the multi-source resource data is compared with the preset delay time;
[0019] If the push delay of the multi-source resource data is less than or equal to the preset delay, the real-time performance of the matching result adjustment is determined to meet the requirements.
[0020] Furthermore, if the push delay of the multi-source resource data is longer than the preset delay, it is determined that the real-time performance of the matching result adjustment does not meet the requirements.
[0021] Furthermore, if the real-time performance of the matching result adjustment does not meet the requirements, the sensitivity of real-time demand recognition is determined based on the user demand matching deviation rate.
[0022] Furthermore, based on the user demand matching deviation rate, a discrete decision threshold for adjusting the matching result is determined, including:
[0023] Compare the user demand matching deviation rate with the preset deviation rate;
[0024] If the user demand matching deviation rate is less than or equal to the preset deviation rate, it is determined that the real-time demand recognition sensitivity meets the requirements.
[0025] Furthermore, if the user demand matching deviation rate is greater than the preset deviation rate, it is determined that the recognition sensitivity of the real-time demand does not meet the requirements, and the discrete judgment threshold for adjusting the matching result is reduced.
[0026] Furthermore, the reduction in the discrete judgment threshold of the matching result adjustment is determined by the difference between the user demand matching deviation rate and the preset deviation rate.
[0027] Compared with existing technologies, the beneficial effects of this invention are as follows: The method of this invention determines the accuracy of the matching results based on the user acceptance rate of the matching results. Because user needs are characterized by sudden and discrete changes, and the matching model's response to real-time demand changes is lagging, the generated matching results deviate from the user's current actual needs. This reduces the user's acceptance rate of the matching results, failing to accurately reflect the true effectiveness of the matching results, and consequently affecting the system's assessment and response to the degree of demand satisfaction. By determining the accuracy of the matching results, user satisfaction can be transformed into a quantifiable evaluation indicator, promptly identifying matching distortion problems caused by algorithm defects or data quality issues, and triggering targeted optimization of the model training strategy or algorithm. Furthermore, the method determines the real-time nature of the matching result adjustment based on the push delay of multi-source resource data. Due to limitations in multi-source data synchronization efficiency and network transmission, the end-to-end delay from data update to matching result generation and push to the user is too high, affecting the system's response to dynamic demands and resource needs. The rapid response capability to changes in the source cannot obtain effective matching results within the time window. By determining the real-time nature of the matching result adjustment, the system response performance can be transformed into a quantifiable timeliness indicator. This allows for timely identification of response delays caused by data synchronization lags or improper resource scheduling, triggering targeted optimizations to the data stream processing architecture or dynamic priority scheduling. The discrete judgment threshold for matching result adjustment can be adjusted based on the user demand matching deviation rate. Because model updates lag behind dynamically changing user demands and resource states, the model cannot capture effective feature changes in the real-time data stream in a timely manner, leading to a systematic deviation between the model output and the user's true intent, causing significant failure of the matching results. By reducing the discrete judgment threshold for matching result adjustment, the magnitude of demand changes required to trigger the matching result adjustment can be reduced, allowing the model to trigger matching result updates when it detects small changes in demand or resource states. This improves the sensitivity of matching result adjustment to real-time demand changes and enhances the accuracy of matching result generation.
[0028] Furthermore, the method described in this invention determines the accuracy of the generated matching results by setting a preset adoption rate. Since user needs are characterized by sudden and discrete changes, and the matching model's response to real-time changes in needs is lagging, the generated matching results deviate from the user's current actual needs. This reduces the user's adoption rate of the matching results, failing to accurately reflect the true effectiveness of the matching results, and consequently affecting the system's assessment and response to the degree of need fulfillment. By determining the accuracy of the matching results, user satisfaction can be transformed into a quantifiable evaluation indicator, promptly identifying matching distortion problems caused by algorithm defects or data quality issues, and triggering targeted optimization of the model training strategy or algorithm, further improving the accuracy of the generated matching results.
[0029] Furthermore, the method described in this invention determines the real-time performance of matching result adjustments by setting a preset delay duration. Due to limitations in multi-source data synchronization efficiency and network transmission, the end-to-end delay from data update to matching result generation and push to the user is too high, affecting the system's ability to respond quickly to dynamic demands and resource changes, and making it impossible to obtain effective matching results within the time window. By determining the real-time performance of matching result adjustments, the system response performance can be transformed into a quantifiable timeliness indicator, promptly identifying response delay issues caused by data synchronization lag or improper resource scheduling, triggering targeted optimizations to the data stream processing architecture or dynamic priority scheduling, and further improving the accuracy of matching result generation.
[0030] Furthermore, the method of the present invention adjusts the minimum number of complete cycles in the processing process by setting a preset delay duration. Since the model update lags behind the dynamically changing user needs and resource status, the model cannot capture the effective feature changes in the real-time data stream in a timely manner, thereby causing a systematic deviation between the model output and the user's true intention, resulting in significant failure of the matching results. By reducing the discrete judgment threshold for adjusting the matching results, the magnitude of the demand change required to trigger the adjustment of the matching results can be reduced, so that the model can trigger the update of the matching results when it detects a small change in demand or resource status, thereby improving the sensitivity of the matching results adjustment to real-time demand changes and further improving the accuracy of the generated matching results. Attached Figure Description
[0031] Figure 1 This is an overall flowchart of the information matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention;
[0032] Figure 2 This is a flowchart illustrating the process of determining whether the accuracy of the generated matching result meets the requirements in the information matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention.
[0033] Figure 3This is a flowchart illustrating the process of determining whether the real-time performance of the information matching method based on multi-source heterogeneous data fusion in an embodiment of the present invention meets the requirements for adjusting the matching result.
[0034] Figure 4 This is a flowchart illustrating the process of determining the discrete judgment threshold for adjusting the matching result in the information matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0036] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0037] Please see Figure 1 As shown, it is an overall flowchart of the information matching method based on multi-source heterogeneous data fusion in an embodiment of the present invention.
[0038] This invention discloses an information matching method based on multi-source heterogeneous data fusion, comprising:
[0039] Collect multi-source resource data from several data sources, and sequentially clean, denoise, extract information, match entities, and fuse the multi-source resource data to construct a resource knowledge graph. Then, extract features from the resource knowledge graph to output resource attribute features.
[0040] The initial model is trained based on the resource attribute features and user needs to obtain a demand matching model, and a matching result is generated based on the demand matching model and historical user needs.
[0041] The matching results are optimized and adjusted based on real-time user needs to obtain optimized matching results, and the optimized matching results and their corresponding multi-source resource data are pushed to the user terminal.
[0042] Obtain the user acceptance rate of the matching results, and determine whether the accuracy of the generated matching results meets the requirements based on the user acceptance rate of the matching results;
[0043] If the accuracy of the generated matching result does not meet the requirements, the real-time performance of the matching result adjustment is determined by the push delay time of the multi-source resource data.
[0044] If the real-time performance of the matching result adjustment does not meet the requirements, then the discrete judgment threshold for the matching result adjustment is determined based on the user demand matching deviation rate.
[0045] Specifically, the data sources include a user address information database, a directory database of testing institutions, and a standard database for construction machinery.
[0046] Specifically, multi-source resource data includes user geographic location information, testing agency certification information, and the types of construction machinery to which the standards apply.
[0047] Specifically, information extraction involves identifying and extracting resource-related information elements from cleaned and denoised multi-source resource data, and then converting these information elements into structured data representations.
[0048] Specifically, entity matching involves aligning and comparing structured data and information elements from different data sources, and associating information elements that point to the same actual resource entity.
[0049] Specifically, a resource knowledge graph is a structured semantic model constructed by processing multi-source resource data to uniformly represent resource entities and their attributes and relationships.
[0050] Specifically, resource attribute characteristics include regional location characteristics, certification coverage completeness, and standard implementation status.
[0051] Specifically, the process of training an initial model based on resource attribute features and user needs to obtain a demand matching model involves using user needs and actual matching results as supervision signals to optimize the parameters of the initial model, enabling the model to learn the correlation between resource attribute features and user needs. During the training process, the model parameters are iteratively updated and validation constraints are introduced to suppress overfitting until the matching results output by the model meet the preset requirements in terms of accuracy and stability, thereby obtaining the demand matching model.
[0052] Specifically, the initial model is a model framework that has ranking capabilities and supports real-time feature fusion and online learning.
[0053] Specifically, the demand matching model can be a dual-tower model, a multi-task learning model, or a vector retrieval model, with the preferred embodiment being the multi-task learning model.
[0054] Specifically, user needs include searching for testing institutions in the user's area, searching for institutions that can carry out specified testing items, and searching for the latest published engineering machinery standards.
[0055] Specifically, the process of generating matching results based on the demand matching model and historical user demands involves converting historical user demands into a demand feature sequence, inputting it along with resource attribute features into the demand matching model, calculating the matching degree between user demands and each candidate resource, sorting the candidate resources based on the matching degree, and generating matching results consistent with historical demand preferences.
[0056] Specifically, the matching results include a list of testing agencies that match geographical locations, a list of suitable agencies, and valid engineering machinery standards that meet timeliness requirements.
[0057] Specifically, the process of optimizing and adjusting the matching results based on real-time user needs to obtain optimized matching results involves the following steps: After obtaining the initial matching results, the user needs collected in real time are characterized to obtain a real-time demand feature vector. The real-time demand feature vector is then compared with the resource attribute features of each candidate resource in the matching results to obtain a real-time matching deviation index. Based on the matching deviation index, the matching results that do not meet the real-time needs are sorted and adjusted or candidates are removed. The matching results that meet the real-time needs are then weighted and enhanced, and finally, the optimized matching results are output.
[0058] Specifically, the corresponding multi-source resource data includes the coordinates of the testing organization, the scope of projects undertaken by the testing organization, and the publication time of the engineering machinery standards.
[0059] Specifically, the discrete decision threshold for adjusting the matching result is the minimum effective discrete change threshold used to determine whether the user demand characteristics at two adjacent time points trigger an adjustment of the matching result.
[0060] Specifically, the magnitude of change can be quantified by the distance between user demand feature vectors, including Euclidean distance, weighted distance, and distance based on cosine similarity, with the preferred embodiment being distance based on cosine similarity.
[0061] In implementation, the method of this invention determines the accuracy of the matching results based on the user acceptance rate of the matching results. Because user needs are characterized by sudden and discrete changes, and the matching model's response to real-time demand changes is lagging, the generated matching results deviate from the user's current actual needs. This reduces the user acceptance rate of the matching results, failing to accurately reflect the true effectiveness of the matching results, and consequently affecting the system's assessment and response to the degree of demand satisfaction. By determining the accuracy of the matching results, user satisfaction can be transformed into a quantifiable evaluation indicator, promptly identifying matching distortion problems caused by algorithm defects or data quality issues, and triggering targeted optimization of the model training strategy or algorithm. The method also determines the real-time nature of the matching result adjustment based on the push delay of multi-source resource data. Due to limitations in multi-source data synchronization efficiency and network transmission, the end-to-end delay from data update to matching result generation and push to the user is too high, affecting the system's rapid response to dynamic demand and resource changes. The system's response capability is insufficient to obtain effective matching results within the time window. By determining the real-time nature of matching result adjustments, the system's response performance can be transformed into a quantifiable timeliness indicator. This allows for timely identification of response delays caused by data synchronization lags or improper resource scheduling, triggering targeted optimizations to the data stream processing architecture or dynamic priority scheduling. The discrete judgment threshold for matching result adjustments can be adjusted based on the user demand matching deviation rate. Because model updates lag behind dynamically changing user demands and resource states, the model cannot capture effective feature changes in the real-time data stream in a timely manner, leading to a systematic deviation between the model output and the user's true intent, causing significant matching result failure. By reducing the discrete judgment threshold for matching result adjustments, the magnitude of demand changes required to trigger matching result adjustments can be reduced, allowing the model to trigger matching result updates when detecting minor changes in demand or resource states. This improves the sensitivity of matching result adjustments to real-time demand changes and enhances the accuracy of matching result generation.
[0062] Please continue reading. Figure 2 The diagram shown is a logical flowchart illustrating the process by which the accuracy of the generated matching result of the information matching method based on multi-source heterogeneous data fusion in an embodiment of the present invention meets the requirements.
[0063] Specifically, the accuracy of the generated matching results is determined based on the user acceptance rate, including:
[0064] Compare the user acceptance rate of the matching results with the preset acceptance rate;
[0065] If the user acceptance rate of the matched result is greater than the preset acceptance rate, the accuracy of the generated matching result is determined to meet the requirements. Specifically, if the user acceptance rate of the matched result is less than or equal to the preset acceptance rate, the accuracy of the generated matching result is determined to fail to meet the requirements.
[0066] Understandably, in information matching methods based on multi-source heterogeneous data fusion, using a preset adoption rate to characterize the accuracy of the generated matching results involves transforming the qualitative assessment of the accuracy of the generated matching results into a quantifiable user behavior feedback indicator, and using the preset adoption rate as a threshold boundary to determine whether the matching results meet the accuracy requirements. The preset adoption rate can be set according to actual working conditions. The setting of the preset adoption rate aims to ensure the accuracy and usability of the generated matching results. Optionally, the preset adoption rate is determined through a limited number of experiments by evaluating the effect of different adoption rates on the generation of matching results. The determined preset adoption rate should be neither too small nor cause excessive interference to the generation process of the matching results. For example, the preset adoption rate is generally selected within the range of [58%, 62%].
[0067] Preferably, the preset adoption rate is 60% in this preferred embodiment.
[0068] Specifically, the user adoption rate of the matching results is the ratio of the number of matching results actually adopted by users to the total number of matching results generated.
[0069] In practice, the method described in this invention determines the accuracy of the generated matching results by setting a preset adoption rate. Because user needs are characterized by sudden and discrete changes, and the matching model's response to real-time changes in needs is lagging, the generated matching results deviate from the user's current actual needs. This reduces the user's adoption rate of the matching results, failing to accurately reflect the true effectiveness of the matching results, and consequently affecting the system's assessment and response to the degree of need fulfillment. By determining the accuracy of the matching results, user satisfaction can be transformed into a quantifiable evaluation indicator, promptly identifying matching distortion problems caused by algorithm defects or data quality issues, and triggering targeted optimization of the model training strategy or algorithm, further improving the accuracy of the generated matching results.
[0070] Please continue reading. Figure 3 The diagram shown is a logical flowchart illustrating the process of determining whether the real-time performance of the matching result adjustment meets the requirements of the information matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention.
[0071] Specifically, if the accuracy of the generated matching results does not meet the requirements, the real-time performance of the matching result adjustment is determined based on the push delay time of multi-source resource data.
[0072] Specifically, the determination of whether the real-time performance of matching result adjustments meets the requirements based on the push delay duration of multi-source resource data includes:
[0073] The push delay time of the multi-source resource data is compared with the preset delay time;
[0074] If the push delay of the multi-source resource data is less than or equal to the preset delay, the real-time performance of the matching result adjustment is determined to meet the requirements.
[0075] Specifically, if the push delay of the multi-source resource data is longer than the preset delay, the real-time performance of the matching result adjustment is deemed unsatisfactory.
[0076] Understandably, in information matching methods based on multi-source heterogeneous data fusion, the core logic of using a preset delay duration to characterize the real-time performance of matching result adjustments is to transform the qualitative assessment of the real-time performance of matching result adjustments into a quantifiable time indicator, and to use the preset delay duration as a threshold boundary for determining whether the matching result adjustments meet real-time requirements. The preset delay duration can be set according to actual working conditions. The setting of the preset delay duration aims to ensure the accuracy and usability of the generated matching results. Optionally, the preset delay duration is determined through a limited number of experiments by evaluating the effect of different push delay durations on the generation of matching results. The determined preset delay duration should be neither too small nor cause excessive interference to the matching result generation process. For example, the preset delay duration is generally selected in the range of [0.9s, 1.1s].
[0077] Preferably, the preset delay duration is 1 second.
[0078] Specifically, the push delay of multi-source resource data is the difference between the actual time and the target time for the optimized and adjusted multi-source resource data to be pushed to the user-side interface.
[0079] In practice, the method described in this invention determines the real-time performance of matching result adjustments by setting a preset delay duration. Due to limitations in multi-source data synchronization efficiency and network transmission, the end-to-end delay from data update to matching result generation and push to the user is too high, affecting the system's ability to respond quickly to dynamic demands and resource changes, and making it impossible to obtain effective matching results within the time window. By determining the real-time performance of matching result adjustments, the system response performance can be transformed into a quantifiable timeliness indicator, promptly identifying response delay issues caused by data synchronization lag or improper resource scheduling, triggering targeted optimizations to the data stream processing architecture or dynamic priority scheduling, and further improving the accuracy of matching result generation.
[0080] Please continue reading. Figure 4 As shown, it is a logical flowchart of the discrete judgment threshold process for determining the matching result adjustment in the information matching method based on multi-source heterogeneous data fusion according to an embodiment of the present invention.
[0081] Specifically, if the real-time adjustment of the matching result does not meet the requirements, the recognition sensitivity of the real-time demand is determined based on the user demand matching deviation rate.
[0082] Specifically, the discrete decision threshold for adjusting the matching result is determined based on the user demand matching deviation rate, including:
[0083] Compare the user demand matching deviation rate with the preset deviation rate;
[0084] If the user demand matching deviation rate is less than or equal to the preset deviation rate, it is determined that the real-time demand recognition sensitivity meets the requirements.
[0085] Specifically, if the user demand matching deviation rate is greater than the preset deviation rate, it is determined that the recognition sensitivity of the real-time demand does not meet the requirements, and the discrete judgment threshold for adjusting the matching result is reduced.
[0086] Understandably, in information matching methods based on multi-source heterogeneous data fusion, the preset deviation rate characterizes the processing stability of operational status data. The core logic is to transform identification sensitivity into a quantifiable demand matching deviation index. When the identification sensitivity fails to meet requirements, the judgment threshold is adjusted to improve the responsiveness of the matching results to real-time demand changes. The preset deviation rate can be set according to actual working conditions. The setting of the preset deviation rate aims to ensure the accuracy and practicality of the generated matching results. Optionally, the preset deviation rate is determined through a limited number of experiments by evaluating the effect of the signal-to-noise ratio of different vibration signals on the generation of matching results. The determined preset deviation rate should be neither too small nor cause excessive interference to the generation process of matching results. For example, the preset deviation rate is generally selected within the range of [9%, 11%].
[0087] Preferably, the preset deviation rate is 10% in the preferred embodiment.
[0088] Specifically, the user demand matching deviation rate is the ratio of the number of matches in which the matching results failed to meet the user's actual needs to the total number of matches.
[0089] Specifically, the reduction in the discrete judgment threshold of the matching result adjustment is determined by the difference between the user demand matching deviation rate and the preset deviation rate.
[0090] Specifically, when the difference between the user demand matching deviation rate and the preset deviation rate is within 2%, the discrete judgment threshold for adjusting the matching result is reduced to 0.9 times the original value. When the difference between the user demand matching deviation rate and the preset deviation rate exceeds 2%, the discrete judgment threshold for adjusting the matching result is reduced by 0.01 for every 1% increase beyond the original value, in addition to the reduction to 0.9 times the original value. For example, when the difference between the user demand matching deviation rate and the preset deviation rate is 3%, the current discrete judgment threshold for adjusting the matching result is 0.15, and the reduced discrete judgment threshold for adjusting the matching result is 0.15×0.9-0.01×1=0.125.
[0091] In practice, the method described in this invention adjusts the minimum number of complete cycles in the processing process by setting a preset delay duration. Because the model update lags behind the dynamically changing user needs and resource status, the model cannot capture effective feature changes in the real-time data stream in a timely manner, thereby causing a systematic deviation between the model output and the user's true intention, resulting in significant failure of the matching results. By reducing the discrete judgment threshold for adjusting the matching results, the magnitude of demand change required to trigger the adjustment of the matching results can be reduced, so that the model can trigger the update of the matching results when it detects a small change in demand or resource status, thereby improving the sensitivity of the matching results adjustment to real-time demand changes and further improving the accuracy of the generated matching results.
[0092] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. An information matching method based on multi-source heterogeneous data fusion, characterized in that, include: Collect multi-source resource data from several data sources, and sequentially clean, denoise, extract information, match entities, and fuse the multi-source resource data to construct a resource knowledge graph. Then, extract features from the resource knowledge graph to output resource attribute features. The initial model is trained based on the resource attribute features and user needs to obtain a demand matching model, and a matching result is generated based on the demand matching model and historical user needs. The matching results are optimized and adjusted based on real-time user needs to obtain optimized matching results, and the optimized matching results and their corresponding multi-source resource data are pushed to the user terminal. Obtain the user acceptance rate of the matching results, and determine whether the accuracy of the generated matching results meets the requirements based on the user acceptance rate of the matching results; If the accuracy of the generated matching result does not meet the requirements, the real-time performance of the matching result adjustment is determined by the push delay time of the multi-source resource data. If the real-time performance of the matching result adjustment does not meet the requirements, then the discrete judgment threshold for the matching result adjustment is determined based on the user demand matching deviation rate.
2. The information matching method based on multi-source heterogeneous data fusion according to claim 1, characterized in that, The accuracy of the generated matching results is determined based on the user acceptance rate, including: Compare the user acceptance rate of the matching results with the preset acceptance rate; If the user acceptance rate of the matching result is greater than the preset acceptance rate, it is determined that the accuracy of the generated matching result meets the requirements.
3. The information matching method based on multi-source heterogeneous data fusion according to claim 2, characterized in that, If the user acceptance rate of the matching result is less than or equal to the preset acceptance rate, it is determined that the accuracy of the generated matching result does not meet the requirements.
4. The information matching method based on multi-source heterogeneous data fusion according to claim 3, characterized in that, If the accuracy of the generated matching results does not meet the requirements, the real-time performance of the matching result adjustment is determined based on the push delay time of multi-source resource data.
5. The information matching method based on multi-source heterogeneous data fusion according to claim 4, characterized in that, The real-time performance of matching result adjustments based on the push delay duration of multi-source resource data is determined to meet the requirements, including: The push delay time of the multi-source resource data is compared with the preset delay time; If the push delay of the multi-source resource data is less than or equal to the preset delay, the real-time performance of the matching result adjustment is determined to meet the requirements.
6. The information matching method based on multi-source heterogeneous data fusion according to claim 5, characterized in that, If the push delay of the multi-source resource data is longer than the preset delay, it is determined that the real-time performance of the matching result adjustment does not meet the requirements.
7. The information matching method based on multi-source heterogeneous data fusion according to claim 6, characterized in that, If the real-time performance of the matching result adjustment does not meet the requirements, the sensitivity of real-time demand recognition is determined based on the user demand matching deviation rate.
8. The information matching method based on multi-source heterogeneous data fusion according to claim 7, characterized in that, The discrete decision threshold for adjusting the matching result is determined based on the user demand matching deviation rate, including: Compare the user demand matching deviation rate with the preset deviation rate; If the user demand matching deviation rate is less than or equal to the preset deviation rate, it is determined that the real-time demand recognition sensitivity meets the requirements.
9. The information matching method based on multi-source heterogeneous data fusion according to claim 8, characterized in that, If the user demand matching deviation rate is greater than the preset deviation rate, it is determined that the recognition sensitivity of the real-time demand does not meet the requirements, and the discrete judgment threshold for adjusting the matching result is reduced.
10. The information matching method based on multi-source heterogeneous data fusion according to claim 9, characterized in that, The reduction in the discrete judgment threshold of the matching result adjustment is determined by the difference between the user demand matching deviation rate and the preset deviation rate.