Scientific and technological resource data extraction system based on multi-source data fusion
By designing a scientific and technological resource data extraction system based on multi-source data fusion, the problems of information asymmetry and underutilization of resources in the existing system are solved, efficient resource matching and personalized management are achieved, and the accuracy of resource reliability evaluation is improved.
Patent Information
- Application Number
- CN202510305227.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The existing scientific and technological resource data extraction systems lack efficient information release and demand matching mechanisms, resulting in information asymmetry, scientific and technological resources are not fully utilized, and the generated resource analysis reports cannot deeply explore the correlation and usage trends between resources.
A scientific and technological resource data extraction system based on multi-source data fusion is designed, including information release module, intelligent matching module, data tracking module, quality evaluation module and report generation module. Quality evaluation is carried out through neural networks, combined with improved BP neural network and PSO optimization algorithms, to achieve improved accuracy in resource reliability classification.
It effectively improves the efficiency of data extraction of scientific and technological resources, reduces labor costs, realizes personalized data management and data extraction of scientific and technological resources, and improves the accuracy of resource reliability assessment.
Smart Images

Figure CN120234757A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data extraction, and particularly to a scientific and technological resource data extraction system based on multi-source data fusion. Background Art
[0002] In the prior art, scientific and technological resource data extraction systems face various technical challenges: existing systems often lack an efficient information release and demand matching mechanism. There is often information asymmetry between the resource information released by scientific and technological institutions and the demands of scientific and technological resource requesters, resulting in scientific and technological resource requesters having difficulty quickly finding the required scientific and technological resources, and scientific and technological resources may also be wasted due to underutilization. Moreover, when generating resource analysis reports, existing systems can often only provide basic statistical information and are unable to deeply explore valuable information such as the associations between resources and usage trends. At the same time, quality assessment reports also often lack targeted optimization suggestions and cannot provide substantial help to scientific and technological resource requesters. These problems limit the efficiency, accuracy of the system, and the experience of scientific and technological resource requesters. For example, in CN118228162A, multi-source data is not collected, and neural network technology is not used for target data extraction, which urgently requires those skilled in the art to solve the corresponding technical problems. Summary of the Invention
[0003] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a scientific and technological resource data extraction system based on multi-source data fusion.
[0004] To achieve the above object of the present invention, the present invention provides a scientific and technological resource data extraction system based on multi-source data fusion, including:
[0005] An information release module: used for scientific and technological institutions to release scientific and technological resource information, scientific and technological resource requesters to submit scientific and technological data demands, and display the usage data of scientific and technological resources, where the usage data includes one or any combination of download volume, citation frequency, and update status;
[0006] An intelligent matching module: used for matching the demands of scientific and technological resource requesters with scientific and technological resources;
[0007] A data tracking module: used for real-time tracking of the usage data of scientific and technological resources, including download volume, citation frequency, and update status;
[0008] A quality assessment module: evaluating the reliability of scientific and technological resources through a neural network to obtain a quality assessment value; a report generation module: generating a resource analysis report and / or a quality assessment report; the resource analysis report includes one or any combination of resource type, usage frequency, and associated research fields, and the quality assessment report includes a quality assessment value, resource level, and optimization suggestions;
[0009] Among them, the information release module and the data tracking module are both connected to the intelligent matching module, the intelligent matching module is connected to the quality evaluation module, and the report generation module is connected to the quality evaluation module and the data tracking module.
[0010] Preferably, the information release module includes:
[0011] A science and technology institution information release unit configured to allow science and technology institutions to release project information, technical documents, and dataset metadata;
[0012] A science and technology resource demander demand release unit configured to allow science and technology resource demanders to submit science and technology data demands, where the demands include data retrieval requests and technical support consultations;
[0013] An information display unit configured to classify and display the resources released by the science and technology institutions and the science and technology demands submitted by the science and technology resource demanders, and support science and technology resource demanders to retrieve based on keywords or tags.
[0014] Preferably, the intelligent matching module includes a condition matching unit and / or a personalized matching unit:
[0015] Condition matching unit: Match the science and technology demands of science and technology resource demanders with the conditions of science and technology resources through a condition matching formula. The condition matching formula is as follows:
[0016]
[0017] Among them, M pd represents the matching degree between science and technology resource demander p and science and technology resource d in terms of conditions;
[0018] w k is the weight coefficient of the kth condition, representing the relative importance of each condition when a science and technology resource demander matches with a resource in science and technology resource matching;
[0019] n represents the number of conditions;
[0020] C pdk represents the matching score between science and technology resource demander p and science and technology resource d on the kth condition;
[0021] m represents the number of resource quality indicators, and the resource quality indicators include one or any combination of data accuracy, citation frequency, and update timeliness;
[0022] N pl represents the demand value of science and technology resource demander p on the lth resource quality indicator;
[0023] R dl represents the resource value provided by science and technology resource d on the lth resource quality indicator;
[0024] N max,l and N min,l respectively represent the maximum and minimum demands of the l-th resource quality index among all technology resource demanders and technology resources;
[0025] ∈ is a positive number used to ensure that the denominator will not be zero;
[0026] λ is the weight coefficient of the geographical location distance penalty term;
[0027] GeoDist is a function representing the geographical distance;
[0028] L p and L d respectively represent the geographical locations of technology resource demanders and technology resources;
[0029] Personalized matching unit: Match the historical retrieval records of technology resource demanders with technology resource characteristics through a personalized matching formula. The personalized matching formula is as follows:
[0030]
[0031] where P ph represents the matching degree between the historical retrieval record of technology resource demander p and the characteristics h of technology resources;
[0032] v l represents the personalized feature weight coefficient of the l-th resource quality index, and the weight coefficient represents the relative importance of different resource quality indexes;
[0033] H phl represents the characteristic value of the l-th resource quality index;
[0034] γ·Sim(PH p ,HH h ) represents the similarity reward term between the personalized technology resource historical data PH p of technology resource demander p and the technology resource historical data HH h ;
[0035] γ is the weight coefficient of the similarity reward term;
[0036] Sim is a similarity calculation function;
[0037] PH p represents the personalized technology resource historical data of technology resource demander p, such as information on the medical history, examination results, medication records, etc. of technology resource demanders.
[0038] HH h represents the set technology resource data;
[0039] η·TimeElapsed(T last_p ,T update_h ) represents the time penalty term elapsed since the last update of the scientific and technological resource data by the scientific and technological resource demander p;
[0040] T last_p is the time when the scientific and technological resource demander last updated the data;
[0041] T update_h is the latest update time of the historical data of scientific and technological resources;
[0042] TimeElapsed represents calculating the time difference;
[0043] η is the weight coefficient of the time penalty term.
[0044] Personalized matching refers to the process of tailoring matching strategies and resource recommendations for the scientific and technological resource demander according to information such as the personalized needs, preferences, and historical behaviors of the scientific and technological resource demander. The core of personalized matching lies in understanding the real needs of the demander and providing resource recommendations that conform to their personalized characteristics.
[0045] Preferably, it further includes;
[0046] If the intelligent matching module only has a conditional matching unit, C pdh = M pd ;
[0047] If the intelligent matching module only has a personalized matching unit, C pdh = P ph ;
[0048] If the intelligent matching module includes a conditional matching unit and a personalized matching unit, C pdh is calculated through the following formula:
[0049] C pdh = α·M pd + β·P ph
[0050] where C pdh represents the comprehensive matching degree between the scientific and technological resource demander p and the scientific and technological resource d (and its associated personalized data h);
[0051] α represents the weight coefficient used to balance the importance of conditional matching;
[0052] β represents the weight coefficient used to balance the importance of personalized matching.
[0053] Preferably, the data tracking module includes:
[0054] A resource usage tracking unit, configured to collect resource usage data in real time through log analysis, including download volume, access duration, and citation records;
[0055] An update status tracking unit, configured to record the version update information of resources and compare it with the subscription requirements of technology resource requesters to generate a resource dynamic reminder report.
[0056] Preferably, the neural network is an improved BP neural network. The improved BP neural network performs backpropagation of errors and updates the errors by referring back to the previous layer based on the errors of different layers. The expression of the improved BP neural network is as follows:
[0057]
[0058] Where, represents the error correction vector of the l-th layer regarding the accuracy of data extraction;
[0059] η(t) represents the adaptive learning rate function;
[0060] Y risk represents the expected output vector of data extraction accuracy;
[0061] represents the actual output vector of data extraction accuracy of the l-th layer;
[0062] f′() represents the derivative vector of the activation function;
[0063] represents the output vector of the l-th layer regarding the accuracy of data extraction;
[0064] λ″ represents the regularization coefficient;
[0065] represents the weight vector of data extraction accuracy of the l-th layer;
[0066] || ||2 represents the L2 norm;
[0067] sign() represents the sign function, used to indicate the positive or negative of the weight.
[0068] Preferably, it further includes: optimizing and improving the weights and biases of the BP neural network using the PSO optimization algorithm; in order to seek the optimal solution of the risk assessment model, thereby improving its performance. Specifically, the PSO algorithm regards the weights and biases of the BP neural network as the particle positions in the search space, and finds the optimal or approximately optimal particle positions (i.e., the optimal weights and biases) through iterative search. In this process, the PSO algorithm updates the particle velocity and position according to the fitness value of the particle (usually the error function of the BP neural network). Through continuous iterative search, the PSO algorithm can gradually converge to the optimal solution or approximately optimal solution, thereby obtaining a BP neural network model with better performance.
[0069] The PSO optimization algorithm is expressed as follows:
[0070]
[0071] Where, represents the particle velocity vector of the science and technology resource demander p at time t (representing the search direction of data extraction accuracy);
[0072] represents the particle velocity vector of the science and technology resource demander p at time t-1;
[0073] W(t) represents the inertia weight vector that changes with time;
[0074] ⊙ represents the element-wise multiplication operator;
[0075] C ind 、C glb represent the individual and global acceleration coefficient vectors respectively;
[0076] σ() represents a non-linear activation function, such as sigmoid;
[0077] R ind 、R glb represent the individual and global random number vectors respectively;
[0078] H p represents the current data extraction accuracy data vector of the science and technology resource demander p;
[0079] represents the individual optimal data extraction accuracy data vector;
[0080] represents the global optimal data extraction accuracy data vector.
[0081] In summary, due to the adoption of the above technical solutions, the system of the present invention integrates multi-source data of scientific and technological resource information and the scientific and technological data requirements submitted by scientific and technological resource demanders, combines artificial intelligence and big data analysis technologies, provides intelligent scientific and technological resource management services, can effectively improve the efficiency of scientific and technological resource data extraction, reduce labor costs, and realize personalized scientific and technological resource data management and data extraction. In addition, by combining the improved BP neural network and PSO optimization algorithm, the accuracy of resource reliability classification is improved.
[0082] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0084] Figure 1 is a schematic block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0086] The present invention provides a scientific and technological resource data extraction system based on multi-source data fusion. The multi-source data is data from multiple different channels, different types, and different formats. These data come from databases, social media, website logs, transaction records, etc. Specifically, in the scientific and technological resource data extraction system, it includes scientific and technological resource information released by scientific and technological institutions, demand information submitted by scientific and technological resource demanders, and usage data of scientific and technological resources (such as download volume, citation frequency, update status). These data exist in different forms, including structured data (such as tables in a database), semi-structured data (such as data in XML or JSON format), and unstructured data (such as text, images, videos, etc.). In the scientific and technological resource data extraction system based on multi-source data fusion, the fusion of multi-source data is to integrate data from different sources, different types, and different formats into a unified and consistent data set.
[0087] The system of the present invention is as Figure 1 shown and includes:
[0088] Information Release Module: Used for scientific and technological institutions to release scientific and technological resource information, scientific and technological resource demanders to submit scientific and technological data requirements, and to display the usage data of scientific and technological resources, where the usage data includes one or any combination of download volume, citation frequency, and update status;
[0089] Intelligent Matching Module: Used to match the needs of scientific and technological resource demanders with scientific and technological resources; before matching, it is necessary to process the scientific and technological resource information of the information release module and the scientific and technological data requirements submitted by scientific and technological resource demanders. According to the characteristics of the data sources, they are divided into structured data (such as relational databases), semi-structured data (such as CSV files), and unstructured data (such as text, images). Then, missing value and outlier processing are carried out in sequence, and data in different formats are converted into a unified format, such as converting text data into structured data. Finally, the data from different data sources are merged to form a unified data set.
[0090] Data Tracking Module: Used to track the usage data of scientific and technological resources in real time, including download volume, citation frequency, and update status;
[0091] Quality Assessment Module: Assess the reliability of scientific and technological resources through a neural network to obtain a quality assessment value; Report Generation Module: Generate a resource analysis report and / or a quality assessment report; the resource analysis report includes one or any combination of resource type, usage frequency, and associated research fields, and the quality assessment report includes the quality assessment value, resource level, and optimization suggestions;
[0092] Among them, the information release module and the data tracking module are both connected to the intelligent matching module, the intelligent matching module is connected to the quality assessment module, and the report generation module is connected to the quality assessment module and the data tracking module.
[0093] Preferably, the information release module includes:
[0094] Scientific and Technological Institution Information Release Unit, configured to allow scientific and technological institutions to release project information, technical documents, and dataset metadata;
[0095] Scientific and Technological Resource Demander Requirement Release Unit, configured to allow scientific and technological resource demanders to submit scientific and technological data requirements, where the requirements include data retrieval requests and technical support consultations;
[0096] Information Display Unit, configured to classify and display the resources released by the scientific and technological institutions and the scientific and technological requirements submitted by scientific and technological resource demanders, and support scientific and technological resource demanders to retrieve based on keywords or tags.
[0097] Preferably, the intelligent matching module includes a condition matching unit and / or a personalized matching unit:
[0098] Condition Matching Unit: Initially match the technology requirements of technology resource demanders with the conditions of technology resources through a condition matching formula, and the condition matching formula is as follows:
[0099]
[0100] Among them, M pd represents the matching degree between technology resource demander p and technology resource d in terms of conditions;
[0101] w k is the weight coefficient of the k-th condition, representing the relative importance of each condition when matching technology resource demanders with resources in technology resource matching;
[0102] n represents the number of conditions; each matching condition represents a specific screening factor, and these factors involve relevant attributes of technology resources, such as the professional direction of experts in the research field, the facility conditions of scientific research institutions or laboratories, the amount of project funding or cooperation funds, etc.
[0103] C pdk represents the matching score between technology resource demander p and technology resource d on the k-th condition;
[0104] m represents the number of resource quality indicators, and each resource quality indicator represents a specific requirement or resource, such as the accuracy, integrity, timeliness, authority of data, as well as the availability, accessibility of resources and the satisfaction of technology resource demanders, etc.
[0105] N pl represents the demand value of technology resource demander p on the l-th resource quality indicator;
[0106] R dl represents the resource value provided by technology resource d on the l-th resource quality indicator;
[0107] N max,l 、N min,l respectively represent the maximum and minimum demand values of the l-th resource quality indicator among all technology resource demanders and technology resources; |N pl -R dl | This range is used to normalize the absolute difference to make it a relative value between 0 and 1.
[0108] ∈ is a positive number, used to ensure that the denominator will not be zero;
[0109] λ is the weight coefficient of the geographical location distance penalty term;
[0110] GeoDist is a function representing the geographical distance;
[0111] L p ,Ld respectively represent the geographical locations of technology resource demanders and technology resources;
[0112] w k is the weight coefficient;
[0113] S ijk is the matching score between technology resource demander i and technology resource j on the k-th condition;
[0114] where represents the total weighted matching score of technology resource demander p and technology resource d on n conditions. represents the product of the relative matching degrees of technology resource demander p and technology resource d on m key demand / resource quality indicators; by calculating the differences in the values of the demander and the resource on each indicator and normalizing considering the maximum and minimum values of the indicators, the matching degree between them is measured. λ·GeoDist(L p ,L d ) represents the distance penalty term of technology resource demander p and technology resource d in terms of geographical location. Condition matching refers to the process of matching technology resource demanders and technology resource providers according to a preset number of conditions or criteria. These conditions are usually based on the clear demands of the demanders, the attributes of the resource providers, and specific matching rules. The core of condition matching lies in ensuring that the matching results meet the preset set of conditions, thus meeting the basic requirements of the demanders.
[0115] The condition matching formula comprehensively considers the matching degrees of technology resource demanders and technology resources on multiple conditions and resource quality indicators, as well as the influence of data location, through weighted summation, product operation, and geographical location distance penalty term. And by adjusting the weight coefficient and penalty term, it flexibly reflects the importance of different conditions and resource quality indicators in the matching process, as well as the influence of geographical location distance on the matching results. Therefore, this formula can reasonably calculate the matching degree between technology resource demanders and technology resources.
[0116] Personalized matching unit: Match the historical retrieval records of technology resource demanders with technology resource characteristics (technology resource characteristics: resource form, quality indicators, application fields, etc.) through the personalized matching formula. The personalized matching formula is as follows:
[0117]
[0118] where, P ph represents the matching degree between the historical retrieval record of technology resource demander p and technology resource characteristic h;
[0119] v lDenote the personalized feature weight coefficient of the l-th resource quality indicator; in the context of personalized matching degree, the weight coefficient represents the relative importance or influence of this resource quality indicator on meeting the needs or preferences of specific technology resource demanders. These weight coefficients are usually dynamically adjusted according to factors such as the personal characteristics, historical behaviors, and preference settings of technology resource demanders to ensure that the resource matching process can more accurately reflect the personalized needs of technology resource demanders.
[0120] H phl Denote the eigenvalue of the l-th resource quality indicator;
[0121] γ is the weight coefficient of the similarity reward term;
[0122] Sim is a similarity calculation function; used to measure the similarity degree between two data sets (here it is the personalized technology resource historical data of technology resource demanders and the set technology resource data).
[0123] PH p Denote the personalized technology resource historical data of technology resource demander p, including personalized information related to technology resources such as usage records or preference data. These data reflect the interests, needs, and behavior patterns of the demander. For example, information such as cloud service and data storage history, hardware and device usage history, etc. By analyzing the historical data of the demander, the system can understand the preferences and needs of the demander, so as to recommend more technology resources that meet their expectations.
[0124] HH h Denote the historical data of the characteristic data of each resource in the personalized technology resources, and these data include various attributes, quality indicators, application fields, etc. of the resources, which are used to comprehensively describe the characteristics and values of the resources; the historical data of technology resource characteristics HH h Can be understood as a dynamic snapshot of the technology resource library, which records the characteristic information of the resources at different time points. These information are of great significance for understanding the evolution trend of the resources, evaluating the quality and value of the resources. At the same time, by comparing the data at different time points, potential problems and improvement directions of the resources can also be found.
[0125] T last_p Is the time when the technology resource demander last updated the data;
[0126] T update_h Is the latest update time of the technology resource historical data;
[0127] TimeElapsed denotes the calculated time difference;
[0128] η is the weight coefficient of the time penalty term.
[0129] Among them, denotes the total weighted matching score of the historical retrieval records of the science and technology resource requester p and the science and technology resource characteristics h on m personalized characteristics. By weighted summation, all resource quality indicators are comprehensively considered to obtain a comprehensive personalized matching score. γ·Sim(PH p ,HH h ) represents the personalized science and technology resource historical data PH of the science and technology resource requester p p and the set science and technology resource data HH h similarity reward term; the purpose of setting the similarity reward term is to improve the matching degree of science and technology resources similar to the historical data of the science and technology resource requester. η·TimeElapsed(T last_p ,T update_h ) represents the time penalty term elapsed since the science and technology resource requester p last updated the science and technology resource data; subtracting the time penalty term aims to reduce the matching degree of science and technology resources that have not been updated for a long time. This formula comprehensively considers the matching degree between the historical retrieval records of the science and technology resource requester and the science and technology resource characteristics, as well as the influence of time factors through weighted summation, similarity reward term and time penalty term. The personalized matching formula can flexibly reflect the importance of different resource quality indicators in the personalized matching process and the influence of time difference on the matching result by adjusting the weight coefficient and penalty term. Therefore, this formula can reasonably calculate the matching degree between the historical retrieval records of the science and technology resource requester and the science and technology resource characteristics.
[0130] Preferably, it further includes;
[0131] If the intelligent matching module only has a conditional matching unit, C pdh =M pd ;
[0132] If the intelligent matching module only has a personalized matching unit, C pdh =P ph ;
[0133] If the intelligent matching module includes a conditional matching unit and a personalized matching unit, C pdh is calculated by the following formula:
[0134] C pdh =α·M pd +β·P ph
[0135] where C pdh represents the comprehensive matching degree of the science and technology resource requester p and the science and technology resource d (and its associated personalized data h);
[0136] α represents the weight coefficient used to balance the importance of conditional matching;
[0137] β represents the weight coefficient used to balance the importance of personalized matching.
[0138] Before matching through the above steps, it is necessary to perform data extraction and processing on the scientific and technological resource information of the information release module and the scientific and technological data requirements submitted by scientific and technological resource demanders. According to the characteristics of the data extraction sources, they are divided into structured data (such as relational databases), semi-structured data (such as CSV files), and unstructured data (such as texts, images). Missing value and outlier processing are carried out in sequence, and then data in different formats are converted into a unified format, such as converting text data into structured data. Finally, the data in different data sources are merged to form a unified data set. Relevant scientific and technological data are extracted from these data sources, and subsequent analysis and processing are carried out through the extracted scientific and technological data throughout the process.
[0139] Preferably, the data tracking module includes:
[0140] A resource usage tracking unit configured to collect resource usage data in real time through log analysis, including download volume, access duration, and citation records;
[0141] An update status tracking unit configured to record the version update information of the resource and compare it with the subscription requirements of scientific and technological resource demanders to generate a resource dynamic reminder report.
[0142] Preferably, the neural network is an improved BP neural network. The improved BP neural network performs backpropagation of errors and updates the errors by referring to the errors of different layers and returning to the previous layer; the expression of the improved BP neural network is as follows:
[0143]
[0144] Where, represents the error correction vector of the l-th layer regarding the accuracy of data extraction;
[0145] η(t) represents the adaptive learning rate function;
[0146] Y risk represents the expected output vector of data extraction accuracy;
[0147] represents the actual output vector of data extraction accuracy of the l-th layer;
[0148] f′() represents the derivative vector of the activation function;
[0149] represents the output vector of the l-th layer regarding the accuracy of data extraction;
[0150] λ″ represents the regularization coefficient;
[0151] represents the weight vector for the data extraction accuracy of the l-th layer;
[0152] || ||2 represents the L2 norm;
[0153] sign() represents the sign function, which is used to indicate the positive or negative of the weight.
[0154] Although the traditional BP neural network can effectively handle non-linear mapping problems, when dealing with large-scale and high-dimensional health risk data, it often faces problems such as slow convergence speed, easy to fall into local minima, and overfitting. The present invention introduces an adaptive learning rate function η(t) to enable the BP neural network to more effectively process large-scale health risk data, improve the accuracy and efficiency of data extraction. A regularization and sparsification strategy is also adopted to enhance the generalization ability of the model, making the BP neural network more robust when dealing with new data sources and improving the robustness of the data extraction system. Moreover, by optimizing the error correction term, the ability of the BP neural network to capture data features is enhanced, making the extracted scientific and technological resource data more accurate and comprehensive.
[0155] Preferably, it further includes: using the PSO optimization algorithm to optimize the weights and biases of the improved BP neural network; to seek the optimal solution of the quality assessment module, thereby improving its performance. Specifically, the PSO algorithm regards the weights and biases of the BP neural network as the particle positions in the search space, and finds the optimal or approximate optimal particle positions (i.e., the optimal weights and biases) through iterative search. In this process, the PSO algorithm updates the particle velocity and position according to the fitness value of the particle (usually the error function of the BP neural network). Through continuous iterative search, the PSO algorithm can gradually converge to the optimal solution or approximate optimal solution, thereby obtaining a BP neural network model with better performance.
[0156] The PSO optimization algorithm is expressed as follows:
[0157]
[0158] where, represents the particle velocity vector of the scientific and technological resource demander p at time t (representing the search direction of data extraction accuracy);
[0159] represents the particle velocity vector of the scientific and technological resource demander p at time t-1;
[0160] W(t) represents the inertia weight vector that changes with time;
[0161] ⊙ represents the element-wise multiplication operator;
[0162] C ind 、C glbrepresent the individual and global acceleration coefficient vectors respectively;
[0163] σ() represents a non - linear activation function, such as sigmoid;
[0164] R ind 、R glb represent the individual and global random number vectors respectively;
[0165] H p represents the current data extraction accuracy data vector of the technology resource demander p;
[0166] represents the individual optimal data extraction accuracy data vector;
[0167] represents the global optimal data extraction accuracy data vector.
[0168] The particle swarm optimization algorithm (PSO) has advantages such as fast convergence speed and easy implementation when solving optimization problems. However, when dealing with complex and multi - modal optimization problems, it also faces problems such as premature convergence and low search accuracy. And this application solves the above problems. An inertial weight vector W(t) that changes with time is designed and dynamically adjusted according to the number of iterations or convergence situation to balance the global search and local search capabilities. During the parameter optimization process of the technology resource data extraction system, dynamic inertial weight adjustment can enable the PSO algorithm to maintain a high global search ability in the initial stage and focus on local fine - search in the later stage, thereby improving the accuracy and efficiency of parameter optimization. A non - linear activation function and a random number vector are also introduced to make the PSO algorithm avoid falling into local optimal solutions during the search process, improving the globality and accuracy of the search. And the update strategies for the individual optimal solution and the global optimal solution are optimized to ensure that the changing trend of the optimal solution can be accurately and efficiently captured during the iteration process. The optimized update strategy can enable the PSO algorithm to converge to the global optimal solution more quickly, improving the overall performance of the data extraction system.
[0169] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and purposes of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. A scientific and technological resource data extraction system based on multi-source data fusion, characterized in that: include: Information release module: used for scientific and technological institutions to release scientific and technological resource information, scientific and technological resource demanders to submit scientific and technological data needs, and to display the usage data of scientific and technological resources, which includes one or any combination of download volume, citation frequency, and update status; Intelligent matching module: used to match the needs of technology resource demanders with technology resources; Data tracking module: used to track the usage data of scientific and technological resources in real time, including download volume, citation frequency and update status; Quality assessment module: evaluates the reliability of scientific and technological resources through neural networks to obtain quality assessment values; Report generation module: generates resource analysis report and / or quality assessment report; the resource analysis report includes one or any combination of resource type, usage frequency, and related research field; the quality assessment report includes quality assessment value, resource level, and optimization suggestions; Among them, the information release module and the data tracking module are connected to the intelligent matching module, the intelligent matching module is connected to the quality assessment module, and the report generation module is connected to the quality assessment module and the data tracking module.
2. A scientific and technological resource data extraction system based on multi-source data fusion according to claim 1, characterized in that: The information publishing module includes: A science and technology institution information publishing unit, configured for science and technology institutions to publish project information, technical documents and dataset metadata; A demand publishing unit for technology resource demanders, configured to allow technology resource demanders to submit technology data demands, including data retrieval requests and technical support consultations; The information display unit is configured to display the resources released by the scientific and technological institutions and the scientific and technological needs submitted by scientific and technological resource demanders in a classified manner, and to support scientific and technological resource demanders to search based on keywords or tags.
3. A scientific and technological resource data extraction system based on multi-source data fusion according to claim 1, characterized in that: The intelligent matching module includes a condition matching unit and / or a personalized matching unit: Condition matching unit: The technology needs of technology resource demanders are preliminarily matched with the conditions of technology resources through the condition matching formula. The condition matching formula is as follows: Among them, M pd It represents the matching degree between the technological resource demander p and the technological resource d in terms of conditions; w k is the weight coefficient of the kth condition; n represents the number of conditions; C pdk It means that this is the matching score between the technology resource demander p and the technology resource d on the kth condition; m represents the number of resource quality indicators, wherein the resource quality indicators include one or any combination of data accuracy, number of citations, and update timeliness; N pl represents the demand value of scientific and technological resource demander p on the lth resource quality indicator; R dl represents the resource value provided by scientific and technological resource d on the lth resource quality indicator; N max,l 、N min,l They represent the maximum and minimum demand values of the lth resource quality indicator among all technology resource demanders and technology resources; ∈ is a positive number to ensure that the denominator will not be zero; λ is the weight coefficient of the geographical location distance penalty term; GeoDist is a function representing geographic distance; L p ,L d They represent the geographical locations of technology resource demanders and technology resources respectively; Personalized matching unit: matches the historical search records of technology resource demanders with the characteristics of technology resources through a personalized matching formula. The personalized matching formula is as follows: Among them, P ph It represents the matching degree between the historical search records of the technology resource demander p and the technology resource feature h; v l represents the personalized feature weight coefficient of the lth resource quality indicator; H phl represents the characteristic value of the lth resource quality indicator; γ is the weight coefficient of the similarity reward item; Sim is a similarity calculation function; PH p Represents the personalized scientific and technological resource historical data of the scientific and technological resource demander p; HH h Indicates the set scientific and technological resource data; T last_p It is the time when the technology resource demander last updated the data; T update_h It is the latest update time of the historical data of scientific and technological resources; TimeElapsed means calculating the time difference; η is the weight coefficient of the time penalty term.
4. A scientific and technological resource data extraction system based on multi-source data fusion according to claim 3, characterized in that: Also includes; If the intelligent matching module only has conditional matching units, C pdh =M pd ; If the intelligent matching module only has a personalized matching unit, C pdh =P ph ; If the intelligent matching module includes a condition matching unit and a personalized matching unit, C pdh Calculated by the following formula: C pdh =α·M pd +β·P ph Among them, C pdh It represents the comprehensive matching degree between the technology resource demander p and the technology resource d; α represents the weight coefficient used to balance the importance of condition matching; β represents the weight coefficient used to balance the importance of personalized matching.
5. The scientific and technological resource data extraction system based on multi-source data fusion according to claim 1 is characterized in that: The data tracking module includes: A resource usage tracking unit, configured to collect resource usage data in real time through log analysis, including download volume, access duration, and citation records; The update status tracking unit is configured to record version update information of resources and compare it with the subscription requirements of technology resource demanders to generate a resource dynamic reminder report.
6. A scientific and technological resource data extraction system based on multi-source data fusion according to claim 1, characterized in that: The neural network is an improved BP neural network, which uses the improved BP neural network to back propagate the error and return to the previous layer to update the error by referring to the errors of different layers; the expression of the improved BP neural network is as follows: in, represents the error correction vector of the lth layer regarding the accuracy of data extraction; η(t) represents the adaptive learning rate function; Y risk represents the desired data extraction accuracy output vector; Represents the actual data extraction accuracy output vector of layer l; f′() represents the derivative vector of the activation function; Represents the output vector of the lth layer regarding the accuracy of data extraction; λ″ represents the regularization coefficient; Represents the weight vector of the accuracy of data extraction at the lth layer; ||||2 represents L2 norm; sign() represents the sign function, which is used to indicate the positive or negative value of the weight.
7. The scientific and technological resource data extraction system based on multi-source data fusion according to claim 1 is characterized in that: It also includes: using PSO optimization algorithm to optimize the weights and biases of the improved BP neural network; The PSO optimization algorithm is expressed as follows: in, represents the particle velocity vector of the technology resource demander p at time t; represents the particle velocity vector of the technology resource demander p at time t-1; W(t) represents the inertia weight vector that changes with time; ⊙ represents the element-wise multiplication operator; C ind , C glb denote the individual and global acceleration coefficient vectors, respectively; σ() represents a nonlinear activation function; R ind , R glb denote individual and global random number vectors respectively; H p The data vector representing the current data extraction accuracy of the technology resource demander p; represents the individual optimal data extraction accuracy data vector; Represents the global optimal data extraction accuracy data vector.
Citation Information
Patent Citations
Big data analysis method and system based on science and technology enterprise innovation resource investment
CN118228162A
Cited By
Scientific and technological resource management and service system based on big data
CN120782395A