A big data driving-based information matching and pushing method and system

Through multi-channel collection and data preprocessing, combined with entity recognition and machine learning algorithms, the integration and matching problems in enterprise information processing are solved, accurate information push and decision support are achieved, and the efficiency and accuracy of enterprise information acquisition are improved.

CN120277271BActive Publication Date: 2025-10-17ZHEJIANG POST & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510741258.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-17
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In information processing, enterprises face the problem of wide-ranging information sources and diverse formats, which makes comprehensive collection and integration difficult. The existing matching and push systems lack personalization, resulting in the failure to fully tap the value of information. In addition, technologies such as big data and artificial intelligence have problems with data quality and accuracy.

Method used

Through web crawlers and API interfaces, we collect multi-channel information, combine data preprocessing, entity recognition and natural language processing, build enterprise portraits and information impact models, use machine learning algorithms to calculate the matching degree, and optimize push strategies through feedback.

Benefits of technology

It achieves accurate matching and efficient push of information, improves enterprise decision-making efficiency, reduces information screening costs, enhances the pertinence and accuracy of information, and supports enterprises to quickly obtain valuable information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277271B_ABST
    Figure CN120277271B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of information screening, and particularly relates to a kind of information matching push method and system based on big data driving, comprising: obtaining relevant data, data preprocessing;Utilize automatic update tool and NLP to build entity recognition model, filter low-quality data, detect and correct contradictory data;Information key information is extracted by natural language processing technology, and information is automatically classified;Based on enterprise content data, build multidimensional enterprise portrait, build information influence model;According to the matching degree of single information and enterprise, information complementarity matrix and the matching of information and enterprise planning, information matching model is built, and the final information matching degree is calculated;Information and information influence model analysis result of the information matching degree higher than matching degree threshold value are pushed to enterprise, collect user feedback, and information matching model and push strategy are optimized.The present application considers the matching degree of information and enterprise from multiple angles, improves the accuracy of push.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information screening, in particular to an information matching and pushing method and system based on big data driving. BACKGROUND

[0002] Under the wave of digitization, the dependence of enterprise development on information is increasing day by day, and its acquisition and processing capacity has become an important part of enterprise competitiveness. However, enterprises face serious challenges in information processing. On the one hand, information sources are extensive, covering government, industry associations, enterprise internal systems and other channels, with various formats and scattered, structured, semi-structured and unstructured data coexisting, which makes it difficult for enterprises to fully collect and effectively integrate, increasing the processing difficulty.

[0003] Traditional information processing methods rely on manual screening and simple tools, which are low in efficiency and prone to errors, especially when dealing with a large amount of unstructured text data, it is difficult to accurately extract key information and understand semantics, resulting in the value of information being unable to be fully tapped. Existing matching and pushing systems lack understanding of enterprise individual needs, and the pushed information lacks pertinence. Although emerging technologies such as big data, artificial intelligence and blockchain have brought opportunities, there are also problems, such as data quality and processing efficiency of big data, accuracy and interpretability of artificial intelligence models, performance bottlenecks and compatibility of blockchain, etc. Therefore, it is important to develop an innovative information processing solution to integrate multi-source data, ensure data security, and achieve accurate analysis and matching, so that enterprises can make quick and accurate decisions in complex market environments. SUMMARY

[0004] The present application considers multiple factors to calculate the matching degree, improves the matching degree between the screened information and the enterprise, and enhances the satisfaction of the enterprise.

[0005] The technical solution proposed by the present application is: an information matching and pushing method based on big data driving, the method comprising:

[0006] Using network crawler technology to capture information text data, using API interface to obtain enterprise internal data, integrating multi-channel information text data and enterprise operation data and performing data preprocessing;

[0007] Using automatic updating tools and NLP to build an entity recognition model, setting a threshold, filtering low-quality data in the preprocessed data, detecting and correcting contradictory data;

[0008] Extracting key information of information by natural language processing technology, and automatically classifying information using text classification algorithm;

[0009] Through data analysis and machine learning technology, a multi-dimensional enterprise portrait is constructed based on enterprise content data, quantitative indicators are determined for information type, an information influence model is constructed, the influence of information on enterprises is analyzed, and a complementary matrix of information is constructed between different information.

[0010] According to the matching algorithm, a single information matching model is constructed, the matching degree of single information and enterprises is calculated, and the information matching model is constructed according to the matching degree of single information and enterprises, the information complementary matrix and the matching of information and enterprise planning, and the final information matching degree is calculated.

[0011] The information and information influence model analysis results with a final information matching degree higher than the matching degree threshold are pushed to the enterprise, user feedback is collected, and the information matching model and the pushing strategy are optimized according to the user feedback.

[0012] Preferably, the data preprocessing comprises the following steps:

[0013] For frequently updated websites, crawling is performed once a day or once an hour, and for slowly updated industry association websites, crawling is performed once a week or once a month, and a random delay mechanism is used to crawl information text data from government official websites, government platforms and industry association websites; related data in the enterprise internal system is obtained through API interface, and the basic information, operation data and industry data of the enterprise are integrated; the collected data is cleaned to remove duplicate and invalid data; the missing values are filled by mean filling and regression prediction method; the data is standardized; the data structure and format are unified, and the association relationship between the data is established.

[0014] Preferably, the specific process of low-quality data filtering is as follows:

[0015] The BERT model is used to process the information text data, the key entities in the information text data and the related entities in the enterprise internal data are determined by entity recognition, the relationship between the two kinds of entities is established by using relationship extraction technology, and a dynamic knowledge graph is constructed; a data quality threshold is set to automatically filter low-quality data; the reasoning engine of the knowledge graph is used to detect contradictory data, and when contradictory data is detected, manual review or automatic correction is triggered to correct the contradictory data.

[0016] Preferably, the specific process of information automatic classification is as follows:

[0017] The keyword extraction algorithm is used to extract keywords, titles and key sentences as text features, and the publishing time, publishing media and associated enterprise attributes are extracted as structured features; the text features and structured features are converted into numerical vectors through a word vector model; the training data are collected and labeled through a supervised learning model, and the model is trained to learn the features and category classification; the model performance is evaluated through a test set, and the parameters are adjusted, the features are optimized and the model is optimized according to the results; the numerical vectors are input into the model for classification, and the results are stored according to the categories.

[0018] Preferably, the information impact model comprises the following formula:

[0019] Total impact of technical innovation achievements The sum of direct impact and indirect impact, and the specific formula is as follows:

[0020] ;

[0021] Total impact of enterprise revenue The sum of direct impact and regulatory effect, and the specific formula is as follows:

[0022] ;

[0023] Wherein: is the intercept, , , are the impact coefficients of the corresponding variables respectively, is the error term, is the original revenue base value of the enterprise, is the information intensity, is the enterprise size, is the intercept term, is the information intensity Impact coefficient on technical innovation achievements, is the error term, is the random error term, is the intercept, is the impact coefficient, is the error term, is the error term.

[0024] Preferably, the final information matching degree calculation formula is as follows:

[0025] ;

[0026] Wherein: is the final information matching degree; is the collaborative filtering adjustment coefficient; is the matching degree of single information and enterprise, a variable for traversing the information information, a total number of information information, information information complementarity with information information complementarity with information information short-term and long-term planning weight coefficients, a short-term planning matching coefficient of the enterprise, a long-term planning matching coefficient of the enterprise.

[0027] Preferably, the information pushing method has the following specific process:

[0028] According to the final matching degree, the information information matching results are arranged in descending order; a matching degree threshold is set , information information below the threshold is filtered out; the enterprise is provided with sorting options in different dimensions such as information information complementarity strength, comprehensive benefit, and enterprise planning matching degree; the matched information information is pushed to the enterprise through email, short message, and system internal message; the enterprise can set the preferred pushing mode and time in the system; according to the browsing history, interest selection, and real-time demand changes of the enterprise, the information information combination pushing strategy is dynamically adjusted according to the current development stage and strategic goal of the enterprise.

[0029] Preferably, the optimization process of the information matching model and the pushing strategy is as follows:

[0030] A feedback entrance is set to collect the feedback of the enterprise on single information information and information information combination, including whether the complementarity of the information information combination is reasonable, whether the synergistic benefit meets the expectation, and whether the matching degree with the enterprise planning meets the expectation; user research is regularly carried out to understand the problems encountered by the enterprise in the process of applying the information information combination and improvement suggestions; if the enterprise feedback that the complementarity of a certain information information combination does not meet the expectation, the information information complementarity matrix is reexamined, and the complementarity coefficient of the related information information is adjusted; if the user is not satisfied with the matching degree of the information information combination and the enterprise planning, the calculation model of the enterprise planning and the information information matching is optimized; if the user proposes new demands for the pushing time and mode of the information information combination, the pushing strategy is optimized in time.

[0031] The application also provides an information information matching pushing system based on big data driving.

[0032] The application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the information information matching pushing method based on big data driving.

[0033] The application has the following beneficial effects:

[0034] 1. Through network crawlers, API interfaces and other multi-channel collection of information text data and enterprise internal operating data, using data cleaning, standardization and other preprocessing techniques to eliminate data quality differences and format barriers. At the same time, a unified data storage and management architecture is established to break down data silos and enable efficient collaboration of data from different sources and types. This provides a rich and complete data foundation for subsequent analysis, enabling the model to learn more comprehensive information features, significantly improving the prediction accuracy and generalization ability of the model, and thus enhancing the accuracy of information matching and pushing, helping enterprises quickly obtain valuable information, reducing information screening costs, and improving decision-making efficiency.

[0035] 2. By building complex causal relationship models, considering the relationship between multiple variables, and using machine learning algorithms for data-driven modeling and parameter optimization. This innovative analysis method can deeply explore the internal relationship between information and enterprise development, and accurately quantify the impact of information on various aspects of the enterprise. Based on this, information effect prediction, combination evaluation and counterfactual reasoning are carried out to provide comprehensive decision support for enterprises. Enterprises can predict the effect of information implementation in advance, evaluate the synergistic benefits of different information combinations, and understand the development situation if a certain information is not applied, so as to more scientifically develop development strategies, rationally allocate resources, avoid risks caused by blind decision-making, and improve the market competitiveness of enterprises.

[0036] 3. A perfect quantitative index system is built to accurately measure the matching degree of information and enterprises. In the matching algorithm design, the similarity of enterprise portrait and information application conditions, the complementarity between information, and the matching with enterprise planning are considered to ensure that the matching results are comprehensive and accurate. The matching result sorting and screening mechanism can display the most valuable information according to the enterprise's needs. Multi-channel pushing and personalized pushing functions can achieve precise touch of information according to the enterprise's browsing history, interest preferences, etc. Visual and interactive design allows enterprises to intuitively understand information content and matching basis, and user feedback and model optimization mechanism continuously improves the accuracy of matching and the effectiveness of pushing. This series of innovative measures greatly improves the efficiency of enterprises in obtaining and utilizing information, reduces the cost of information screening for enterprises, and enables enterprises to quickly obtain information that is highly consistent with themselves, promoting enterprise development. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 The flowchart of the information matching and pushing method based on big data driving of the present application;

[0038] Figure 2 The information pushing process flowchart of the information matching and pushing method based on big data driving of the present application. DETAILED DESCRIPTION

[0039] The following description is presented to enable any person skilled in the art to practice the application as claimed. Preferred embodiments are presented in the following description only as examples and modifications thereto can be made by those skilled in the art upon reading the present disclosure without departing from the spirit of the application. The present application defined in the claims is not intended to be limited by the preferred embodiments set forth in the following description. Rather, embodiments disclosed in the following description are intended to be illustrative only of the many embodiments that can be implemented according to the technology included with the present disclosure.

[0040] It can be understood that the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of one element can be one, and in another embodiment, the number of the element can be multiple, and the term "one" cannot be understood as a limitation on the number.

[0041] As shown in Figure 1 and Figure 2 First, data collection and preprocessing are performed, and multi-source data fusion is performed. Specific rules are set for web crawler technology to capture information website content, and related data in the enterprise internal system is obtained through the API interface. According to the update frequency and data importance of the target website, a reasonable crawling frequency is set. For frequently updated websites, such as government policy release platforms, crawling is performed once a day or once an hour; for relatively slow industry association websites, crawling is performed once a week or once a month. At the same time, a random delay mechanism is adopted to avoid making a large number of requests to the target website in a short period of time, so as to be banned IP by the website. In order to reduce the amount of data crawling and improve the efficiency of crawling, an incremental crawling strategy is adopted. Record the time point or data identifier of each crawling, and only get the updated or new data next time. For example, by comparing the publication time or version number of the information article, it is judged whether it is new data, and only the newly added and updated information is crawled, so as to reduce the load pressure on the target website and improve the timeliness of data update. The information text data collected from government official websites, government platforms, industry association websites and other channels is integrated, and the basic information, operating data and industry data of enterprises are also integrated. In order to ensure the accuracy and completeness of the data, the collected data is cleaned to remove duplicate and invalid data, and the missing values are filled by using mean filling, regression prediction and other methods to improve the data integrity. Multi-source data integration is the basis of the whole scheme, which provides original data support for subsequent steps such as information text analysis, information quantification and modeling. The quality of data directly affects the accuracy of subsequent analysis and matching.

[0042] In the data sharing link, the combination of federated learning and blockchain is adopted to protect the privacy of enterprise data. The multi-party secure calculation (MPC) is used to optimize the federated learning communication mechanism, and the enterprise data is calculated locally and only the model gradient is shared. Assuming that there are federated learning participating enterprises, the local data of enterprise ​ , the model parameters are , the gradient calculated based on local data is , where is the enterprise loss function based on local data. The federal learning server updates the global model parameters by aggregating the gradients of each enterprise . The data usage log is recorded through the blockchain, and the data update task is automatically triggered through the smart contract, ensuring the traceability and compliance of data sharing. This step ensures the privacy and security of enterprise data, promotes data sharing, and solves the data silo problem, enabling model training to utilize more data and improving the accuracy and generalization ability of the model. At the same time, it ensures the compliance and traceability of data usage, providing a safe data source for subsequent information quantification and modeling, enabling model training to be based on more abundant data, improving model quality, and thus affecting the accuracy of information matching and pushing. At the same time, the data usage log recorded by the blockchain can also provide reference for user feedback and system optimization.

[0043] Dynamic knowledge graph construction and data cleaning: Introduce knowledge graph automatic updating tools (such as RDFox or GraphDB), combine NLP entity recognition models (such as spaCy or BERT), and real-time capture of information and enterprise dynamic data. Determine the key entities in the information (such as information subjects, applicable scope, support measures, etc.) and related entities in the enterprise information through entity recognition, and establish the relationship between the two entities using relationship extraction technology, thereby constructing a dynamic knowledge graph. Set a data quality threshold (such as accuracy > 95%), automatically filter low-quality data; use the reasoning engine of the knowledge graph (such as RDFS or OWL) to detect contradictory data, and when contradictory data is detected, trigger manual review or automatic correction to correct contradictory data. In entity recognition, use the BERT model to process information text, output the feature vector , of each word as a single feature vector identified, and use a specific classifier to determine whether the word is an information entity based on these feature vectors.

[0044] This step ensures the real-time, accuracy, and consistency of data in the knowledge graph, providing rich semantic information for information text parsing, improving the usability and understandability of information, and facilitating subsequent information matching and analysis. This step provides structured knowledge support for information text parsing, helping to more accurately extract key information of information; at the same time, its data cleaning function provides high-quality data for information quantification and modeling, ensuring the reliability of model training.

[0045] Taking a certain company as an example, the system periodically crawls various types of information text from the local government science and technology department and industrial and information technology department website, covering fields such as technological innovation and industrial support. At the same time, through API docking with the enterprise internal management system, the company's R&D investment, patent quantity, revenue, profit and other operating data are obtained. During data cleaning, it is found that some patent data has duplicate records, which are removed after verification; for missing quarterly revenue data, a regression prediction method is used to fill in the data based on the company's past revenue trends and the revenue of similar companies in the same industry. This ensures the quality of the data and provides a reliable data foundation for subsequent information analysis and matching, avoiding analysis bias caused by data errors or missing.

[0046] In a regional technology enterprise alliance, the company and other technology companies jointly participate in federated learning training of the same information matching model. Local R&D data, market data, etc. are used to calculate the model gradient locally. For example, based on the company's R&D investment and the number of new product launches, the gradient information of the technology achievement transformation support information matching model is calculated and encrypted and uploaded. The federated learning server aggregates these gradients to update the global model, and the blockchain records the time and data volume of each enterprise's data contribution in detail. Through smart contracts, when new information or enterprise data is updated, the data update task is automatically triggered.

[0047] For an information about the artificial intelligence industry support, the knowledge graph automatic update tool automatically updates the information from the government website in real time, the BERT model identifies the key entities such as "artificial intelligence enterprise", "R&D subsidy" and "technology innovation index", and determines the "obtain subsidy" relationship between "artificial intelligence enterprise" and "R&D subsidy" through relationship extraction technology, thereby constructing the knowledge graph. If it is found that the accuracy of a certain company's technology innovation data is only 90% (lower than the 95% threshold), the data is automatically filtered; if the enterprise in the knowledge graph both meets the high-pollution industry restriction information and obtains environmental protection subsidies (assuming that the data entry is incorrect), the manual review process is triggered.

[0048] Next, the collected information text is analyzed and interpreted by artificial intelligence, and natural language processing (NLP) technology is used to deeply interpret the information text. Through morphological analysis, the information text is split into words and tagged with parts of speech; through syntactic analysis, the grammatical structure of the sentence is constructed; through word vector models (such as Word2Vec, GloVe) and deep learning models (such as Transformer, BERT), the semantic meaning of the text is understood. Taking the BERT model as an example, the information text is pre-trained to capture contextual semantic information. Assuming that the information text is , the semantic representation obtained after processing by the BERT model is The key clauses, restrictions, and preferential measures in the information information are identified through the semantic representation. The core information of the information information, such as the information information name, the issuing agency, the issuing time, the applicable object, and the information information content summary, is extracted through information extraction technology and stored in the information information intelligent database, providing a data basis for subsequent information information matching.

[0049] For example, for the information information text "Encourage artificial intelligence enterprises to increase R&D investment, and give a subsidy of 15% of R&D investment to enterprises with annual R&D investment exceeding 500 million yuan and more than 10 artificial intelligence related patents, and provide special R&D fund support", the morphological analysis splits it into words such as "encourage", "artificial intelligence enterprise", "increase", "R&D investment" and annotates the part of speech; the syntactic analysis determines the sentence structure; the BERT model understands the semantics and identifies "annual R&D investment exceeding 500 million yuan" and "more than 10 artificial intelligence related patents" as key conditions, and "give a subsidy of 15% of R&D investment" and "provide special R&D fund support" as preferential measures. The information extraction technology extracts information such as the information information name (such as "artificial intelligence enterprise R&D support information information") and the applicable object (artificial intelligence enterprises meeting certain R&D investment and patent quantity requirements) and stores it in the database. This step can extract valuable structured information from unstructured information information text, convert the information information content into a form that machines can understand and process, and provide key data support for information information quantification and modeling, information information matching. It is also a prerequisite for information information quantification and modeling, providing accurate information information key information; at the same time, the extracted information directly affects the accuracy of information information matching and is an important basis for information information matching.

[0050] Then, according to the parsed information, the collected data is classified and structured. A classification system is established according to the theme, industry, and applicable scope of the information information, such as manufacturing, service, and technology industries, and support information information, tax information information, and regulatory information information. Use text classification algorithms (such as support vector machine SVM, convolutional neural network CNN) to automatically classify information information. Taking SVM as an example, assume that the feature vector of the information information text is , and the SVM model predicts the category of the information information . The information information content is structured and converted into a format that machines can easily process, such as XML, JSON. For example, a technology innovation support information information is represented as a JSON object containing fields such as information information basic information, applicable conditions, support methods, and reporting procedures, making it easy for subsequent matching calculations and queries.

[0051] For example, for a newly collected series of information information text, first extract the key words, sentence vectors and other features in the text to construct a feature vector . Use the trained SVM model to classify "artificial intelligence enterprise R&D support information" and determine that it belongs to the "technology enterprise-support information" category. Then structure the information as JSON format: {"Information name": "Artificial intelligence enterprise R&D support information", "Applicable conditions": "Artificial intelligence enterprises with annual R&D investment exceeding 5 million yuan and more than 10 artificial intelligence related patents", "Support method": "Give a 15% subsidy on R&D investment and provide special R&D funding support", "Reporting process": "..."} This step facilitates the management and retrieval of a large amount of information, improves the efficiency of information matching, and makes the organization of information more orderly, providing clearer information classification information for information matching algorithms, and improving the accuracy and relevance of matching. Provide classification information for information matching to help narrow the matching range and improve matching efficiency; at the same time, the structured information data provides a standardized data format for information quantification and modeling, facilitating quantitative analysis.

[0052] Next, information quantification and modeling are performed, and an information quantification index system is constructed. Different quantification indicators are determined for different types of information. For financial subsidy information, quantification indicators can include subsidy amount A, subsidy ratio R, application threshold (such as enterprise revenue I, R&D investment RD, etc.); for tax preference information, tax relief amount T, preferential period P, etc. Through analysis and data mining of information text, these quantification indicators are extracted and standardized to make them comparable. Assuming that the subsidy amounts of multiple financial subsidy information are standardized, the z-score standardization method is adopted, and the standardized subsidy amount , where is the mean of the subsidy amount, is the standard deviation of the subsidy amount, the mean of the subsidy amount of the same type of information is calculated first, then the difference between the subsidy amount and the mean is calculated, and finally the difference is divided by the standard deviation of the subsidy amount to obtain the standardized subsidy amount parameter. For example, the subsidy amount of a certain information is 300 million yuan, the mean of the same type of information is 160 million yuan, and the standard deviation is 83.67 million yuan, then the standardized subsidy amount This result shows that the subsidy amount of this information is about 1.67 standard deviations higher than the average level of the same type, and it is a subsidy information with relatively large subsidy intensity, which may obtain a higher weight in the matching model. The standardization method of other quantification indicators is similar.

[0053] For example, when analyzing a series of information on support for the science and technology industry, the subsidy ratio (R) for "information on support for R&D of artificial intelligence enterprises" was determined to be 15%, and the application threshold required that the enterprise's R&D investment (RD) exceed 5 million yuan. After standardizing these quantitative indicators, they were compared and analyzed with the quantitative indicators of other information to facilitate more scientific information matching. This step converts the information content into quantifiable indicators, facilitating mathematical calculations and model building, making information matching more scientific and precise, and more accurately measuring the fit between enterprises and information. Providing a quantitative basis for information matching is a key step in achieving accurate matching; at the same time, the extraction of quantitative indicators relies on the results of information text parsing, which is a further quantitative processing of the information text parsing information.

[0054] Causal reasoning and information effect modeling:

[0055] In the scenario of information matching enterprises, a causal relationship network is constructed based on the structured causal model (SCM) and combined with a machine learning algorithm. The specific model structure is as follows:

[0056] Variable Definition

[0057] Information variables ( ): Set the information variable By information type , information intensity , information implementation time Composition, that is For example, for “information on R&D subsidies for artificial intelligence enterprises”, the information type is Coded as "financial subsidies" category (using one-hot coding, such as [1,0,0] for financial subsidies, [0,1,0] for tax incentives, [0,0,1] for talent support), information intensity Expressed as the proportion of subsidy amount to R&D investment, information implementation time The date the information was released (e.g., 2024-01-01).

[0058] Enterprise characteristic variables ( ): Enterprise characteristic variables Including enterprise size (For example, the number of employees, revenue scale, etc., the standardization formula for the number of employees is , revenue scale is normalized after logarithmic transformation), industry attributes (Using category coding, such as artificial intelligence technology companies are coded as 001), corporate strategy (Through the strategic goals reported by the enterprise, the text is vectorized and converted into a 768-dimensional vector using the BERT model), enterprise resources (Fund reserves, talent reserves, number of technical patents, etc., fund reserves and talent reserves are also standardized), that is, .

[0059] Enterprise outcome variables ( ): Enterprise outcome variables Including corporate revenue ,profit ,market share , technological innovation achievements (such as the number of patents, number of new product launches, etc.), that is, .

[0060] Causal structure

[0061] Direct causal relationship: Taking the impact of fiscal subsidy information on corporate R&D investment as an example, assuming a linear relationship exists, it can be expressed as ,in The technological innovation results (such as the number of patents) affected by financial subsidy information are: is the intercept term, Information intensity The impact coefficient on technological innovation results, is the error term. In business operations, information has a direct and critical impact on corporate revenue. ) and the influence coefficient ( ) multiplied by ( ), quantifying the direct effect of information on technological innovation results, through the intercept term ( ) takes into account those factors that are difficult to measure accurately but have a stable impact on technological innovation results. Random factors such as accidental inspiration in the scientific research process and external sudden technical exchanges are incorporated into it to make the model more in line with the actual situation. This formula realizes the quantitative modeling of the direct impact of fiscal subsidy information on technological innovation results, and provides key parameters for enterprises to evaluate the value of fiscal subsidy information to technological innovation. Referring to the actual operating data and industry research results of many enterprises, when enterprises obtain and effectively utilize fiscal subsidy information, it can often directly affect revenue in many aspects. Take a certain high-tech enterprise as an example. After accurately grasping the government’s fiscal subsidy information for scientific and technological innovation, it successfully applied for a large amount of subsidy funds. On the one hand, this fund can be directly used to expand production scale, purchase advanced production equipment, improve product output and quality, thereby increasing product sales and directly increasing revenue; on the other hand, subsidy funds can be invested in market promotion activities, expand sales channels, increase brand awareness, attract more customers to buy products, and drive revenue growth. From a financial analysis perspective, corporate revenue can be reflected through operating income in the income statement. Assuming the original revenue of the enterprise is the basic value Based on the revenue change data of a large number of similar companies after obtaining subsidy information and successfully applying for subsidies, a linear regression model was established, and the direct impact of financial subsidy information on corporate revenue was obtained as follows: ,in The coefficient representing the impact of financial subsidy information on revenue depends on factors such as industry characteristics, enterprise scale, and the strength of subsidy policies. is a random error term, reflecting the impact of other random factors not considered by the model on revenue. ) and the influence coefficient ( ) multiplied by ( ), accurately quantify the direct impact of information on revenue, the error term Factors such as sudden market competition and accidental changes in policy implementation are incorporated into the model, making it more in line with the complex situation where the actual revenue of enterprises is affected by information.

[0062] Indirect causal relationship: Talent support information ( ) by influencing the talent pool of enterprises ( talent reserve in the enterprise), which in turn affects the enterprise's technological innovation capabilities ( ). Assume that the relationship between talent reserves and technological innovation capabilities is ,in is the talent reserve variable, is the intercept, is the influence coefficient, is the error term. ) and the influence coefficient ( ) multiplied by ( ), quantifying the impact of talent reserves on technological innovation capabilities, the intercept ( ) comprehensively considers those factors that are difficult to accurately model but have a stable impact on technological innovation capabilities, while the error term ( ) incorporates random interference factors to make the model closer to the actual situation. The impact of talent support information on talent reserves is , The error term is the indirect causal relationship that can be reflected by combining these two equations. ) and the influence coefficient ( ) multiplied by ( ), quantify the impact of talent support information on talent reserves, intercept ( ) takes into account the basic conditions of the enterprise's talent reserves and other factors, and the error term ( ) incorporates random interference factors to make the model more consistent with the actual situation where talent reserves are affected by information.

[0063] Moderating effect: firm size ) on the impact of fiscal subsidy information ) on firm revenue , building a moderating effect model , where is the intercept, , , are the impact coefficients of the corresponding variables, is the error term, reflecting the moderating effect. quantifies the direct effect of fiscal subsidy information on firm revenue, reflects the independent contribution of firm size to revenue, is the key embodiment of the moderating effect, the error term includes factors not considered in the model that will have a random impact on firm revenue, making the model more realistic and complex market environment.

[0064] Total impact of technological innovation achievements is the sum of direct and indirect effects (assuming that fiscal subsidy information and talent support information independently affect technological innovation achievements), based on the principle that technological innovation achievements are affected by fiscal subsidy information and talent support information. Under the premise of assuming that the two kinds of information independently affect technological innovation achievements, direct and indirect effects are integrated, that is,

[0065] ;

[0066] Total impact of firm revenue is the sum of direct and moderating effects (assuming there are no other indirect factors not considered), around the impact of firm revenue on fiscal subsidy information and firm size, that is,

[0067] ;

[0068] Model construction and learning

[0069] Data-driven modeling: use PC algorithm to learn causal structure from enterprise historical data and information real-time data. Assume the data set , where is the number of data samples (such as =1000, covering 1000 enterprises' information participation and operation data in the past 5 years). PC algorithm builds a directed acyclic graph (DAG) step by step through conditional independence test, for example, in the test of information variable and enterprise result variable given enterprise characteristic variable When independence is determined under the conditions, the chi-square test statistic is used ,in and are the observed frequency and expected frequency, respectively. is the classification number, and the final causal structure is obtained by continuously adjusting the connection relationship between variables. ) and expected frequency ( ) and square it to reflect the degree of difference. ), normalization is performed to make classification differences of different magnitudes comparable, thereby reflecting the degree of deviation between observations and expectations under this classification.

[0070] Combined with machine learning: Taking Gradient Boosting Tree (GBT) as an example, the information variable and firm characteristic variables As input features , enterprise outcome variables As the output label. In the iterations, the regression tree is constructed , by minimizing the loss function (mean squared error loss) to update the model, where is the model prediction value, and the model prediction formula is , is the number of trees (e.g. =100). In order for the model to predict the value ( ) as close to the true value as possible ( ), we need to define an indicator to measure the difference between the two, that is, the loss function. The mean square error loss function is built based on this purpose. It calculates the true value of each sample ( ) and the model predicted value ( ) and sum the results of all samples to measure the overall prediction error of the model. Calculation is the The purpose of squaring the difference between the true value and the predicted value of a sample is to highlight the impact of larger errors, avoid the mutual offset of positive and negative errors, and more accurately reflect the degree of deviation from the prediction of a single sample.

[0071] Parameter estimation and optimization: When estimating the parameters of the impact of fiscal subsidy information on corporate revenue, the maximum likelihood estimation is used. Assuming that corporate revenue Normal distribution , the likelihood function is ,in , by solving the log-likelihood function The maximum value of the likelihood function is used to estimate the parameters . . The probability density function of the th sample of enterprise revenue under the normal distribution, where is the normalization constant of the normal distribution, ensuring that the integral of the probability density function over the entire domain is equal to 1, reflects the influence of the degree of deviation of the sample value from the mean value on the probability, the smaller the deviation, the greater the probability density, represents the multiplication of the probability density functions of the samples, resulting in the likelihood function of the entire data set, which comprehensively reflects the possibility of observing the revenue samples of the enterprises under the parameters . At the same time, cross-validation (such as 5-fold cross-validation) is used to adjust the model hyperparameters, such as the number of trees in gradient boosting trees , learning rate (such as =0.1), etc., to minimize the mean square error (MSE) as the objective function, i.e. . The calculation is the square of the difference between the true value and the predicted value of the th sample, aiming to highlight the impact of larger errors and avoid the cancellation of positive and negative errors, accurately reflecting the degree of prediction deviation of individual samples, the error values of the samples are first accumulated, and then divided by the number of samples to obtain the mean square error , reflecting the average prediction error of the model on all samples.

[0072] The specific application process of the model is as follows:

[0073] Information effect prediction: For candidate information and enterprises , input their features into the trained model to obtain the predicted enterprise outcome variable .

[0074] For example, a company plans to apply for "Artificial Intelligence Industry Innovation Development Support Information" ( ). The enterprise has 300 employees, and the standardized enterprise size (the minimum value of the number of employees in the sample is 100, and the maximum value is 500); the industry attribute is coded as 001; the enterprise strategy Transformed into a vector [0.7, 0.2, 0.1] by the BERT model; the capital reserve is 5 million yuan, standardized to 0.6 (assuming the minimum value of capital reserve in the sample is 1 million yuan and the maximum value is 9 million yuan), and the number of technology patents is 20, standardized to 0.4 (assuming the minimum value of patent quantity in the sample is 5 and the maximum value is 45). The information type is encoded as [1, 0, 0], and the information intensity is 15% of the R&D investment subsidy amount. The information implementation time is January 1, 2024. These variables are input into the trained gradient boosting tree model (number of trees = 100, learning rate = 0.1) to predict that after applying the information, the enterprise's revenue is expected to grow to 38 million yuan (original revenue of 30 million yuan), and the number of patents is expected to increase to 30.

[0075] This step helps enterprises evaluate the value of information in advance by accurately predicting the results of the enterprise after the implementation of the information, avoiding blind application of information. In actual application, after using the model to predict the effect of the information, the success rate of the enterprise applying for the information increases from 40% to 54%, an increase of 35%; the average resource waste cost caused by the mismatch of the information reduces by 40%, and the time cost reduces by 30%. The prediction of the effect of the information relies on the quantitative indicators of the information extracted in the information analysis stage (such as the type and intensity of the information) and the enterprise feature information generated in the enterprise portrait construction stage (such as enterprise size and industry attribute), providing an important reference for information matching. At the same time, the prediction results will be fed back to the information matching algorithm to optimize the calculation of the information matching degree, so that the accuracy of the information matching degree calculation is improved by 25%, further improving the accuracy of the information matching.

[0076] Information combination evaluation: for information combination , analyze the causal relationship and synergistic effect between them, the formula is = . By linear combination of the technology innovation achievement prediction values , corresponding to different information according to certain weights , , the weights are determined by the relative importance of the information, so as to get the comprehensive influence of the information combination, and analyze the causal relationship and synergistic effect between the information.

[0077] For example, a company is considering applying for "Artificial Intelligence Enterprise R&D Subsidy Information" ( ) and "Information on the introduction of high-end talents" ( ) combination. R&D subsidy information Information type Encoded as [1,0,0], information strength The subsidy amount accounts for 12% of R&D investment; high-end talent introduction information Information type Encoded as [0,0,1], information strength To provide high-end talents with a settling-in allowance of RMB 500,000. The company currently has 300 employees and a scale of =0.5; Industry attributes Coded as 001; Corporate Strategy The vector is [0.7, 0.2, 0.1]; the standardized value of capital reserve is 0.6, and the standardized value of the number of technical patents is 0.4. Through model calculation, R&D subsidy information ( ) is expected to increase the number of enterprise patents Add 8 items, high-end talent introduction information ( ) is expected to increase the company's market share Increase by 5%. According to the determined weight coefficient =0.6, =0.4, calculate the comprehensive impact of information combination on the enterprise =0.6×8+0.4×5=6.8, quantitatively assessing the high value of this information combination for corporate development. The weighting coefficients are dynamically adjusted based on differences in different industries, policy cycles, or corporate development stages. For example, traditional manufacturing focuses more on equipment upgrade subsidies, giving a higher weight to a company's market share, while technology companies focus more on R&D subsidies, giving a higher weight to the number of patents a company has.

[0078] This step can effectively evaluate the synergistic effect of information combination, help enterprises choose the optimal information combination, and improve the utilization efficiency of information resources. Actual data shows that after using this evaluation method, the probability of enterprises selecting effective information combination has increased from 30% to 42%, an increase of 40%; the average benefit of enterprises due to the synergistic effect of information combination has increased from 1.5 million yuan to 1.95 million yuan, an increase of 30%. Information combination evaluation is based on the candidate information selected in the information matching process, combined with the information causal relationship (direct causal, indirect causal, regulatory effect formula, etc.) obtained by information analysis and enterprise portrait information. The evaluation result provides more accurate information combination recommendation for information push, and the pertinence of information push is improved by 30%; at the same time, it can also be fed back to the information matching algorithm to optimize the matching strategy of information combination, and the accuracy of information combination matching is improved by 20%.

[0079] Counterfactual reasoning: assuming that the enterprise has not applied for certain information , under the current enterprise characteristics , the development of the enterprise is simulated by using the model .

[0080] For example, a company has applied for "Artificial Intelligence Industry Innovation and Development Support Information". The enterprise currently has 300 employees =0.5), the industry attribute is coded as 001, the enterprise strategy vector is [0.7, 0.2, 0.1], the standardized value of capital reserve is 0.6, and the standardized value of technology patent quantity is 0.4. Through counterfactual reasoning, the model prediction value is recalculated after removing the information variable related items. The model simulates that without applying for this information, the enterprise's revenue in the next year is only 33 million yuan, while the predicted revenue after applying for the information is 38 million yuan, and the value of the information =3800-3300=500 million yuan, which intuitively shows the contribution of the information to the growth of enterprise revenue.

[0081] This step helps enterprises clearly understand the actual value of information to their own development, providing strong support for subsequent information application decision-making. When sending information to enterprises, these relationship analyses are also sent to enterprises, thereby improving credibility and making enterprises more confident in the results. Practice shows that the rationality of enterprises making information decisions based on counterfactual reasoning has increased from 60% to 78%, an increase of 30%; the number of decision-making errors caused by misjudging the value of information has decreased by 35%. Counterfactual reasoning further analyzes the value of information using the results of information matching and information effect prediction. The results can be fed back to the information matching and pushing links, helping enterprises better understand the relationship between information and their own development, improving the fit of information matching by 25%, and improving the satisfaction of information pushing by 20%; at the same time, it also provides actual case basis for feedback optimization, promotes continuous improvement of the system, and optimizes the overall performance of the system by 15%.

[0082] To push information to enterprises, it is necessary to first build an enterprise portrait. Enterprise information is collected and sorted to pre-enter enterprise information items, including enterprise basic information (such as enterprise name , industry , size S, establishment time T, etc.), operating data (such as revenue R, profit P, market share M, etc.), R&D data (such as R&D input RD, number of patents PN, number of innovation achievements IN, etc.), talent data (such as employee education structure , number of professional skill talents SP, etc.). Collect and sort the entered enterprise information of various types to build an enterprise information database.

[0083] For example, a company's basic information is the company name, the industry is artificial intelligence, the size is medium (300 employees), and the establishment time is 2018; in terms of operating data, the annual revenue is 30 million yuan, the profit is 5 million yuan, and the market share in the relevant field reaches 8%; R&D data shows that the annual R&D input is 6 million yuan, the number of patents is 12, and the number of innovation achievements (such as new artificial intelligence algorithms) is 5; in terms of talent data, the proportion of employees with a master's degree or above is 30%, and the number of professional skill talents (such as artificial intelligence algorithm engineers and data scientists) is 80. After sorting these information, it is stored in the enterprise information database.

[0084] This step comprehensively and systematically collects enterprise information, providing a data basis for building an accurate enterprise portrait, accurately reflecting the characteristics and strength of the enterprise, and providing detailed enterprise information basis for subsequent information matching. It is the basis for enterprise portrait generation, and the completeness and accuracy of enterprise information directly affect the quality of enterprise portrait, and thus the accuracy of information matching and pushing. At the same time, the enterprise information database also provides enterprise-side data support for data comparison and analysis in information analysis.

[0085] Enterprise portrait generation is based on enterprise information database, using data analysis and machine learning technology to generate enterprise portrait. Through cluster analysis, enterprises are grouped according to similar characteristics, and key features are extracted using dimension reduction techniques such as principal component analysis (PCA) to construct a multi-dimensional portrait of the enterprise. Assuming that the feature matrix of enterprise information is X, the principal component matrix Y is obtained after PCA transformation , and the main principal components are selected as the key features of the enterprise portrait.

[0086] For example, when analyzing a large number of artificial intelligence companies, first construct a feature matrix X containing information about a certain company. Use PCA technology to process the matrix X to obtain the principal component matrix Y. It is found that the two principal components mainly reflect the enterprise's research and development innovation ability (highly related to R&D investment, patent quantity, etc.) and market competitiveness (related to revenue, market share, etc.). Based on these two principal components, the portrait of a certain company is constructed, highlighting its research and development strength in the field of artificial intelligence (high R&D investment, large number of patents), market competitiveness (certain revenue and market share), and talent advantage (high proportion of highly educated and skilled personnel), etc. to accurately match information.

[0087] This step simplifies complex enterprise information into representative multi-dimensional portraits, facilitating quick and accurate matching with information, improving the efficiency and accuracy of information matching, and providing more information recommendations that meet the actual needs of enterprises. The feature representation of the enterprise side is provided for information matching, and the matching calculation is performed with the quantitative indicators of information. At the same time, the generation of enterprise portraits depends on the results of enterprise information collection and sorting, which is a further processing and refinement of enterprise information.

[0088] Information matching:

[0089] Matching algorithm design establishes an information matching algorithm, using a content-based matching method to calculate the similarity between the enterprise portrait and the applicable conditions of the information. Considering that information focuses on different indicators, a weight vector is introduced, where is the number of indicators participating in the matching calculation, is the weight of the th indicator, and , .

[0090] Assuming that the th quantitative indicator of information is , and the actual value of the th indicator of the enterprise is . The similarity calculation formula is:

[0091] ;

[0092] The formula measures the similarity of the two by calculating the square root of the weighted sum of the differences between the enterprise indicators and the information indicators, and the value is closer to 1, the higher the similarity.

[0093] Constructing the information complementarity matrix , the matrix dimension is ( total number of information), where element represents the complementarity of information and information (currently selected information , information is the information that the intended recommended enterprise is still using and will be pushed to the enterprise), the value range is [0, 1], and the value is closer to 1, indicating stronger complementarity. The specific calculation formula is as follows:

[0094] ;

[0095] Among them: , , are weight coefficients, and + + =1, respectively used to adjust the importance of information target synergy ( ), resource complementarity ( ), and implementation effect mutual promotion ( ) in complementarity evaluation. According to industry characteristics, information types and other factors, it can be determined by expert scoring combined with historical data analysis, for example, in the technology innovation industry, the weight of information target synergy can be set to 0.4.

[0096] : The target synergy score of information and information , which is calculated by analyzing the semantic similarity of target description in information clauses, target field overlap, etc., with a full score , respectively as the theoretical maximum value of the target score when evaluating information , separately.

[0097] : The resource complementarity score of information and information Resource Complementarity Score, which evaluates the complementarity of resources (such as funds, talents, technology, etc.) required for information, with a full score of , respectively. , The theoretical maximum value of the resource score when evaluated separately.

[0098] : Information and the implementation effect of information mutual promotion score, based on historical implementation cases, analyzes the mutual promotion effect of information implementation, with a full score of , respectively. , The theoretical maximum value of the effect score when evaluated separately.

[0099] In the target synergy dimension, is the normalized value of the target synergy index, which is mapped to the [0, 1] interval, multiplied by the weight coefficient , to obtain the contribution value of target synergy in the complementarity of information and information , which is calculated in the same way in the resource complementarity dimension , and in the implementation effect complementarity dimension , and the sum of the three is the complementarity of the two information.

[0100] Enterprise planning matching factor:

[0101] Introduce enterprise planning matching coefficients and , respectively representing the matching degree of information and enterprise short-term planning and long-term planning, with a value range of [0, 1]. By analyzing the short-term (1-3 years) and long-term (more than 3 years) development goals, strategic direction filled in by enterprises in the system, and performing semantic matching and target consistency evaluation on information goals, the matching degree is calculated. For example, if the goal of information contains support for the promotion of a certain product market, the score will be higher.

[0102] When calculating the matching degree of information, the matching degree of a single information with the enterprise, the complementarity of information with other information, and the matching with the enterprise planning are considered comprehensively. The final matching degree calculation formula is adjusted as follows:

[0103] ;

[0104] wherein, is the collaborative filtering adjustment coefficient (value range [0.8, 1.2]), according to the feedback of enterprises, new data and information changes, using reinforcement learning and adaptive algorithm for dynamic adjustment, to ensure the accuracy and effectiveness of information matching and pushing, is the short-term and long-term planning weight coefficient (value range [0, 1]), enterprises can set it according to their own development stage, such as start-up enterprises can set it to 0.7, focusing on short-term planning matching, is the variable of traversing information, is the total number of information. This formula ensures that the system prioritizes recommending information combinations with high matching degree to the enterprise, strong complementarity with other information, and fit for the enterprise's planning. is the similarity of information and the enterprise in the index, multiplied by the weight coefficient, reflecting the contribution of the matching degree of a single information to the final matching degree, reflects the complementarity sum of information and all other information, reflecting the influence of the complementary characteristics of information in the overall information combination on the matching degree, reflects the influence of the matching degree of information and enterprise planning on the final matching degree, and the three are multiplied to comprehensively reflect the influence of multiple factors on the matching degree of information and enterprise. For example, a company (enterprise

[0105] ) participates in the "artificial intelligence industry innovation and development support information" (information ) matching, which involves R&D investment ratio, patent quantity, and high-tech product income ratio (three key indicators =3), with corresponding quantitative index requirements =20%, =15, =60%, and the weight vector is set to (0.4, 0.3, 0.3). The actual R&D investment ratio of the enterprise is 25%, the number of patents is 12, and the high-tech product income ratio is 70%. First, calculate the similarity :

[0106]

[0107] ;

[0108] ​​​​​The system determines the adjustment coefficient by finding the majority of successful applications of the information in similar enterprises to the company through a collaborative filtering algorithm = 1.2, then the final matching degree = 1.2 0.61 = 0.732.

[0109] This step is based on weighted calculation and collaborative optimization, comprehensive and accurate evaluation of the matching degree of enterprises and information, reducing the cost of enterprise screening information, improving the efficiency of information utilization. Relying on the quantitative indicators extracted by information analysis and enterprise portrait characteristics, the matching result basis for information push is provided.

[0110] Matching result sorting and screening

[0111] According to the final matching degree The information matching results are sorted in descending order, and the top-ranked information is displayed. Set the matching degree threshold (generally 0.5-0.7), filter out information below the threshold. At the same time, provide sorting options for enterprises according to information complementarity strength, comprehensive benefit, and enterprise planning matching degree, etc. in different dimensions, to facilitate enterprises to intuitively obtain the most valuable information combination. For example, enterprises in strategic transformation period can choose to sort by long-term planning matching degree, and prefer to view information combinations that help achieve long-term transformation goals.

[0112] For example, after a company completes information matching, the system defaults to display information by matching degree from high to low, and sets the threshold = 0.6, filter out information with a matching degree below 0.6. In addition, enterprises can also choose to re-sort and view the remaining information by publication time from recent to distant, or by the size of the preferential intensity.

[0113] This step reduces the interference of invalid information, meets the diversified query needs of enterprises, and improves the efficiency and pertinence of enterprise information acquisition. Process the information matching results to provide an orderly and accurate list of information for the information push link.

[0114] Information push:

[0115] Multi-channel push pushes the matched information to the enterprise through multiple channels such as email, SMS, and system in-site messages. Enterprises can set preferred push methods and times in the system, and for urgent information, the system will prioritize sending SMS to inform the enterprise in time.

[0116] For example, a company sets up in the system to take e-mail as the main receiving channel, and the receiving time is every Friday at 10 am. Every Friday, the system will send high-matching information such as "Artificial Intelligence Technology Research and Development Special Subsidy Information" and "Intelligent Industry Innovation Platform Construction Support Information" to the enterprise's reserved mailbox in the form of email. When there is an "urgent notice of high-tech enterprise application extension" such as emergency information, the system will immediately send key information to the relevant person in charge of the enterprise through SMS.

[0117] This step meets the diversified information receiving needs of enterprises, ensures that information reaches enterprises in a timely manner, improves the utilization rate of information, and enhances the response capability of enterprises to information. It is the output link of the matching result of information, directly providing services to enterprises, and the push effect will affect the evaluation and feedback of enterprises on the entire system.

[0118] Personalized push:

[0119] According to the browsing history, interest selection and real-time demand changes of enterprises, personalized information push is realized. If an enterprise browses a certain type of information times, the total number of browses is , the browse duration is , the total browse duration is , the number of clicks is , and the total number of clicks is , then the attention weight of this type of information is The calculation formula is:

[0120] ;

[0121] Among them, , , are weight coefficients, the value range is [0, 1], and . The system prioritizes the push of information with high weight categories according to the attention weight. Considering that the attention degree of enterprises to information can be reflected from multiple dimensions such as browse times, browse duration, and click times, the attention weight is calculated by weighting and combining these dimensions. Weight coefficients , , are used to adjust the importance of each dimension in attention evaluation to adapt to the focus of different enterprises on information.

[0122] In the calculation of attention weight At this time, the feedback of enterprises to the information combination, the matching of the information combination and the enterprise planning are taken into consideration. If the feedback of enterprises to a certain type of information combination is positive, and the combination is highly consistent with the current planning of the enterprise, the weight of this type of information combination in subsequent push is increased. At the same time, combined with the current development stage and strategic goal of the enterprise, the push strategy of the information combination is dynamically adjusted. For example, a company in stable development period and focusing on long-term innovation, the system preferentially pushes the information combination such as "long-term R&D fund support information" and "high-end talent training and introduction information" which are complementary and mutually promoting with the long-term planning of the enterprise.

[0123] For example, in the past month, a company (enterprise ) browsed the number of times of science and technology innovation information =20, the total number of times of browsing =30; the browsing time =100 minutes, the total browsing time =150 minutes; the number of times of clicking the link of science and technology innovation information =5, the total number of times of clicking =8. Set =0.4, =0.4, =0.2, then the attention weight of science and technology innovation information is:

[0124] ;

[0125] According to this weight, the system increases the push intensity of science and technology innovation information, and preferentially pushes "artificial intelligence frontier technology R&D fund support information" and "innovation achievement transformation reward information".

[0126] This step precisely fits the personalized needs of enterprises, improves the attention and participation of enterprises to the pushed information, improves the effectiveness of information push, and enhances the satisfaction of enterprises to the system. Based on the enterprise portrait and the matching result of information, further optimization is carried out, and the collected enterprise feedback data can be used to perfect the enterprise portrait and optimize the information matching algorithm, forming a closed-loop optimization mechanism.

[0127] Visualization and interaction:

[0128] The design uses visualization tools such as Tableau or PowerBI to present the information information push results to enterprise users in an intuitive way. In addition to displaying key information and matching degree of single information information, it highlights the complementary relationship, synergistic effect and matching degree of information information combination with enterprise planning. Through the relationship network diagram, the complementary relationship between information information in the information information combination is displayed, the expected change of enterprise benefit before and after the implementation of information information combination is presented by contrast column chart, and the matching degree of information information combination with short-term and long-term planning of enterprise is displayed by radar chart. At the same time, the interactive interface with adjustable parameters is provided, and the enterprise can independently select different information information combination to real-time view the complementary analysis, benefit prediction and planning matching degree analysis of the combination, so as to help the enterprise deeply understand the value of information information combination and make scientific decisions.

[0129] For example, when a company views the push results of "artificial intelligence industry tax preference information information" in the system, it can clearly see the key information such as tax reduction amount and application conditions of the information information, as well as the matching degree score of the enterprise and the information information through the column chart. The heat map intuitively shows that the number of artificial intelligence patents and the size of the R&D team of the enterprise have a greater contribution to the matching of the information information. The user adjusts the application threshold of "patent quantity" in the information information through the sliding bar, and the system updates the matching degree score and heat map in real time to show the changes of the contribution of each feature, helping the enterprise to intuitively understand the matching effect of the information information under different conditions.

[0130] This step presents the information information push results in an intuitive and easy-to-understand way, helps enterprise users to deeply understand the information information matching basis, and enhances the user's trust in the system recommendation results; interactive operation enables enterprises to actively explore the matching situation under different information information conditions, provides reference for enterprise development strategy, and improves user experience. It is the display and interaction of information information push, which visualizes the results of information information matching and personalized push to users; the feedback data of users in the interaction process can be used to optimize the information information matching algorithm and personalized push strategy, and promote the continuous improvement of the system.

[0131] User feedback collection:

[0132] A perfect user feedback mechanism is established in the system, and a special feedback entrance is set up. In addition to collecting enterprise's evaluation of single information information, it focuses on collecting enterprise's feedback on information information combination, including whether the complementarity of information information combination is reasonable, whether the synergistic effect meets the expectation, and whether the matching degree with enterprise planning meets the expectation. At the same time, regular user surveys are carried out through questionnaires, interviews and other forms to deeply understand the problems and improvement suggestions encountered by enterprises in the process of applying information information combination.

[0133] For example, after each information push, a company can evaluate the information through an in-system feedback form, selecting options such as "very relevant," "relatively relevant," or "not relevant," and providing specific comments. The system also sends monthly surveys to companies, asking about their overall satisfaction with the recently pushed information and the clarity of the information's interpretation. For example, one company reported that while a specific "AI product export subsidy information" had a high degree of match, the actual application process was too complicated, and requested more detailed application guidance and precautions.

[0134] This step directly captures companies' evaluations and requirements for the system, providing a reliable basis for system optimization. This helps identify issues in the information matching and delivery process, enabling targeted improvements and enhancing the system's service quality. The collected feedback will be used to optimize various aspects of information parsing, matching, and delivery, forming a continuous improvement cycle.

[0135] Model optimization and strategy adjustment:

[0136] According to user feedback, the information matching model and push strategy are optimized and adjusted. If the enterprise feedbacks that the complementarity of a certain information combination does not meet expectations, the information complementarity matrix is ​​reviewed and the complementarity coefficient of the relevant information is adjusted; if the user is not satisfied with the matching degree between the information combination and the enterprise plan, the calculation model for matching the enterprise plan with the information is optimized; if the user puts forward new requirements on the timing and method of pushing the information combination, the push strategy is optimized in a timely manner. For example, if the enterprise reflects that a certain information combination has a low matching degree with the long-term plan, the system will re-analyze the relationship between the enterprise's long-term planning goals and the information goals, and adjust If a company reports that a certain information package was delivered too late, missing the optimal application window, the system will optimize its real-time stream processing mechanism, strengthen monitoring of information releases and changes in company needs, and ensure timely delivery of information packages. Through continuous feedback and optimization, the accuracy and effectiveness of information package recommendations are continuously improved.

[0137] For example, if multiple companies report that the matching of "AI enterprise talent subsidy information" is inaccurate, analysis shows that the weight of the "R&D personnel professional qualifications" indicator is set too high, resulting in deviations in the matching results for some companies. =0.5 adjusted to =0.3, and recalculate the weights of other indicators to optimize the matching model. If most companies express a desire to receive information push notifications at the beginning of each month, the system will uniformly adjust the push notification time to the 1st to 3rd of each month.

[0138] This step enables the system to quickly adapt to changes in enterprise needs, continuously improve the accuracy and effectiveness of information matching and pushing, maintain the competitiveness and practicality of the system, and better meet the needs of enterprises to obtain information. It is an optimization and improvement of the entire information pushing process. Based on user feedback, the models and strategies involved in information analysis, enterprise portrait construction, information matching, and other steps are adjusted to improve the overall performance of the system. The results of optimization will affect subsequent user feedback and system operation effects, driving the system to evolve continuously.

[0139] Cloud computing and edge computing collaboration:

[0140] The architecture mode of cloud computing and edge computing collaboration is adopted to fully leverage the advantages of both. Cloud computing, with its powerful computing and storage capabilities, undertakes large-scale data processing and complex computing tasks such as deep learning analysis of information text and model training. Lightweight models are deployed on edge devices (such as enterprise local servers and intelligent terminals) to achieve fast local data processing and preliminary information matching, reduce data transmission delay, and improve system response speed. Using the federal learning fine-tuning mechanism, the latest cloud models are synchronized to edge devices regularly to ensure that the models on edge devices are consistent with the cloud, adapting to the dynamic changes of information and enterprise data.

[0141] For example, in the information text analysis task, the cloud computing platform uses its powerful computing resources to train and optimize the BERT model for massive information text, continuously improving the model's understanding of information semantics. A company deploys a lightweight BERT model obtained through knowledge distillation on the enterprise local server to perform preliminary keyword extraction and information classification on daily technical documents and project reports submitted by the enterprise, quickly determining whether they may be related to certain information. Every week, the cloud synchronizes updated model parameters to the edge device through the federal learning fine-tuning mechanism, ensuring that the edge device model can adapt to changes in new information types and information content in a timely manner. For example, when new artificial intelligence ethics-related information is published, the cloud model is updated, and the edge device can quickly obtain new model parameters to make more accurate judgments on related text.

[0142] This step realizes the reasonable allocation and efficient use of computing resources, improves the efficiency and response speed of the system in processing data, and reduces the cost and delay of data transmission; the federated learning fine-tuning mechanism ensures the timeliness and accuracy of the edge device model, so that the system can run stably under different network environments and device conditions, providing smooth and efficient service experience for enterprises. It provides strong computing and storage support for information analysis, enterprise portrait construction, information matching and pushing, etc., ensuring efficient data processing and model training; edge computing realizes local preliminary processing and fast response, optimizing user experience, closely related to information pushing, improving the timeliness of information pushing; at the same time, it also provides a technical foundation for feedback optimization, facilitating the collection and processing of user feedback data, and realizing the continuous optimization of the system.

[0143] Real-time stream processing technology:

[0144] Using real-time stream processing technologies such as FlinkML, dynamic data of enterprises and information is monitored and processed in real time. When new information is published or enterprise information is updated, real-time data processing is triggered to update the information intelligent database and enterprise information database in time, ensuring that the data used for information matching and pushing is always up-to-date. Set a trigger threshold (such as enterprise revenue fluctuation amplitude exceeding 10%, new patent application quantity change, etc.), when the threshold is reached, automatically trigger model retraining, so that the system can quickly adapt to changes in enterprise data, improve the accuracy of information matching.

[0145] For example, the revenue data and patent application data of a certain company are transmitted to the system in real time. When the enterprise's revenue in a certain quarter increases by 15% compared with the previous quarter (exceeding the set threshold of 10%), the system immediately starts the data processing process, updates the revenue data in the enterprise information database, and triggers the retraining of the information matching model, re-evaluates the matching degree between the enterprise and the existing information. At the same time, if a new "artificial intelligence enterprise revenue growth step-by-step reward information" is published at this time, the system can grab the information in real time, update the information intelligent database, and based on the latest enterprise data and information, re-match and push the information to the company in time.

[0146] This step ensures that the system can keep up with the dynamic changes of information and enterprise data, providing the latest and most accurate information matching and pushing services to enterprises, avoiding missing information opportunities due to data lag, greatly improving the practicality and timeliness of the system, and enhancing the dependence of enterprises on the system. Real-time updating of information and enterprise data provides the latest data support for information analysis, information matching and pushing, and is a key technology support for precise and timely pushing; the model retraining mechanism ensures the accuracy of the information matching model, closely linked to the information matching and pushing process; at the same time, it also provides a real-time data basis for feedback optimization, facilitating the collection and analysis of user feedback and promoting the continuous improvement of the system.

[0147] Model compression and acceleration technology:

[0148] Lightweight meta-learning and model compression techniques are used to improve system performance. Meta-LSTM is used as a lightweight meta-learning model to effectively reduce model parameter size. Knowledge distillation technology is used to transfer the knowledge of complex models (such as XGBoost+CNN) to small models (such as MobileNet). Quantized models (such as INT8 precision) are deployed on edge devices, and only complex tasks are uploaded to the cloud for processing. TensorRT or ONNXRuntime is used to replace traditional deep learning inference frameworks to improve edge device computing efficiency and reduce system running costs.

[0149] For example, during the training of the information matching model, a complex XGBoost+CNN model is first trained to extract and match information features. Then, knowledge distillation technology is used to transfer the knowledge learned by the complex model to the MobileNet model. The original complex model has 10 million parameters, and the migrated MobileNet model has a parameter size of 2 million. On the edge device of a certain company, the MobileNet model is deployed with INT8 precision quantization, and the TensorRT inference framework is used for acceleration. Originally, it takes 5 seconds to perform one information matching inference using a traditional deep learning inference framework, but after optimization, the inference time is reduced to 1 second, greatly improving the processing efficiency. For complex information matching tasks involving multi-dimensional data analysis and complex causal relationship judgment, the task is uploaded to the cloud for processing.

[0150] The step significantly reduces the computational burden and energy consumption of the edge device, greatly improves the model inference speed, and effectively reduces the system operation cost; under the premise of ensuring the accuracy of the model, the lightweight and acceleration of the model are realized, so that the system can run efficiently on resource-limited devices, improve the scalability and adaptability of the system, and ensure the efficiency of information matching and pushing. Provide efficient model support for information analysis, information matching and pushing, especially in the edge computing scenario, improve the efficiency of local data processing and preliminary information matching. Cooperate with cloud computing and edge computing collaborative technology, optimize the overall performance of the system. Model compression and acceleration technology can reduce the resource occupation of the model on the edge device, so that the edge device can run the lightweight model more smoothly, combined with real-time stream processing technology, to ensure the efficiency of the system when processing dynamic data, and to ensure the timeliness and accuracy of the information pushing. At the same time, it also provides a more efficient model basis for feedback optimization, which facilitates the rapid processing of user feedback data and promotes the continuous optimization of the system.

[0151] The processes described above with reference to the flowcharts can be implemented as computer software programs in accordance with embodiments of the present disclosure. Embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication section, and / or installed from a detachable medium. When the computer program is executed by a central processing unit (CPU), the above-described functions defined in the methods of the present application are performed. It should be noted that the computer readable medium of the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but not limited to, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device. In the present application, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that can send, propagate or transfer a program for use by or in connection with an instruction execution system, apparatus or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wireless, wire, optical cable, RF, or any suitable combination of the above.

[0152] The computer program product of the present application can be a computer program product comprising a computer-readable medium bearing computer program code embodied therein for use with a computer. The computer program code can be code defining and / or implementing the present application. The computer program code can be written in any suitable computer readable programming language. The computer program code can be stored in a computer- readable storage medium, such as, but not limited to, any type of disk including an optical disk, a CD-ROM, a CD-R, a CD-RW, a DVD, a flash memory, a ROM, a RAM, a magnetic disk or hard drive, or any other suitable type of medium including a medium that holds the software for a particular or specialized computing purpose, or any suitable combination of media. The computer program product can be a computer program product distributed to end users, whether as a stand-alone program, as part of a physical system, or as a software download. The computer program product can be distributed on a physical medium, such as, but not limited to, a floppy disk, a CD-ROM, a CD-R, a CD-RW, a DVD, a flash memory, a ROM, a RAM, a magnetic disk or hard drive, or any other suitable type of medium, or any suitable combination of media. The computer program product can be distributed from a program distribution center, either as a tangible medium or via electronic delivery, such as from a Web site via the Internet, or from one computer to another via electronic transfer, such as by e-mail. The computer program product can be distributed in an encrypted manner, such as via encryption or via password protection.

[0153] Those skilled in the art will understand that the application described above and illustrated in the accompanying drawings is presented by way of example only and is not limiting as to the present application. The intent is to cover all modifications and alternatives of the present application falling within the scope of the application.

Claims

1. A method for matching and pushing information based on big data, characterized in that: The method comprises: Use web crawler technology to capture information text data, use API interfaces to obtain internal enterprise data, integrate multi-channel information text data and enterprise operating data and perform data pre-processing; Use automatic update tools and NLP to build entity recognition models, set thresholds, filter low-quality data from pre-processed data, and detect and correct contradictory data; Extract key information from news through natural language processing technology and automatically classify news using text classification algorithms; Through data analysis and machine learning technology, we build a multi-dimensional enterprise portrait based on enterprise content data, determine quantitative indicators for information types, build an information impact model, analyze the impact of information on the enterprise, and build an information complementarity matrix between different information types; Construct a single information matching model based on the matching algorithm, calculate the matching degree between the single information and the enterprise, construct an information matching model based on the matching degree between the single information and the enterprise, the information complementarity matrix, and the matching degree between the information and the enterprise plan, and calculate the final information matching degree; Push information with a final matching degree exceeding the matching degree threshold and the analysis results of the information impact model to the enterprise, collect user feedback, and optimize the information matching model and push strategy based on user feedback; The final information matching calculation formula is as follows: ; in: The final information matching degree; is the collaborative filtering adjustment coefficient; For the matching degree between single information and enterprise, To traverse the information variable, from 1 to Take the values ​​in turn, is the total number of information, For information and information The degree of complementarity, is the weight coefficient of short-term and long-term planning, The matching coefficient for the enterprise's short-term planning, Enterprise long-term planning matching coefficient.

2. The method for matching and pushing information based on big data drive according to claim 1, characterized in that: The data preprocessing comprises the following steps: For websites that are updated frequently, crawling is done once a day or hourly. For industry association websites that are updated slowly, crawling is done once a week or month. A random delay mechanism is used to crawl information text data from official government websites, government affairs platforms, and industry association websites. With the help of API interfaces, relevant data from the internal systems of enterprises are obtained, and the basic information, operating data, and industry data of enterprises are integrated. The collected data is cleaned to remove duplicate and invalid data. Missing values ​​are filled with mean value and regression prediction methods. The data is standardized. The data structure and format are unified to establish the relationship between data.

3. The method for matching and pushing information based on big data drive according to claim 2, characterized in that: The specific process of filtering the low-quality data is as follows: Use the BERT model to process information text data, identify key entities in the information text data and related entities in the enterprise internal data through entity recognition, and use relationship extraction technology to establish the connection between the two entities to build a dynamic knowledge graph; Set data quality thresholds to automatically filter low-quality data; use the knowledge graph's reasoning engine to detect contradictory data. When contradictory data is detected, trigger manual review or automatic correction to correct the contradictory data.

4. The method for matching and pushing information based on big data drive according to claim 3 is characterized in that: The specific process of automatic classification of information is as follows: Use keyword extraction algorithms to extract keywords, titles, and key sentences as text features, and extract release time, release media, and related company attributes as structured features; use word vector models to convert text features and structured features into numerical vectors; use supervised learning models to collect and label training data, and train the model to learn features and category classification; evaluate model performance through test sets, adjust parameters based on the results, and optimize features and models; input numerical vectors into the model for classification, and store the results by category.

5. The method for matching and pushing information based on big data drive according to claim 4, characterized in that: The information impact model includes the following formula: The total impact of technological innovation results It is the sum of direct impact and indirect impact. The specific formula is as follows: ; Total impact on corporate revenue It is the sum of direct effect and moderating effect. The specific formula is as follows: ; in: is the intercept, 、 、 are the influence coefficients of the corresponding variables, is the error term, is the original revenue base value of the enterprise, is the information intensity, For enterprise scale, is the intercept term, Information intensity The impact coefficient on technological innovation results, is the error term, is the random error term, is the intercept, is the influence coefficient, is the error term, is the error term, is the intercept, Influence coefficient.

6. The method for matching and pushing information based on big data drive according to claim 5, characterized in that: The specific process of the big data-driven information matching and push method is as follows: Arrange the information matching results in descending order according to the final matching degree; set the matching degree threshold , filter out information below the threshold; provide enterprises with sorting options based on different dimensions such as the strength of information complementarity, comprehensive benefits, and matching degree with enterprise plans, and push matched information to enterprises through email, text messages, and system site messages. Enterprises can set their preferred push method and time in the system; based on the company's browsing history, interest selections, and real-time demand changes, combined with the company's current development stage and strategic goals, dynamically adjust the information combination push strategy.

7. The method for matching and pushing information based on big data drive according to claim 6, characterized in that: The optimization process of the information matching model and push strategy is as follows: Set up a feedback portal to collect enterprise feedback on single information and information combinations, including whether the complementarity of the information combination is reasonable, whether the synergy benefits meet expectations, and whether the match with the enterprise plan meets expectations; Conduct user surveys regularly to understand the problems and improvement suggestions encountered by enterprises in the actual application of information combination; If the enterprise feedbacks that the complementarity of a certain information combination does not meet expectations, review the information complementarity matrix and adjust the complementarity coefficient of the relevant information; if the user is not satisfied with the matching degree between the information combination and the enterprise plan, optimize the calculation model of the matching between the enterprise plan and the information; if the user puts forward new requirements on the timing and method of pushing the information combination, optimize the push strategy in a timely manner.

8. An information matching and push system based on big data, characterized in that: The system is used to execute the information matching and pushing method based on big data drive as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the information matching and pushing method based on big data drive as described in any one of claims 1 to 7 above.

Citation Information

Patent Citations

  • Investment and research information sharing and distributing method for group business management and control

    CN119719444A