Information message matching and pushing method and system based on big data driving
Through the big data-driven information matching method, multi-channel data is integrated to build corporate portraits and information impact models, the integration and matching problems of enterprises when dealing with multi-channel information information is solved, accurate information push and decision-making support are achieved, and corporate competitiveness is enhanced.
Patent Information
- Application Number
- CN202510741258.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-05
AI Technical Summary
When enterprises process multi-channel and multi-format information, it is difficult for them to fully collect and effectively integrate it. The existing matching and push systems lack personalization, resulting in the under-exploration of information value and data quality and processing efficiency problems.
Using a big data-driven method, information information is obtained through network crawlers and API interfaces, combined with data preprocessing, natural language processing and machine learning, a corporate portrait and information impact model is built, information matching is calculated, and personalized information information is pushed through multiple channels.
It improves the accuracy and efficiency of information matching, enhances corporate decision-making support, reduces information screening costs, ensures that information and information are highly compatible with the company, and enhances market competitiveness.
Smart Images

Figure CN120277271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information screening, and particularly relates to a method and system for matching and pushing information based on big data drive. Background Art
[0002] In the digital wave, the development of enterprises is increasingly dependent on information, and its acquisition and processing capabilities have become an important part of the competitiveness of enterprises. However, enterprises face severe challenges in information processing. On the one hand, the sources of information are extensive, covering multiple channels such as governments, industry associations, and enterprise internal systems. The formats are diverse and scattered, with structured, semi-structured, and unstructured data coexisting. This makes it difficult for enterprises to comprehensively collect and effectively integrate, increasing the processing difficulty.
[0003] Traditional information processing methods mostly rely on manual screening and simple tools, with low efficiency and easy errors. Especially when dealing with a large amount of unstructured text data, it is difficult to accurately extract key information and understand semantics, resulting in the inability to fully exploit the value of information. Existing matching and pushing systems have insufficient understanding of the personalized needs of enterprises, and the pushed information lacks pertinence. Although emerging technologies such as big data, artificial intelligence, and blockchain have brought opportunities, there are also problems, such as data quality and processing efficiency problems in big data, accuracy and interpretability problems of artificial intelligence models, performance bottlenecks and compatibility problems of blockchain, etc. Therefore, researching and developing an innovative information processing solution to integrate multi-source data, ensure data security, and achieve accurate parsing and matching is crucial for enterprises to make rapid and accurate decisions in a complex market environment. Summary of the Invention
[0004] The present invention comprehensively considers various factors to calculate the matching degree, improves the matching degree between the screened information and the enterprise, and enhances the satisfaction of the enterprise.
[0005] The technical solution proposed by the present invention is: a method for matching and pushing information based on big data drive, the method comprising: Using web crawler technology to capture information text data, obtaining enterprise internal data through API interfaces, integrating multi-channel information text data and enterprise operation data, and performing data preprocessing; Using an automatic update tool and NLP to build an entity recognition model, setting a threshold, filtering low-quality data in the preprocessed data, and detecting and correcting contradictory data; Extracting key information of information through natural language processing technology, and automatically classifying information using text classification algorithms; Through data analysis and machine learning techniques, construct a multi-dimensional enterprise portrait based on enterprise content data, determine quantitative indicators for information types, construct an information impact model, analyze the impact of information on the enterprise, and construct an information complementarity matrix between different information; Construct a single information matching model according to the matching algorithm, calculate the matching degree between the single information and the enterprise, and construct an information matching model based on the matching degree between the single information and the enterprise, the information complementarity matrix, and the matching between the information and the enterprise plan, and calculate the final information matching degree; Push the information with the final information matching degree higher than the matching degree threshold and the analysis results of the information impact model to the enterprise, collect user feedback, and optimize the information matching model and the push strategy according to the user feedback.
[0006] Preferably, the data preprocessing includes the following steps: For websites with frequent updates, crawl once a day or once an hour. For industry association websites with slow updates, crawl once a week or once a month. Adopt a random delay mechanism to crawl information text data from government official websites, government affairs platforms, and industry association websites; obtain relevant data in the enterprise internal system through the API interface, and integrate the basic information, business data, and industry data of the enterprise; clean the collected data to remove duplicate and invalid data; fill in missing values using the mean filling and regression prediction methods; standardize the data; unify the data structure and format, and establish the association relationship between data.
[0007] Preferably, the specific process of low-quality data filtering is as follows: Use the BERT model to process the information text data, determine the key entities in the information text data and the relevant entities in the enterprise internal data through entity recognition, use the relationship extraction technology to establish the connection between the two entities, and construct a dynamic knowledge graph; set the data quality threshold to automatically filter low-quality data; use the inference engine of the knowledge graph to detect contradictory data. When contradictory data is detected, trigger manual review or automatic correction to correct the contradictory data.
[0008] Preferably, the specific process of automatic classification of information is as follows: Use the keyword extraction algorithm to extract keywords, titles, and key sentences as text features, and extract the release time, release media, and associated enterprise attributes as structured features; convert the text features and structured features into numerical vectors through the word vector model; collect training data and label through the supervised learning model, and train the model to learn the features and category classification; evaluate the model performance through the test set, adjust parameters, optimize features and models according to the results; input the numerical vectors into the model for classification, and store the results by category.
[0009] Preferably, the information influence model includes the following formula: Total influence of technological innovation achievements Is the sum of the direct influence and the indirect influence. The specific formula is as follows: ; Total influence of enterprise revenue Is the sum of the direct influence and the moderating effect influence. The specific formula is as follows: ; Where: Is the intercept, , , Are the influence coefficients of the corresponding variables respectively, Is the error term, Is the original revenue base value of the enterprise, Is the information intensity, Is the enterprise scale, Is the intercept term, Is the information intensity The influence coefficient on technological innovation achievements, Is the error term, Is the random error term, Is the intercept, Is the influence coefficient, Is the error term, Is the error term.
[0010] Preferably, the formula for calculating the final information matching degree is as follows: ; Where: Is the final information matching degree; Is the collaborative filtering adjustment coefficient; Is the matching degree between a single piece of information and the enterprise, Is the variable for traversing information, Is the total number of information, Is the information And the information The degree of complementarity, Is the short-term and long-term planning weight coefficient, Is the enterprise short-term planning matching coefficient, Enterprise long-term planning matching coefficient.
[0011] Preferably, the specific process of the information push method is as follows: Arrange the information matching results in descending order according to the final matching degree; Set the matching degree threshold , filter out information below the threshold; provide the enterprise with sorting options in different dimensions such as the complementary strength of information, comprehensive benefits, and the degree of matching with the enterprise's plan, and push the matched information to the enterprise via email, SMS, and in-system messages. The enterprise can set the preferred push method and time in the system; based on the enterprise's browsing history, interest checkboxes, and real-time demand changes, combined with the enterprise's current development stage and strategic goals, dynamically adjust the combined push strategy of information.
[0012] Preferably, the optimization process of the information matching model and push strategy is as follows: Set up a feedback entry to collect the enterprise's feedback on single information and combined information, including whether the complementarity of the combined information is reasonable, whether the synergy effect meets the expectation, and whether the degree of matching with the enterprise's plan meets the expectation; regularly conduct user research to understand the problems and improvement suggestions encountered by the enterprise in the actual application of combined information; if the enterprise feedbacks that the complementarity of a certain combined information does not meet the expectation, re-examine the information complementarity matrix and adjust the complementarity coefficient of relevant information; if the user is not satisfied with the degree of matching between the combined information and the enterprise's plan, optimize the calculation model for the matching of the enterprise's plan and information; if the user puts forward new requirements for the push timing and method of the combined information, optimize the push strategy in a timely manner.
[0013] The present invention also provides a big data-driven information matching and pushing system, which is used to execute the big data-driven information matching and pushing method described above.
[0014] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the big data-driven information matching and pushing method described above.
[0015] Advantages of the present invention: 1. Collect information text data and enterprise internal operation data through multiple channels such as web crawlers and API interfaces, and use preprocessing technologies such as data cleaning and standardization to eliminate data quality differences and format barriers. At the same time, establish a unified data storage and management architecture to break data islands and enable efficient collaboration of data from different sources and different types. This provides a rich and complete data foundation for subsequent analysis, enables the model to learn more comprehensive information features, significantly improves the prediction accuracy and generalization ability of the model, and further enhances the accuracy of information matching and pushing, helping enterprises quickly obtain valuable information, reducing information screening costs, and improving decision-making efficiency.
[0016] 2. By constructing a complex causal relationship model, comprehensively considering the relationships among various variables, and using machine learning algorithms for data-driven modeling and parameter optimization. This innovative analysis method can deeply explore the internal connection between information and enterprise development, and accurately quantify the impact of information on all aspects of the enterprise. Based on this, the prediction, combined evaluation, and counterfactual reasoning of information effects provide comprehensive decision-making support for the enterprise. The enterprise can predict in advance the effects after the implementation of information, evaluate the synergistic benefits of different information combinations, and understand the development situation when a certain information is not applied, so as to formulate a development strategy more scientifically, allocate resources reasonably, avoid the risks brought by blind decision-making, and improve the market competitiveness of the enterprise.
[0017] 3. Construct a perfect quantitative index system to accurately measure the matching degree between information and the enterprise. In the design of the matching algorithm, comprehensively consider the similarity between the enterprise portrait and the applicable conditions of the information, the complementarity between information, and the matching with the enterprise plan to ensure that the matching results are comprehensive and accurate. The sorting and screening mechanism of the matching results can display the most valuable information according to the enterprise's needs. The multi-channel push and personalized push functions can achieve the accurate delivery of information according to the enterprise's browsing history, interest preferences, etc. The visualization and interaction design enable the enterprise to intuitively understand the content of the information and the basis for matching, while the user feedback and model optimization mechanism continuously improve the accuracy of matching and the effectiveness of pushing. This series of innovative measures greatly improve the efficiency of the enterprise in obtaining and using information, reduce the cost of the enterprise in screening information, enable the enterprise to quickly obtain information highly compatible with itself, and promote the development of the enterprise. Brief Description of the Drawings
[0018] Figure 1 It is a flowchart of a method for matching and pushing information based on big data drive according to the present invention; Figure 2 It is a flowchart of the information pushing process of a method for matching and pushing information based on big data drive according to the present invention. Detailed Embodiment
[0019] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations. The basic principles defined in the following description can be applied to other implementation schemes, variant schemes, improvement schemes, equivalent schemes, and other technical schemes without departing from the spirit and scope of the present invention.
[0020] It is understood that the term "a" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "a" should not be understood as a limitation on the quantity.
[0021] As Figure 1 and Figure 2 shown, first, data collection and preprocessing are carried out, and multi-source data fusion is performed. Specific rules are set for web crawler technology to capture the content of information web pages and obtain relevant data in the enterprise internal system through the API interface. According to the update frequency and data importance of the target website, a reasonable crawling frequency is set. For websites with frequent updates, such as government policy release platforms, they are crawled once a day or once an hour; for industry association websites with relatively slow updates, they are crawled once a week or once a month. At the same time, a random delay mechanism is adopted to avoid making a large number of requests to the target website in a short period of time and being blocked by the website's IP. To reduce the amount of data crawled and improve the crawling efficiency, an incremental crawling strategy is adopted. Record the time point or data identifier of each crawl. When crawling next time, only obtain updated or new data. For example, by comparing the release time or version number of information articles, it is judged whether it is new data, and only new and updated information is crawled, reducing the load pressure on the target website and improving the timeliness of data update. Integrate the information text data collected from multiple channels such as government official websites, government affairs platforms, and industry association websites, and at the same time integrate data such as the basic information, business data, and industry data of the enterprise. To ensure the accuracy and integrity of the data, the collected data is cleaned to remove duplicate and invalid data, and methods such as mean filling and regression prediction are used to fill in the missing values, thereby improving the data integrity. Multi-source data integration is the basis of the entire solution, providing raw data support for subsequent steps such as information text parsing, information quantification, and modeling. The quality of the data directly affects the accuracy of subsequent analysis and matching.
[0022] Adopt the combination of federated learning and blockchain to ensure enterprise data privacy in the data sharing link. Use multi-party secure computing (MPC) to optimize the federated learning communication mechanism. Enterprise data is encrypted and calculated locally, and only model gradients are shared. Assume that there are enterprises participating in federated learning. The local data of enterprise is , the model parameters are , and the gradient calculated based on the local data is , where is the loss function of enterprise based on the local data. The federated learning server aggregates the gradients of each enterprise To update global model parameters. Data usage logs are recorded through blockchain, and data update tasks are automatically triggered through smart contracts to ensure the traceability and compliance of data sharing. This step protects the privacy and security of enterprise data and promotes data sharing, thereby solving the problem of data silos, enabling model training to use more data, and improving the accuracy and generalization of the model. At the same time, it ensures the compliance and traceability of data use, provides a secure data source for subsequent information quantification and modeling, enables model training to be based on richer data, improves model quality, and thus affects the accuracy of information matching and push. At the same time, the data usage logs recorded by the blockchain can also provide a reference for user feedback and system optimization.
[0023] Build dynamic knowledge graphs and clean data: Introduce knowledge graph automatic update tools (such as RDFox or GraphDB), combined with NLP entity recognition models (such as spaCy or BERT), to capture news information and enterprise dynamic data in real time. Through entity recognition, determine the key entities in the information (such as the subject of the information, scope of application, support measures, etc.) and the related entities in the enterprise information, and use relationship extraction technology to establish the connection between the two entities, so as to build a dynamic knowledge graph. Set data quality thresholds (such as accuracy > 95%) to automatically filter low-quality data; use the reasoning engine of the knowledge graph (such as RDFS or OWL) to detect contradictory data. When contradictory data is detected, trigger manual review or automatic correction to correct the contradictory data. In entity recognition, use the BERT model to process the information text and output the feature vector of each word. , For the single feature vectors identified, a specific classifier is used to determine whether the word is an information entity based on these feature vectors.
[0024] This step ensures the real-time, accuracy and consistency of the data in the knowledge graph, provides rich semantic information for information text analysis, improves the availability and comprehensibility of information, and facilitates subsequent information matching and analysis. This step provides structured knowledge support for information text analysis, helping to more accurately extract key information; at the same time, its data cleaning function provides high-quality data for information quantification and modeling, ensuring the reliability of model training.
[0025] Taking a certain company as an example, the system uses web crawlers to regularly capture various information texts from the websites of local government science and technology departments and industry and information technology departments, covering fields such as technological innovation and industrial support. At the same time, by docking with the API of the enterprise's internal management system, it obtains business data such as the enterprise's R & D investment, number of patents, revenue, and profit. During the data cleaning process, it is found that some patent data has duplicate records, which are deduplicated after verification; for the missing revenue data of a certain quarter, according to the enterprise's past revenue trends and the revenue situations of similar enterprises in the same industry, a regression prediction method is used for filling. Thus, the quality of the data is ensured, providing a reliable data basis for subsequent information parsing and matching, and avoiding analysis biases caused by data errors or omissions.
[0026] In a technology enterprise alliance within a region, this company and many other technology enterprises jointly participate in federated learning to train the same information matching model. Use local R & D data, market data, etc. to encrypt and calculate the model gradients locally. For example, based on the data of its own R & D investment and the number of new product launches, calculate the gradient information about a certain technology achievement transformation support information matching model and upload it encrypted. The federated learning server aggregates these gradients to update the global model, and the blockchain details information such as the time and data volume of each enterprise's contributed data. Through smart contracts, when there is new information or enterprise data update, the data update task is automatically triggered.
[0027] For an information about artificial intelligence industry support, the knowledge graph automatic update tool retrieves the information from the government website in real time. The BERT model identifies key entities such as "artificial intelligence enterprises", "R & D subsidies", and "technological innovation indicators", and determines the "obtaining subsidy" relationship between "artificial intelligence enterprises" and "R & D subsidies" through relation extraction technology, thus constructing a knowledge graph. If it is found that the accuracy of a certain technological innovation data of a company is only 90% (lower than the 95% threshold), the data is automatically filtered; if there is a contradictory information in the knowledge graph that the enterprise both meets the information of high-pollution industry restrictions and obtains environmental protection subsidies (assuming it is caused by data entry errors), an artificial review process is triggered.
[0028] Next, parse the collected information texts, conduct information interpretation and analysis through artificial intelligence, and use natural language processing (NLP) technology to deeply interpret the information texts. The information texts are split into words through lexical analysis and part-of-speech tagging is performed; the grammatical structure of sentences is constructed through syntactic analysis; the text semantics is understood through word vector models (such as Word2Vec, GloVe) and deep learning models (such as Transformer, BERT). Taking the BERT model as an example, pre-train the information texts to capture context semantic information. Assume the information text is , and the semantic representation is obtained after being processed by the BERT model , identify key terms, restrictive conditions, preferential measures, etc. in the information by means of this semantic representation. Extract the core information of the information, such as the information name, issuing agency, issuing time, applicable objects, content summary of the information, etc. through information extraction technology, and store it in the intelligent database of information for subsequent information matching to provide a data basis.
[0029] For example, for the information text "Encourage artificial intelligence enterprises to increase R & D investment. For enterprises with an annual R & D investment exceeding 5 million yuan and having more than 10 artificial intelligence-related patents, a subsidy of 15% of the R & D investment will be given and special R & D funds will be provided", lexical analysis splits it into words such as "encourage", "artificial intelligence enterprises", "increase", "R & D investment" and marks their parts of speech; syntactic analysis determines the sentence structure; the BERT model understands its semantics and identifies that "annual R & D investment exceeding 5 million yuan" and "having more than 10 artificial intelligence-related patents" are key conditions, and "giving a subsidy of 15% of the R & D investment" and "providing special R & D funds" are preferential measures. Information extraction technology extracts information such as the information name (such as "R & D support information for artificial intelligence enterprises") and applicable objects (artificial intelligence enterprises that meet specific R & D investment and patent quantity requirements) and stores them in the database. This step can extract valuable structured information from unstructured information texts, convert the content of the information into a form that can be understood and processed by machines, and provide key data support for information quantification and modeling, as well as information matching. It is also a prerequisite for information quantification and modeling, providing accurate key information of the information; at the same time, the extracted information directly affects the accuracy of information matching and is an important basis for information matching.
[0030] Then, classify and structurally represent the collected data according to the parsed information. Establish a classification system based on dimensions such as the theme, industry, and applicable scope of the information. For example, classify by industry into manufacturing, service, technology, etc., and classify by information type into support information, tax information, regulatory information, etc. Use text classification algorithms (such as support vector machine SVM, convolutional neural network CNN) to automatically classify the information. Taking SVM as an example, assume that the feature vector of the information text is , through the SVM model predict the category to which the information belongs . Structurally represent the content of the information and convert it into a format that is easy for machines to process, such as XML, JSON. For example, represent an information on scientific and technological innovation support as a JSON object containing fields such as basic information of the information, applicable conditions, support methods, application processes, etc., which is convenient for subsequent matching calculations and queries.
[0031] For example, for a newly collected series of information texts, we first extract features such as keywords and sentence vectors from the text to construct feature vectors. . The trained SVM model is used to classify the "information on R&D support for artificial intelligence enterprises" and determine that it belongs to the category of "technology enterprises-support information". Then the information is structured and represented in JSON format: {"information name": "information on R&D support for artificial intelligence enterprises", "applicable conditions": "artificial intelligence enterprises with annual R&D investment exceeding 5 million yuan and more than 10 artificial intelligence-related patents", "support method": "a subsidy of 15% of R&D investment and special R&D funding support", "application process": "..."}. This step facilitates the management and retrieval of a large amount of information, improves the efficiency of information matching, makes the organization of information more orderly, provides clearer information classification information for information matching algorithms, and improves the accuracy and pertinence of matching. Providing classification information for information matching helps narrow the matching scope and improve matching efficiency; at the same time, the structured information data provides a standardized data format for information quantification and modeling, which is convenient for quantitative analysis.
[0032] Next, the information is quantified and modeled, and a quantitative indicator system for information is constructed. Different quantitative indicators are determined for different types of information. For fiscal subsidy information, quantitative indicators may include subsidy amount A, subsidy ratio R, application threshold (such as corporate revenue I, R&D investment RD and other requirements); for tax incentive information, the tax exemption amount T, preferential period P, etc. can be quantified. Through the analysis and data mining of information text, these quantitative indicators are extracted and standardized to make them comparable. Assuming that the subsidy amounts of multiple fiscal subsidy information are standardized, the z-score standardization method is used. The standardized subsidy amount ,in The average amount of subsidy, is the standard deviation of the subsidy amount. First, calculate the mean of the subsidy amount for similar information, then find the difference between the subsidy amount and the mean, and finally divide the difference by the standard deviation of the subsidy amount to get the standardized subsidy amount parameter. For example, the subsidy amount for a certain information is 3 million yuan, the mean of similar information is 1.6 million yuan, and the standard deviation is 836,700 yuan. Then the standardized subsidy amount is , which indicates that the subsidy amount of this information is about 1.67 standard deviations higher than the average level of the same category, and it is a relatively large subsidy information, which may receive a higher weight in the matching model. The standardization method of other quantitative indicators is similar.
[0033] For example, when analyzing a series of information on the support of science and technology industries, for the "information on the support of artificial intelligence enterprise R & D", the subsidy ratio R is determined to be 15%, and the enterprise R & D investment RD in the application threshold needs to exceed 5 million yuan. After standardizing these quantitative indicators, they are compared and analyzed with the quantitative indicators of other information, so as to match the information more scientifically. This step converts the information content into quantifiable indicators, which is convenient for mathematical calculation and model construction, making the information matching more scientific and accurate, and being able to measure the fit between the enterprise and the information more accurately. Providing a quantitative basis for information matching is the key step to achieve accurate matching; at the same time, the extraction of quantitative indicators depends on the results of information text parsing and is a further quantitative processing of the information parsed from the information text.
[0034] Causal reasoning and information effect modeling: In the scenario of information matching enterprises, a causal relationship network is constructed based on the Structural Causal Model (SCM) and combined with machine learning algorithms. The specific model structure is as follows: Variable definition Information variable ( ) Let the information variable be composed of the information type , information intensity , and information implementation time , that is . For example, for the "information on R & D subsidies for artificial intelligence enterprises", the information type is encoded as the "financial subsidy" category (using one-hot encoding, such as [1,0,0] represents financial subsidy, [0,1,0] represents tax preference, [0,0,1] represents talent support), the information intensity is expressed as the ratio of the subsidy amount to the R & D investment, and the information implementation time is the information release date (such as 2024-01-01).
[0035] Enterprise characteristic variable ( ) The enterprise characteristic variable includes enterprise scale (such as the number of employees, revenue scale, etc. The standardization formula for the number of employees is , and the revenue scale is standardized after logarithmic transformation), industry attribute (using category encoding, such as artificial intelligence technology enterprises are encoded as 001), enterprise strategy (the strategic goal filled in by the enterprise is vectorized by text, and is converted into a 768-dimensional vector using the BERT model), enterprise resources (such as capital reserve, talent reserve, number of technology patents, etc., and the capital reserve and talent reserve are also processed in a standardized manner), that is 。
[0036] Enterprise outcome variables ( ): Enterprise outcome variables include enterprise revenue , profit , market share , technological innovation achievements (such as the number of patents, the number of new products launched, etc.), that is 。
[0037] Causal relationship structure Direct causal relationship: Taking the impact of financial subsidy information on a company's R & D investment as an example, assuming a linear relationship, it can be expressed as , where is the technological innovation achievement (such as the number of patents) affected by the financial subsidy information, is the intercept term, is the intensity of the information, is the impact coefficient of the information on the technological innovation achievement, is the error term. In enterprise operation, information has a direct and crucial impact on enterprise revenue. Through the product of the information intensity ( ) and the impact coefficient ( ), the direct effect of information on technological innovation achievements is quantified. By the intercept term ( ), factors that are difficult to accurately measure but have a stable impact on technological innovation achievements are comprehensively considered. The error term incorporates random factors such as accidental inspirations in the scientific research process and external sudden technological exchanges, making the model more in line with the actual situation. This formula realizes the quantitative modeling of the direct impact of financial subsidy information on technological innovation achievements, providing key parameters for enterprises to evaluate the value of financial subsidy information on technological innovation. Referring to the actual operation data of many enterprises and industry research results, when an enterprise obtains and effectively utilizes financial subsidy information, it can often directly act on revenue in many aspects. Taking a high-tech enterprise as an example, after accurately grasping the financial subsidy information of the government for scientific and technological innovation, it successfully applied for a large amount of subsidy funds. On the one hand, this fund can be directly used to expand the production scale, purchase advanced production equipment, improve product output and quality, thereby increasing product sales and directly boosting revenue; on the other hand, the subsidy funds can be invested in market promotion activities, expand sales channels, improve brand awareness, attract more customers to buy products, and drive revenue growth. From the perspective of financial analysis, enterprise revenue can be reflected through the operating income in the income statement. Let the original revenue of the enterprise be the base value then incorporates random factors such as accidental inspirations in the scientific research process and external sudden technological exchanges, making the model more in line with the actual situation. This formula realizes the quantitative modeling of the direct impact of financial subsidy information on technological innovation achievements, providing key parameters for enterprises to evaluate the value of financial subsidy information on technological innovation. Referring to the actual operation data of many enterprises and industry research results, when an enterprise obtains and effectively utilizes financial subsidy information, it can often directly act on revenue in many aspects. Taking a high-tech enterprise as an example, after accurately grasping the financial subsidy information of the government for scientific and technological innovation, it successfully applied for a large amount of subsidy funds. On the one hand, this fund can be directly used to expand the production scale, purchase advanced production equipment, improve product output and quality, thereby increasing product sales and directly boosting revenue; on the other hand, the subsidy funds can be invested in market promotion activities, expand sales channels, improve brand awareness, attract more customers to buy products, and drive revenue growth. From the perspective of financial analysis, enterprise revenue can be reflected through the operating income in the income statement. Let the original revenue of the enterprise be the base value Based on the revenue change data of a large number of similar enterprises after obtaining subsidy information and successfully applying for subsidies, a linear regression model was established, and the direct impact of fiscal subsidy information on enterprise revenue was obtained as follows: ,in The coefficient that indicates the impact of financial subsidy information on revenue depends on factors such as industry characteristics, enterprise scale, and the strength of subsidy policies. is a random error term, reflecting the impact of other random factors not considered by the model on revenue. ) and the influence coefficient ( ) multiplied by ( ), accurately quantify the direct impact of information on revenue, the error term Factors such as sudden market competition and accidental changes in policy implementation are incorporated into the model, making it more in line with the complex situation where the actual revenue of an enterprise is affected by information.
[0038] Indirect causal relationship: Talent support information ( ) by influencing the talent pool of enterprises ( The talent reserve part in the enterprise), which in turn affects the enterprise's technological innovation capabilities ( ). Assume that the relationship between talent reserve and technological innovation capability is ,in is the talent reserve variable, is the intercept, is the influence coefficient, is the error term. ) and the influence coefficient ( ) multiplied by ( ), quantifying the impact of talent reserves on technological innovation capabilities, intercept ( ) comprehensively considers those factors that are difficult to accurately model but have a stable impact on technological innovation capabilities, while the error term ( ) incorporates random interference factors to make the model closer to the actual situation. The impact of talent support information on talent reserves is , The error term is the indirect causal relationship that can be reflected by combining these two equations. ) and the influence coefficient ( ) multiplied by ( ), quantify the impact of talent support information on talent reserves, intercept ( ) comprehensively considers factors such as the basic status of the enterprise's talent reserve, and the error term ( ) incorporates random interference factors into the model to make it more consistent with the actual situation in which talent reserves are affected by information.
[0039] Moderating effect: The enterprise scale ( ) moderates the impact of financial subsidy information ( ) on the enterprise revenue ( ), and a moderating effect model is constructed, where is the intercept, , , are the influence coefficients of the corresponding variables respectively, is the error term, reflects the moderating effect. Quantifies the direct effect of financial subsidy information on the enterprise revenue, reflects the independent contribution of the enterprise scale to the revenue, is the key manifestation of the moderating effect, The error term incorporates factors that are not considered in the model and will have a random impact on the enterprise revenue, making the model more suitable for the actual complex market environment.
[0040] The total impact of technological innovation achievements is the sum of the direct impact and the indirect impact (assuming that financial subsidy information and talent support information independently affect technological innovation achievements), and is constructed based on the principle that technological innovation achievements are affected by financial subsidy information and talent support information. On the premise of assuming that the two types of information independently affect technological innovation achievements, the direct impact and the indirect impact are integrated, that is ; The total impact of the enterprise revenue is the sum of the direct impact and the moderating effect impact (assuming that there are no other unconsidered indirect impact factors), and is constructed around the enterprise revenue being affected by financial subsidy information and the enterprise scale, that is ; Model construction and learning Data-driven modeling: Use the PC algorithm to learn the causal structure from the enterprise historical data and the real-time data of information. Assume the data set , where is the number of data samples (such as = 1000, covering the information participation and business data of 1000 enterprises in the past 5 years). The PC algorithm gradually constructs a directed acyclic graph (DAG) through conditional independence tests. For example, when testing the independence of the information variable and the enterprise result variable under the condition of the given enterprise characteristic variable , the chi-square test statistic is used, where and are the observed frequency and the expected frequency, respectively. is the number of categories. By continuously adjusting the connection relationships between variables, the final causal structure is obtained. By solving the difference between the observed frequency ( ) and the expected frequency ( ) and squaring it to reflect the degree of difference, then dividing by the expected frequency ( ) for normalization, so that the classification differences of different magnitudes can be compared, thereby reflecting the deviation degree between the observation and the expectation under this classification.
[0041] Combined with machine learning: Taking the Gradient Boosting Tree (GBT) as an example, the information variable and the enterprise feature variable are used as input features , and the enterprise result variable is used as the output label. In the st iteration, a regression tree is constructed, and the model is updated by minimizing the loss function (mean squared error loss), where is the model prediction value, and the model prediction formula is , is the number of trees (e.g., = 100). To make the model prediction value ( ) as close as possible to the true value ( ), an index for measuring the difference between the two needs to be defined, that is, the loss function. The mean squared error loss function is constructed for this purpose. It measures the overall prediction error degree of the model by calculating the square of the difference between the true value ( ) and the model prediction value ( ) of each sample and summing up the results of all samples. Calculating is the square of the difference between the true value and the predicted value of the th sample. Squaring the difference is to highlight the influence of larger errors and avoid the cancellation of positive and negative errors, more accurately reflecting the prediction deviation degree of a single sample.
[0042] Parameter estimation and optimization: When estimating the parameters of the impact of financial subsidy information on enterprise revenue, maximum likelihood estimation is adopted. Assume that the enterprise revenue follows a normal distribution , and the likelihood function is , where , and the parameter is estimated by solving the maximum value of the log-likelihood function . is the th enterprise revenue sample under the normal distribution The probability density function, where is the normalization constant of the normal distribution, ensuring that the integral of the probability density function over the entire domain is equal to 1. reflects the deviation of the sample value from the mean denotes the multiplication of the probability density functions of samples to obtain the likelihood function of the entire dataset , comprehensively reflecting the likelihood of observing the revenue samples of these enterprises under the parameter . Meanwhile, use cross-validation (such as 5-fold cross-validation) to adjust the model hyperparameters, such as the number of trees in gradient boosting trees , learning rate (such as = 0.1), etc., with the goal of minimizing the mean squared error (MSE), that is . Calculates the square of the difference between the true value and the predicted value of the th sample, aiming to highlight the impact of large errors, avoid the cancellation of positive and negative errors, and accurately reflect the prediction deviation degree of a single sample. First, accumulate the error values of samples, and then divide by the sample size to obtain the mean squared error , reflecting the average prediction error situation of the model on all samples.
[0043] The specific application process of the model is as follows: Prediction of information effect: For the candidate information and enterprise , input their features into the trained model to obtain the predicted enterprise result variable .
[0044] For example, a company plans to apply for "Support Information on the Innovation and Development of the Artificial Intelligence Industry" ( ). The enterprise has 300 employees, and the standardized enterprise scale (assuming the minimum value of the number of enterprise employees in the sample is 100 and the maximum value is 500); the industry attribute is encoded as 001; the enterprise strategy Converted into a vector [0.7, 0.2, 0.1] through the BERT model; the capital reserve is 5 million yuan, which is standardized to 0.6 (assuming the minimum capital reserve in the sample is 1 million yuan and the maximum is 9 million yuan), the number of technical patents is 20, which is standardized to 0.4 (assuming the minimum number of patents in the sample is 5 and the maximum is 45). The type of this information Encoded as [1, 0, 0], the intensity of the information Is 15% of the subsidy amount accounting for the R & D investment, the implementation time of the information Is 2024-01-01. Input these variables into the trained gradient boosting tree model (the number of trees = 100, the learning rate = 0.1), and predict that after applying for this information, the company's revenue in the next year Is expected to grow to 38 million yuan (the original revenue is 30 million yuan), and the number of patents Is expected to increase to 30 items.
[0045] This step helps the company evaluate the value of information in advance and avoid blindly applying for information by accurately predicting the company's results after the implementation of the information. In actual applications, after using this model to predict the effect of information, the success rate of the company applying for information has increased from 40% to 54%, an increase of 35%; the average cost of resource waste caused by information mismatch has decreased by 40%, and the time cost has decreased by 30%. The prediction of information effect depends on the quantitative indicators of information extracted in the information analysis stage (such as information type, intensity, etc.) and the enterprise characteristic information generated in the enterprise portrait construction stage (enterprise scale, industry attributes, etc.), providing an important reference basis for information matching. At the same time, the prediction results will be fed back into the information matching algorithm to optimize the calculation of information matching degree, making the accuracy of information matching degree calculation increase by 25% and further improving the accuracy of information matching.
[0046] Combined evaluation of information: For information combinations , analyze the causal relationship and synergy between them, and the formula is = . By linearly combining the predicted values of technological innovation achievements corresponding to different information , , according to certain weights , , the comprehensive impact of the information combination is obtained, and the causal relationship and synergy between information are analyzed based on this.
[0047] For example, a company is considering applying for a combination of "AI enterprise R & D subsidy information" ( ), and "high - end talent introduction information" ( ). The information type of the R & D subsidy information is encoded as [1, 0, 0], and the information intensity is 12% of the subsidy amount accounting for R & D investment; the information type of the high - end talent introduction information is encoded as [0, 0, 1], and the information intensity is a settlement allowance of 500,000 yuan for high - end talents. The company currently has 300 employees, and the enterprise scale = 0.5; the industry attribute is encoded as 001; the enterprise strategy vector is [0.7, 0.2, 0.1]; the standardized value of the capital reserve is 0.6, and the standardized value of the number of technology patents is 0.4. Through model calculation, the R & D subsidy information ( is expected to increase the number of the company's patents by 8 items, and the high - end talent introduction information ( is expected to increase the company's market share by 5%. According to the determined weight coefficients = 0.6, = 0.4, calculate the comprehensive impact of the information combination on the enterprise = 0.6×8 + 0.4×5 = 6.8, and quantitatively evaluate that this information combination has high value for the enterprise's development. The weight coefficients are dynamically adjusted according to differences in different industries, policy cycles, or enterprise development stages. For example, traditional manufacturing industries pay more attention to equipment renewal subsidies, and the weight of the enterprise's market share is higher; technology enterprises focus on R & D subsidies, and the weight of the enterprise's patent quantity is higher.
[0048] This step can effectively evaluate the synergy benefits of information combinations, helping enterprises select the optimal information combination and improve the utilization efficiency of information resources. Actual data shows that after adopting this evaluation method, the probability for enterprises to select an effective information combination has increased from 30% to 42%, a 40% increase; the average benefit brought by the synergy of the information combination for enterprises has increased from the original 1.5 million yuan to 1.95 million yuan, a 30% increase. The evaluation of the information combination is based on the candidate information screened during the information matching process, and is analyzed in combination with the causal relationships of information (formulas such as direct causality, indirect causality, and moderating effects) obtained from information analysis and the enterprise portrait information. The evaluation results provide more accurate information combination recommendations for information push, increasing the targeting of information push by 30%; at the same time, it can also be fed back into the information matching algorithm to optimize the matching strategy of the information combination, improving the accuracy rate of information combination matching by 20%.
[0049] Counterfactual reasoning: Assume that an enterprise has not applied for a certain piece of information , under the current enterprise characteristics , use the model to simulate the development of the enterprise .
[0050] For example, a certain company has applied for the "Support Information for the Innovative Development of the Artificial Intelligence Industry". The enterprise currently has 300 employees =0.5), industry attribute encoded as 001, enterprise strategy vector is [0.7, 0.2, 0.1], the standardized value of capital reserve is 0.6, and the standardized value of the number of technical patents is 0.4. Through counterfactual reasoning, after removing the relevant items of the information variable, the model prediction value is recalculated. The model simulates that when the enterprise does not apply for this piece of information, the revenue of the enterprise in the next year is only 33 million yuan, while the predicted revenue after actually applying for the information is 38 million yuan. Calculate the value of this piece of information =3800 - 3300 = 5 million yuan, intuitively showing the contribution of this piece of information to the revenue growth of the enterprise.
[0051] This step helps enterprises to clearly understand the actual value of information to their own development, and provides strong support for the subsequent information application decision-making of enterprises. When sending information to enterprises, these relationship analyses will be sent to enterprises together, thereby improving credibility and making enterprises more convinced of the push results. Practice shows that the rationality of enterprises' information decision-making based on counterfactual reasoning has increased from 60% to 78%, an increase of 30%; the number of decision-making errors caused by misjudgment of the value of information by enterprises has decreased by 35%. Counterfactual reasoning uses the results of information matching and information effect prediction to further analyze the value of information. The results can be fed back to the information matching and push links, helping enterprises to better understand the relationship between information and their own development, so that the fit of information matching is increased by 25%, and the satisfaction of information push is increased by 20%; at the same time, it also provides actual case basis for feedback optimization, promotes continuous improvement of the system, and optimizes the overall performance of the system by 15%.
[0052] To push information to enterprises, we first need to build enterprise portraits, collect and organize enterprise information, and preset the enterprise information items that need to be entered, including basic enterprise information (such as enterprise name ,industry , scale S, establishment time T, etc.), operating data (such as revenue R, profit P, market share M, etc.), R&D data (such as R&D investment RD, number of patents PN, number of innovative achievements IN, etc.), talent data (such as employee education structure , the number of professional and skilled personnel (SP), etc.) Collect and organize the input of various types of enterprise information and build an enterprise information database.
[0053] For example, the basic information of a company is: company name: company, industry: artificial intelligence, medium-sized (300 employees), founded in 2018; operating data: revenue of 30 million yuan, profit of 5 million yuan, and market share of 8% in the previous year; R&D data: annual R&D investment of 6 million yuan, 12 patents, and 5 innovative achievements (such as new artificial intelligence algorithms); talent data: 30% of employees with a master's degree or above, and 80 professional and skilled talents (such as artificial intelligence algorithm engineers and data scientists). This information is sorted and stored in the enterprise information database.
[0054] This step comprehensively and systematically collects enterprise information, provides a data basis for building an accurate enterprise portrait, accurately reflects the characteristics and strength of the enterprise, and provides detailed enterprise information basis for subsequent information matching. It is the basis for generating enterprise portraits. The integrity and accuracy of enterprise information directly affect the quality of enterprise portraits, and thus affect the accuracy of information matching and push. At the same time, the enterprise information database also provides enterprise-side data support for data comparison and analysis in information parsing.
[0055] Enterprise portrait generation is based on an enterprise information database and uses data analysis and machine learning techniques to generate an enterprise portrait. Enterprises are grouped according to similar characteristics through cluster analysis, and dimensionality reduction techniques such as principal component analysis (PCA) are used to extract key features and construct a multi-dimensional portrait of the enterprise. Assuming that the feature matrix of enterprise information is X, the principal component matrix is obtained after PCA transformation , and the main principal components are selected as the key features of the enterprise portrait.
[0056] For example, when analyzing numerous artificial intelligence enterprises, first construct a feature matrix X containing the information of a certain company. Use PCA technology to process matrix X to obtain the principal component matrix Y. It is found that two of the principal components mainly reflect the enterprise's R & D innovation ability (highly correlated with R & D investment, number of patents, etc.) and market competitiveness (related to revenue, market share, etc.). Based on these two principal components, construct a portrait for a certain company, highlighting its R & D strength (high R & D investment, large number of patents), market competitiveness (certain revenue and market share), and talent advantages (proportion of highly educated and professional skilled talents) in the field of artificial intelligence, etc., in order to accurately match with information.
[0057] This step simplifies complex enterprise information into a representative multi-dimensional portrait, facilitating quick and accurate matching with information, improving the efficiency and accuracy of information matching, and providing information recommendations that better meet the actual needs of enterprises. It provides a feature representation of the enterprise side for information matching and performs matching calculations with the quantitative indicators of information; at the same time, the generation of the enterprise portrait depends on the results of enterprise information collection and collation and is a further processing and refinement of enterprise information.
[0058] Information matching: Design a matching algorithm to establish an information matching algorithm, adopt a content-based matching method, and calculate the similarity between the enterprise portrait and the applicable conditions of the information. Considering the different emphases of information on different indicators, introduce a weight vector , where is the number of indicators participating in the matching calculation, represents the weight of the th indicator, and , .
[0059] Assume that the th quantitative indicator requirement of the information is , and the actual value of the enterprise in the th indicator is . The similarity calculation formula is: ; This formula measures the similarity between the enterprise indicators and the information indicators by calculating the reciprocal of the square root of the weighted sum of squares of the differences between them. The closer the value is to 1, the higher the similarity.
[0060] Construct an information complementarity matrix , the matrix dimension is ( is the total number of information), where the element represents the degree of complementarity between information and information (information is the currently selected information, information is the information that the intended recommended enterprise is still using and implementing, as well as the information to be pushed to the enterprise. The current information is matched one by one with all the information that the enterprise is using and the expected pushed information), and the value range is [0,1]. The closer the value is to 1, the stronger the complementarity. The specific calculation formula is as follows: ; Among them: , , are weight coefficients, and + + = 1, which are used to adjust the importance of the information target coordination degree ( ), the resource complementarity degree ( ), and the implementation effect mutual promotion degree ( ) in the complementarity evaluation. It can be determined by expert scoring combined with historical data analysis according to factors such as industry characteristics and information types. For example, in the scientific and technological innovation industry, the information target coordination degree weight can be set to 0.4.
[0061] : The target coordination degree score between information and information is calculated by analyzing the semantic similarity and target field overlap degree of the target descriptions in the information clauses, with a full score of , are respectively the theoretical maximum values of the target scores of information , when evaluated separately.
[0062] : The resource complementarity degree score between information and information evaluates the complementarity of the resources required for the information (such as funds, talents, technology, etc.), with a full score of , They are information messages respectively 、 The theoretical maximum value of the resource score when evaluated separately
[0063] : Information message And the information message The mutual promotion degree score of the implementation effect, based on the analysis of the mutual promotion effect of the implementation effect of the information message from historical implementation cases, with a full score 、 They are information messages respectively 、 The theoretical maximum value of the effect score when evaluated separately
[0064] In the target synergy degree dimension Is to normalize the target synergy degree index value, map it to the interval [0, 1], and then multiply by the weight coefficient , to obtain the contribution value of the target synergy degree in the complementary degree between the information message And the information message , which is also obtained by the same method in the resource complementarity degree dimension , and obtained by the same method in the implementation effect complementarity degree dimension , and the complementary degree of the two information messages is obtained by adding the three together
[0065] Enterprise planning matching factor Introduce the enterprise planning matching coefficient And , which respectively represent the matching degree between the information message And the enterprise Short-term and long-term plans, with a value range of [0, 1]. It is calculated by analyzing the short-term (1 - 3 years) and long-term (more than 3 years) development goals and strategic directions filled in by the enterprise in the system, and conducting semantic matching and target consistency evaluation with the information message goals. For example, if the short-term plan of the enterprise is to expand the market share of a certain product, and if the goal of the information message Includes support for the market promotion of this product, then The score is higher
[0066] When calculating the matching degree of the information message, comprehensively consider the matching degree between a single information message and the enterprise, the complementarity between the information message and other information messages, and the matching with the enterprise planning. The final matching degree calculation formula is adjusted to ; Among them is the collaborative filtering adjustment coefficient (value range [0.8, 1.2]), which is dynamically adjusted using reinforcement learning and adaptive algorithms according to enterprise feedback, new data, and information changes to ensure the accuracy and effectiveness of information matching and pushing. is the weight coefficient for short-term and long-term planning (value range [0, 1]), which can be set by the enterprise according to its own development stage. For example, an enterprise in the initial stage of entrepreneurship can set to 0.7, focusing on short-term planning matching. is the variable for traversing information. is the total number of information. This formula ensures that the system preferentially recommends information combinations that have a high degree of matching with the enterprise, strong complementarity with other information, and are in line with the enterprise's planning. is the information obtained through the similarity calculation formula and the enterprise The similarity in indicators is multiplied by the weight coefficient to reflect the contribution of the matching degree of a single piece of information itself to the enterprise to the final matching degree. reflects the information The sum of the complementarity degrees with all other information reflects the influence of the complementary characteristics of the information in the overall information combination on the matching degree. reflects the influence of the matching degree between the information and the enterprise's planning on the final matching degree. The three are multiplied to comprehensively reflect the influence of multiple factors on the matching degree between the information and the enterprise.
[0067] For example, a certain company (enterprise ) participates in the matching of "Support Information for the Innovation and Development of the Artificial Intelligence Industry" (information ). This information involves 3 key indicators of the R & D investment ratio, the number of patents, and the proportion of high-tech product income ( = 3), and the corresponding quantitative index requirements = 20%, = 15 items, = 60%. The weight vector (0.4, 0.3, 0.3) is set. The actual R & D investment ratio of the enterprise = 25%, the number of patents = 12 items, and the proportion of high-tech product income = 70%.
[0068] First, calculate the similarity : ; Through the collaborative filtering algorithm, the system finds that most of the enterprises similar to a certain company have successfully applied for this information, and determines the adjustment coefficient = 1.2, then the final matching degree =1.2 0.61=0.732.
[0069] This step is based on weighted calculation and collaborative optimization to comprehensively and accurately evaluate the matching degree between enterprises and information, reduce the cost of enterprises to screen information, and improve the efficiency of information utilization. It relies on the quantitative indicators extracted from information analysis and the characteristics of enterprise portraits to provide matching results for information push.
[0070] Sorting and filtering matching results According to the final matching Sort the information matching results in descending order and display the top ranking information. Set the matching threshold (Generally, the value is 0.5-0.7), filtering out information below the threshold. At the same time, it provides enterprises with sorting options based on different dimensions such as the strength of information complementarity, comprehensive benefits, and matching degree with enterprise planning, so that enterprises can intuitively obtain the most valuable information combination. For example, if an enterprise is in a strategic transformation period, it can choose to sort by matching degree with long-term planning, and give priority to viewing information combinations that help the enterprise achieve long-term transformation goals.
[0071] For example, after a company completes the information matching, the system will display the information by default in descending order of matching degree, and set the threshold =0.6, filtering out information with a matching degree lower than 0.6. In addition, companies can also choose to re-sort the remaining information by the time of information release, from recent to distant, or by the size of the discount.
[0072] This step reduces the interference of invalid information, meets the diversified query needs of enterprises, and improves the efficiency and pertinence of enterprises in obtaining information. The matching results of information are processed to provide an orderly and accurate list of information for the information push link.
[0073] Information push: Multi-channel push pushes matched information to enterprises through multiple channels such as email, SMS, system station messages, etc. Enterprises can set their preferred push method and time in the system. For urgent information, the system will give priority to timely notifying enterprises through SMS.
[0074] For example, a company sets up email as the main receiving channel in the system, and the receiving time is 10 am every Friday. Every Friday, the system will send information with a high degree of matching, such as "special subsidy information for artificial intelligence technology research and development" and "information on support for the construction of intelligent industry innovation platform", to the company's reserved mailbox in the form of email. When urgent information such as "urgent extension notice for high-tech enterprise application" appears, the system will immediately send key information to the relevant person in charge of the company via SMS.
[0075] This step meets the diverse information reception needs of enterprises, ensures that information reaches enterprises in a timely manner, improves the utilization rate of information, and enhances the responsiveness of enterprises to information. It is the output link of the information matching result, directly providing services to enterprises, and the push effect will affect the evaluation and feedback of enterprises on the entire system.
[0076] Personalized push: Based on the enterprise's browsing history, interest selections, and real-time demand changes, personalized information push is realized. Suppose the enterprise The number of times of browsing a certain type of information is , and the total number of browsing times is ; the browsing duration is , and the total browsing duration is ; the number of clicks is , and the total number of clicks is . Then the attention weight of this type of information is calculated as follows: ; Among them, , , are weight coefficients, and their value ranges are all [0, 1], and . The system preferentially pushes information of high-weight categories according to the attention weight. Considering that the degree of enterprise attention to information can be reflected from multiple dimensions such as the number of browsing times, browsing duration, and number of clicks, the attention weight is calculated by weighted combination of the indicators of these dimensions. The weight coefficients , , are used to adjust the importance of each dimension in the attention evaluation to adapt to the different emphases of different enterprises on information attention.
[0077] When calculating the attention weight , the feedback of the enterprise on the information combination and the matching situation between the information combination and the enterprise plan are taken into consideration. If the enterprise gives positive feedback on a certain type of information combination and this combination highly matches the enterprise's current plan, then the weight of this type of information combination in subsequent pushes is increased. At the same time, combined with the current development stage and strategic goals of the enterprise, the information combination push strategy is dynamically adjusted. For example, for a certain company in a stable development period and focusing on long-term innovation, the system preferentially pushes information combinations that are complementary to and mutually promoting the enterprise's long-term plan, such as "long-term R & D fund support information" and "high-end talent cultivation and introduction information".
[0078] For example, in the past month, a certain company (enterprise ) the number of times of browsing science and technology innovation-related information = 20 times, total number of views = 30 times; viewing duration = 100 minutes, total viewing duration = 150 minutes; number of times clicking on the link of science and technology innovation-related news information = 5 times, total number of clicks = 8 times. Set = 0.4, = 0.4, = 0.2, then the attention weight of science and technology innovation-related news information is: ; The system, based on this weight, increases the push intensity of science and technology innovation-related news information, and preferentially pushes "News information on funding support for the R & D of cutting-edge artificial intelligence technologies", "News information on innovation achievement transformation rewards", etc.
[0079] This step precisely fits the personalized needs of enterprises, improves the attention and participation enthusiasm of enterprises for the pushed news information, enhances the effectiveness of news information push, and increases the satisfaction of enterprises with the system. Based on the enterprise portrait and the news information matching results for further optimization, the collected enterprise feedback data can be used to improve the enterprise portrait and optimize the news information matching algorithm, forming a closed-loop optimization mechanism.
[0080] Visualization and interaction: Design uses visualization tools such as Tableau or PowerBI to present the news information push results to enterprise users in an intuitive way. In addition to displaying the key information and matching degree of a single piece of news information, it highlights the complementary relationship, synergistic benefits of the news information combination, and the matching situation with the enterprise plan. Show the complementary association between each piece of news information in the news information combination through a relationship network diagram, present the expected changes in enterprise benefits before and after the implementation of the news information combination using a comparison bar chart, and at the same time show the matching degree of the news information combination with the short-term and long-term plans of the enterprise using a radar chart. At the same time, provide an interactive interface where users can adjust parameters. Enterprises can independently select different news information combinations and view in real-time information such as complementary analysis, benefit prediction, and matching degree analysis with the plan of the combination, facilitating enterprises to deeply understand the value of the news information combination and make scientific decisions.
[0081] For example, when a company views the push results of "Tax Incentive Information for the Artificial Intelligence Industry" in the system, through a bar chart, key information such as the tax exemption amount and application conditions of the information, as well as the matching score between the enterprise and the information, can be clearly seen. The heat map intuitively shows that characteristics such as the number of artificial intelligence patents and the scale of the R & D team of the enterprise contribute more to the matching of this information. By adjusting the application threshold requirement for "the number of patents" in the information through a slider, the system will update the matching score and the heat map in real time, showing the changes in the contributions of various characteristics, and helping the enterprise intuitively understand the matching effects of the information under different conditions.
[0082] This step presents the push results of the information in an intuitive and easy-to-understand way, helps enterprise users deeply understand the basis for information matching, and enhances the trust of users in the system's recommended results; the interactive operation enables the enterprise to actively explore the matching situations under different information conditions, provides a reference for the enterprise to formulate development strategies, and improves the user experience. It is the display and interaction link of information push, visualizing the results of information matching and personalized push to users; the feedback data during the interaction process can be used to optimize the information matching algorithm and personalized push strategy, promoting the continuous improvement of the system.
[0083] User feedback collection: Establish a perfect user feedback mechanism in the system, set up a special feedback entrance, and in addition to collecting the enterprise's evaluation of a single piece of information, focus on collecting the enterprise's feedback on information combinations, including whether the complementarity of the information combination is reasonable, whether the synergy effect meets expectations, and whether the matching degree with the enterprise's plan meets expectations, etc. At the same time, conduct user research regularly, through forms such as questionnaires and interviews, to deeply understand the problems and improvement suggestions encountered by enterprises in the actual application of information combinations.
[0084] For example, after each information push, a company can evaluate the information through the feedback form in the system, and can select options such as "very relevant", "relatively relevant", "not relevant", etc., and can also fill in specific opinions. Every month, the system will also send a questionnaire to the enterprise, asking questions such as the overall satisfaction of the enterprise with the recently pushed information and the clarity of information interpretation. For example, a company feedback that although the matching degree of a certain "Tax Incentive Information for Artificial Intelligence Product Exports" is relatively high, the actual application process is too complicated, and it hopes that the system can provide more detailed application guidance and precautions.
[0085] This step directly obtains the enterprise's evaluation and requirements for the system, providing a true and reliable basis for system optimization. It helps to promptly identify problems in the process of information matching and pushing, and make targeted improvements to enhance the service quality of the system. The feedback information collected will be used to optimize various links such as information parsing, information matching, and information pushing, forming a cycle of continuous system improvement.
[0086] Model Optimization and Strategy Adjustment: According to the user feedback information, optimize and adjust the information matching model and pushing strategy. If an enterprise feedbacks that the complementarity of a certain information combination does not meet the expectation, re-examine the information complementarity matrix and adjust the complementarity coefficient of relevant information; if users are not satisfied with the matching degree between the information combination and the enterprise plan, optimize the calculation model for the matching of enterprise plan and information; if users put forward new requirements for the pushing timing and method of the information combination, promptly optimize the pushing strategy. For example, if an enterprise reflects that the matching degree of a certain information combination with the long-term plan is low, the system re-analyzes the correlation between the enterprise's long-term plan goal and the information goal and adjusts the calculation method; if an enterprise reflects that a certain information combination is pushed too late and misses the best application opportunity, the system optimizes the real-time stream processing mechanism, strengthens the monitoring of information release and enterprise demand changes to ensure the timely pushing of information combinations. Through continuous feedback optimization, continuously improve the accuracy and effectiveness of information combination recommendation.
[0087] For example, if multiple enterprises feedback that the matching of "AI enterprise talent subsidy information" is inaccurate, after analysis, it is found that the weight setting of the "professional qualification of R & D personnel" index is too high, resulting in deviation of the matching results of some enterprises. The system adjusts the weight of this index from =0.5 to =0.3, and recalculates the weights of other indexes to optimize the matching model. Also, if most enterprises express the hope to receive information pushing at the beginning of each month, the system then uniformly adjusts the pushing time to the 1st - 3rd day of each month.
[0088] This step enables the system to quickly adapt to the changes in enterprise requirements, continuously improve the accuracy and effectiveness of information matching and pushing, maintain the competitiveness and practicality of the system, and better meet the enterprise's need to obtain information. It is the optimization and improvement of the entire information pushing process, adjusting the models and strategies involved in information parsing, enterprise portrait construction, information matching, etc. based on user feedback to enhance the overall performance of the system; the optimized results will in turn affect subsequent user feedback and system operation effects, promoting the continuous evolution of the system.
[0089] Cloud Computing and Edge Computing Collaboration: Adopt an architecture mode that synergizes cloud computing and edge computing to give full play to the advantages of both. Cloud computing, with its powerful computing and storage capabilities, undertakes large-scale data processing and complex computing tasks, such as deep learning analysis of information text and model training. Deploy lightweight models on edge devices (such as enterprise local servers and intelligent terminals) to achieve fast local data processing and preliminary information matching, reduce data transmission latency, and improve system response speed. Use the federated learning fine-tuning mechanism to regularly synchronize the latest models from the cloud to edge devices to ensure that the models on edge devices are consistent with the cloud, so as to adapt to the dynamic changes of information and enterprise data.
[0090] For example, in the information text analysis task, the cloud computing platform uses its powerful computing power resources to train and optimize the BERT model for a large amount of information text, continuously improving the model's ability to understand the semantics of information. A company deploys a lightweight BERT model obtained through knowledge distillation on its enterprise local server to perform preliminary keyword extraction and information category classification on data such as technical documents and project reports submitted daily by the enterprise, quickly determining whether it may be related to certain information. Every week, the cloud synchronizes the updated model parameters to the edge device through the federated learning fine-tuning mechanism to ensure that the edge device model can timely adapt to the changes in new information types and information content. For example, when new information related to artificial intelligence ethics norms is released, after the cloud model is updated, the edge device can quickly obtain the new model parameters and make more accurate judgments on relevant texts.
[0091] This step realizes the reasonable allocation and efficient utilization of computing resources, improves the efficiency and response speed of the system in processing data, reduces data transmission costs and latency; the federated learning fine-tuning mechanism ensures the timeliness and accuracy of the edge device model, enabling the system to operate stably under different network environments and device conditions, providing enterprises with a smooth and efficient service experience. It provides powerful computing and storage support for steps such as information parsing, enterprise portrait construction, information matching and pushing, ensuring the efficient progress of data processing and model training; the local preliminary processing and fast response achieved by edge computing optimize the user experience, are closely related to the information pushing link, and improve the timeliness of information pushing; at the same time, it also provides a technical foundation for feedback optimization, facilitating the collection and processing of user feedback data and realizing the continuous optimization of the system.
[0092] Real-time stream processing technology: Using real-time stream processing technologies such as FlinkML, dynamically monitor and process the data of enterprise and information. When new information is released or enterprise information is updated, trigger the data processing process in real time, and update the intelligent database of information and the enterprise information database in a timely manner to ensure that the data on which information matching and pushing are based is always the latest. Set trigger thresholds (such as the enterprise revenue fluctuation exceeding 10%, the change in the number of new patent applications, etc.). When the threshold is reached, automatically trigger the retraining of the model, enabling the system to quickly adapt to the changes in enterprise data and improve the accuracy of information matching.
[0093] For example, the revenue data, patent application data, etc. of a certain company are transmitted to the system in real time. When the enterprise revenue in a certain quarter increased by 15% compared with the previous quarter (exceeding the set 10% threshold), the system immediately automatically starts the data processing process, updates the revenue data in the enterprise information database, and triggers the retraining of the information matching model to re-evaluate the matching degree between the enterprise and the existing information. At the same time, if a new "laddered reward information for the revenue growth of artificial intelligence enterprises" is released at this time, the system can capture the information in real time, update the intelligent database of information, and based on the latest enterprise data and information, re-match and push the information, and recommend the information to the company in a timely manner.
[0094] This step ensures that the system can keep up with the dynamic changes of information and enterprise data in a timely manner, provide the latest and most accurate information matching and pushing services for enterprises, avoid enterprises missing information opportunities due to data lag, greatly improve the practicality and timeliness of the system, and enhance the dependence of enterprises on the system. Updating information and enterprise data in real time provides the latest data support for information parsing, information matching and pushing, which is the key technical support for realizing accurate and timely pushing; the model retraining mechanism ensures the accuracy of the information matching model, which is closely connected with the information matching and pushing links; at the same time, it also provides a real-time data basis for feedback optimization, facilitating the timely collection and analysis of user feedback and promoting the continuous improvement of the system.
[0095] Model compression and acceleration technology: Adopt lightweight meta-learning and model compression technologies to improve system performance. Use Meta-LSTM as a lightweight meta-learning model to effectively reduce the number of model parameters. Combine knowledge distillation technology to transfer the knowledge of complex models (such as XGBoost+CNN) to small models (such as MobileNet). Deploy the quantized model (such as INT8 precision) on edge devices, and only upload complex tasks to the cloud for processing. Use TensorRT or ONNXRuntime to replace the traditional deep learning inference framework to improve the computing efficiency of edge devices and reduce the system operation cost.
[0096] For example, during the training of the information matching model, first, a complex model of XGBoost + CNN is trained to deeply extract and match the features of the information. Then, the knowledge distillation technique is used to transfer the knowledge learned by the complex model to the MobileNet model. The number of parameters of the original complex model is 10 million, and the number of parameters of the transferred MobileNet model is significantly reduced to 2 million. On the edge device of a certain company, the MobileNet model quantized to INT8 precision is deployed and accelerated using the TensorRT inference framework. Originally, it took 5 seconds to perform an information matching inference using the traditional deep learning inference framework. After adopting the optimized solution, the inference time is shortened to 1 second, greatly improving the processing efficiency. For some complex information matching tasks involving multi-dimensional data comprehensive analysis and complex causal relationship judgment, they are uploaded to the cloud for processing.
[0097] This step significantly reduces the computational burden and energy consumption of the edge device, greatly improves the model inference speed, and effectively reduces the system operation cost; on the premise of ensuring the model accuracy, it realizes the lightweight and acceleration of the model, enables the system to operate efficiently on resource-constrained devices, improves the scalability and adaptability of the system, and ensures the efficiency of information matching and pushing. It provides efficient model support for information parsing, information matching, and pushing, especially in the edge computing scenario, improving the efficiency of local data processing and preliminary information matching. Cooperating with the cloud computing and edge computing collaboration technology, it optimizes the overall performance of the system. The model compression and acceleration technology can reduce the resource occupancy of the model on the edge device, enabling the edge device to run the lightweight model more smoothly. Combined with the real-time stream processing technology, it ensures the efficiency of the system in processing dynamic data and guarantees the timeliness and accuracy of information pushing. At the same time, it also provides a more efficient model basis for feedback optimization, facilitating the rapid processing of user feedback data and promoting the continuous optimization of the system.
[0098] Embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. Embodiments disclosed by the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the methods of the present application are executed. It should be noted that the above-mentioned computer-readable medium in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate or transmit a program for use by or in combination with an instruction execution system, apparatus or device. The program codes contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless segment, wire segment, optical cable, RF, etc., or any suitable combination of the above.
[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Without departing from the said principles, the embodiments of the present invention may have any variations or modifications.
Claims
1. A method for matching and pushing information based on big data drive, characterized in that, The method includes the following steps: Using web crawler technology to capture information text data, obtaining enterprise internal data through API interfaces, integrating multi-channel information text data and enterprise operation data, and performing data preprocessing; Using an automatic update tool and NLP to build an entity recognition model, setting a threshold, filtering low-quality data in the preprocessed data, and detecting and correcting contradictory data; Extracting key information of information through natural language processing technology, and automatically classifying information using text classification algorithms; Through data analysis and machine learning technology, constructing a multi-dimensional enterprise portrait based on enterprise content data, determining quantitative indicators according to the type of information, constructing an information impact model, analyzing the impact of information on the enterprise, and constructing an information complementarity matrix between different information; Constructing a single information matching model according to the matching algorithm, calculating the matching degree between a single piece of information and the enterprise, and constructing an information matching model according to the matching degree between the single piece of information and the enterprise, the information complementarity matrix, and the matching between the information and the enterprise plan, and calculating the final information matching degree; Pushing the information with the final information matching degree higher than the matching degree threshold and the analysis results of the information impact model to the enterprise, collecting user feedback, and optimizing the information matching model and the pushing strategy according to the user feedback.
2. The method for matching and pushing information based on big data drive according to claim 1, characterized in that The data preprocessing includes the following steps: For websites with frequent updates, crawl once a day or once an hour. For industry association websites with slow updates, crawl once a week or once a month. Adopt a random delay mechanism to crawl information text data from government official websites, government affairs platforms, and industry association websites; obtain relevant data in the enterprise internal system through API interfaces, and integrate the basic information, operation data, and industry data of the enterprise; clean the collected data to remove duplicate and invalid data; fill in missing values using mean filling and regression prediction methods; standardize the data; unify the data structure and format, and establish the association relationship between data.
3. A method for pushing information matching based on big data drive according to claim 2, characterized in that, The specific process of filtering low-quality data is as follows: Using the BERT model to process information text data, determining key entities in the information text data and relevant entities in the enterprise internal data through entity recognition, using relation extraction technology to establish the connection between the two types of entities, and constructing a dynamic knowledge graph; Setting a data quality threshold to automatically filter low-quality data; using the inference engine of the knowledge graph to detect contradictory data. When contradictory data is detected, trigger manual review or automatic correction to correct the contradictory data.
4. A method for matching and pushing information based on big data drive according to claim 3, characterized in that The specific process of automatically classifying information is as follows: Using the keyword extraction algorithm to extract keywords, titles, and key sentences as text features, and extracting the release time, release media, and associated enterprise attributes as structured features; converting the text features and structured features into numerical vectors through the word vector model; collecting training data and annotating it through the supervised learning model, and training the model to learn the classification of features and categories; evaluating the performance of the model through the test set, adjusting parameters, optimizing features and the model according to the results; inputting the numerical vectors into the model for classification, and storing the results by category.
5. A method for pushing information matching based on big data drive according to claim 4, characterized in that, The information impact model includes the following formula: Total impact of technological innovation achievements It is the sum of the direct impact and the indirect impact, and the specific formula is as follows: ; Total impact on corporate revenue It is the sum of the direct impact and the moderating effect, and the specific formula is as follows: ; Wherein: is the intercept, , , are the influence coefficients of the corresponding variables respectively, is the error term, is the original revenue base value of the enterprise, is the information intensity, is the enterprise scale, is the intercept term, is the information intensity is the influence coefficient on the technological innovation achievement, is the error term, is the random error term, is the intercept, is the influence coefficient, is the error term, is the error term, is the intercept, Influence coefficient.
6. A method for matching and pushing information based on big data drive according to claim 5, characterized in that The calculation formula for the final information matching degree is as follows: ; Wherein: is the final information matching degree; is the collaborative filtering adjustment coefficient; is the matching degree between a single piece of information and the enterprise, is a variable for traversing information, taking values in sequence from 1 to in turn, is the total number of information, is the information and the information complementary degree, is the short-term and long-term planning weight coefficient, is the enterprise short-term planning matching coefficient, Enterprise long-term planning matching coefficient.
7. A method for matching and pushing information based on big data drive according to claim 6, characterized in that, The specific process of the information matching and pushing method driven by big data is as follows: Sort the matching results of information in descending order according to the final matching degree; set a matching degree threshold , and filter out the information whose matching degree is lower than the threshold; provide the enterprise with sorting options in different dimensions such as the complementary strength of information, comprehensive benefits, and the matching degree with the enterprise's plan, and push the matched information to the enterprise via email, SMS, and in-system messages. The enterprise can set the preferred push method and time in the system; dynamically adjust the combined push strategy of information according to the enterprise's browsing history, interest checkboxes, and real-time demand changes, combined with the enterprise's current development stage and strategic goals.
8. A method for matching and pushing information based on big data drive according to claim 7, characterized in that The optimization process of the information matching model and the pushing strategy is as follows: Set up a feedback entry to collect the feedback from enterprises on single information and information combinations, including whether the complementarity of the information combination is reasonable, whether the synergy effect meets the expectation, and whether the matching degree with the enterprise plan meets the expectation; Regularly conduct user research to understand the problems and improvement suggestions encountered by enterprises in the actual application of information combinations; If an enterprise feedbacks that the complementarity of a certain information combination fails to meet the expectation, re-examine the information complementarity matrix and adjust the complementarity coefficient of relevant information; if the user is not satisfied with the matching degree between the information combination and the enterprise plan, optimize the calculation model for the matching of the enterprise plan and the information; if the user puts forward new requirements for the pushing timing and method of the information combination, optimize the pushing strategy in a timely manner.
9. An information matching and pushing system driven by big data, characterized in that, The system is used to execute an information matching and pushing method driven by big data according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement an information matching and pushing method driven by big data according to any one of claims 1-8 above.
Citation Information
Patent Citations
Investment and research information sharing and distributing method for group business management and control
CN119719444A
Patent information push management system and method
CN120067455A
Industrial space suitability optimization method and system
CN120087806A
Method of and system for analyzing, modeling and valuing elements of a business enterprise
US20050119922A1
Cited By
Stock information pushing method
CN122222731A