Enterprise information exposure enhancement method based on multi-source data fusion

By using a multi-source data fusion method to enhance enterprise information exposure, the problems of information silos and event importance differentiation in enterprise information processing have been solved. This method enables intelligent and precise control of information exposure, ensuring logical consistency and timely presentation of information.

CN120849970BActive Publication Date: 2025-12-16NETCONCEPTS NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511375329.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-12-16
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing technologies, when processing scattered and diverse enterprise information, cannot effectively identify fields with different expressions but the same meaning from different sources, leading to the creation of information silos. Furthermore, they cannot distinguish the importance and timeliness of different events, affecting information users' timely judgment of the true situation and potential risks of the enterprise.

Method used

By collecting multi-source data and performing semantic feature vectorization processing, analyzing the degree of semantic matching, reconstructing the semantic matching path of enterprise information, and achieving intelligent and precise control of information exposure through dynamic adjustment of event density and exposure interval time.

Benefits of technology

It accurately identifies and corrects fields with inconsistent expressions but the same meaning from different sources, ensuring logical consistency and integrity in the information integration process. High-value and time-sensitive information is presented in a timely manner, achieving intelligent and precise control of information exposure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849970B_ABST
    Figure CN120849970B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, in particular to an enterprise information exposure enhancement method based on multi-source data fusion, multi-source data is collected and vectorized, the information matching path is optimized by analyzing the semantics, then the density of enterprise related events is calculated, according to the comparison result of the density and the preset threshold, the interval and frequency of information exposure are dynamically adjusted. In the present application, by converting different types of data such as enterprise basic information, financial data, news reports and social media dynamics into computable semantic vectors, the fields with different expressions but the same meaning between different sources can be accurately identified and corrected, then the matching path between information is reconstructed by using state transition rules combined with the context sequence and relevance of information, ensuring the logical consistency and integrity of enterprise information in the integration process, by counting the number of events related to the enterprise within a certain time period and assigning weights, the event density index measuring the dynamic activity level of the enterprise is formed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and in particular to an enterprise information exposure enhancement method based on multi-source data fusion. BACKGROUND

[0002] The technical field of data processing covers the associated technologies of collecting, storing, cleaning, analyzing, visualizing and decision supporting various types of data by using computers and algorithms.

[0003] Among them, the enterprise information exposure enhancement method refers to the case that the enterprise related information sources are scattered and the content forms are various, the enterprise data is obtained from the business registration information, judicial announcement, news report, bidding record, industry association publicity and other information channels, and the enterprise information of different sources is combined and associated by using field matching, keyword search, time sequence sorting, entity name comparison and other methods, so that a relatively complete enterprise information exposure content is formed.

[0004] The prior art mainly relies on field matching, keyword search and other methods when processing scattered and various enterprise information. This processing mode has obvious limitations when facing the consistency problem at the semantic level. Because of the lack of understanding of the deep semantics behind the information, when different sources of data use different terms to describe the same event or entity attribute, for example, a high-level change is reported as core team adjustment in news, and is recorded as board member change in announcement, the system cannot effectively associate due to keyword mismatch, resulting in the generation of information island. In addition, the information integration method is biased towards static stacking. The information set formed by time sequence sorting and entity name comparison does not distinguish the importance and timeliness of different events. All information is treated equally. This may result in a consequence that a major legal lawsuit of an enterprise and a number of insignificant address changes in a short period of time may not highlight the urgency and importance of the latter in the presented information exposure content. The core value of information is diluted by a large amount of low-value data, thereby affecting the timely judgment of the information users on the real situation and potential risks of the enterprise. SUMMARY

[0005] The purpose of the present application is to solve the shortcomings in the prior art, and to provide an enterprise information exposure enhancement method based on multi-source data fusion.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme: the enterprise information exposure enhancement method based on multi-source data fusion comprises the following steps:

[0007] S1: collecting enterprise structured data and unstructured data in multiple business sources, and performing semantic feature vectorization processing to generate semantic vector information of enterprise information fields;

[0008] S2: According to the semantic vector information of the enterprise information field, the semantic matching degree of the same type of enterprise information field in each business source is analyzed, the enterprise information field needing context correction is marked according to the semantic matching degree, and the cross-source enterprise information semantic similarity analysis result is generated;

[0009] S3: Extracting the enterprise information field needing context correction in the cross-source enterprise information semantic similarity analysis result, reconstructing the semantic matching path of the enterprise information, and outputting the enterprise information context optimization matching result;

[0010] S4: Statistics of the number of events associated with the enterprise within a specified time, and association with the enterprise identifier in the enterprise information context optimization matching result, calculation of the enterprise information event density in the target time;

[0011] S5: Comparing the enterprise information event density with the preset frequency control threshold, adjusting the enterprise information exposure interval time parameter corresponding to the enterprise information context optimization matching result, and obtaining the enterprise information exposure enhancement result.

[0012] As a further scheme of the present application, the semantic vector information includes enterprise basic information vector, financial data vector, news report vector, social media dynamic vector, the cross-source enterprise information semantic similarity analysis result includes a list of fields to be corrected, semantic matching degree, field correction mark, the enterprise information context optimization matching result includes optimized field relationship, semantic consistency mark, matching path adjustment record, the enterprise information event density includes event quantity, event distribution within a specified time, and the enterprise information exposure enhancement result includes exposure frequency, exposure interval time.

[0013] As a further scheme of the present application, the semantic vector information of the enterprise information field is obtained by the following steps:

[0014] S111: Collecting and associating enterprise structured data and unstructured data from multiple business sources, wherein the structured data includes enterprise basic information and financial data, the unstructured data includes news reports and social media dynamics, and obtaining a cross-source enterprise information context data set;

[0015] S112: Obtain enterprise industry classification parameters and enterprise information field label parameters, group the enterprise information fields according to the enterprise industry classification parameters, calculate the co-occurrence frequency between the enterprise information fields, and construct an enterprise information context relationship graph according to the co-occurrence frequency;

[0016] S113: According to the enterprise information context relationship graph, input the cross-source enterprise information context data set into the Word2Vec model for semantic feature vectorization processing, and obtain the semantic vector information of the enterprise information field.

[0017] As a further scheme of the present application, the step of obtaining the cross-source enterprise information semantic similarity analysis result is specifically:

[0018] S211: According to the semantic vector information of the enterprise information field, the semantic matching degree of the same type of enterprise information field in each business source is calculated by using the cosine similarity algorithm, and the cross-source enterprise information semantic matching degree data is obtained;

[0019] S212: According to the cross-source enterprise information semantic matching degree data, the enterprise information field with a semantic matching degree lower than a semantic matching degree threshold is identified and marked as an enterprise information field that needs to be context corrected, and context correction marking data is obtained;

[0020] S213: The context correction marking data and the cross-source enterprise information semantic matching degree data are integrated to generate a cross-source enterprise information semantic similarity analysis result.

[0021] As a further scheme of the present application, the step of obtaining the enterprise information context optimization matching result is specifically:

[0022] S311: Obtain the enterprise information field that needs to be context corrected in the cross-source enterprise information semantic similarity analysis result, and call the original data of the enterprise information field that needs to be context corrected from the cross-source enterprise information context data set and establish a corresponding relationship to obtain a context correction field sequence;

[0023] S312: According to the context correction field sequence, the context content order of the enterprise information field in each source and the matching condition of the associated field are combined, the transition probability between the enterprise information fields is calculated according to the state transition rule of the conditional random field model, the arrangement order of the enterprise information field in the semantic matching path is adjusted, and an adjusted matching path is obtained.

[0024] S313: Call the associated field content in the cross-source enterprise information context data set and the adjusted matching path, reconstruct the semantic matching path of the enterprise information according to the optimal path output rule of the conditional random field model, and extract the corresponding enterprise identifier, to generate an enterprise information context optimization matching result.

[0025] As a further scheme of the present application, the step of obtaining the enterprise information event density is specifically:

[0026] S411: Obtain time-labeled event data associated with the enterprise, and associate it with the enterprise identifier in the enterprise information context optimization matching result to obtain an associated event data set;

[0027] S412: Based on the association event data set, the number of events associated with each enterprise is counted within a set time range, and combined with the weighted value of each event to obtain event quantity weighted data;

[0028] S413: According to the event quantity weighted data, the enterprise information event density in the corresponding time range is calculated.

[0029] As a further scheme of the present application, the acquisition step of the enterprise information exposure enhancement result is specifically:

[0030] S511: According to the enterprise information event density, compare with the preset frequency control threshold value to determine whether it is necessary to adjust the enterprise information exposure interval time parameter, and obtain exposure adjustment demand data;

[0031] S512: Based on the exposure adjustment demand data, analyze the event density change amplitude, and when the change exceeds the set smoothing control threshold value, execute step-by-step adjustment, compress or lengthen the exposure frequency by adjusting the exposure interval time, and obtain exposure interval adjustment data;

[0032] S513: Call the exposure interval adjustment data and the exposure parameter in the enterprise information context optimization matching result, so that the adjusted exposure interval will be automatically adjusted according to the event density value within the time window, and generate the enterprise information exposure enhancement result.

[0033] Compared with the prior art, the present application has the advantages and positive effects that:

[0034] In the present application, by converting different types of data such as enterprise basic information, financial data, news reports and social media dynamics into computable semantic vectors, the fields with different expressions but the same meaning between different sources can be accurately identified and corrected, and then combined with the context sequence and relevance of the information, the matching path between the information is reconstructed by using state transition rules, ensuring the logical consistency and integrity of the enterprise information in the integration process. On this basis, by counting the number of events associated with the enterprise within a certain time period and assigning weights, an event density index is formed to measure the dynamic activity level of the enterprise. According to the comparison between the event density change and the preset threshold value, the interval and frequency of information exposure are dynamically adjusted, so that high-value and high-timeliness information can be presented more timely, while the update of regular information maintains a reasonable pace, realizing intelligent and accurate control of information exposure. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 is a schematic diagram of the main steps of the present application;

[0036] Figure 2 is a flowchart of step S1 of the present application;

[0037] Figure 3 Flow chart for step S2 of the present application;

[0038] Figure 4 Flow chart for step S3 of the present application;

[0039] Figure 5 Flow chart for step S4 of the present application;

[0040] Figure 6 Flow chart for step S5 of the present application. DETAILED DESCRIPTION

[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0042] Please refer to Figure 1 The present application provides a technical solution: an enterprise information exposure enhancement method based on multi-source data fusion, comprising the following steps:

[0043] S1: Collecting structured data and unstructured data of enterprises in multiple business sources, and performing semantic feature vectorization processing to generate semantic vector information of enterprise information fields;

[0044] S2: According to the semantic vector information of the enterprise information fields, analyzing the semantic matching degree of the same type of enterprise information fields in each business source, referring to the semantic matching degree to mark the enterprise information fields that need to be context corrected, and generating cross-source enterprise information semantic similarity analysis results;

[0045] S3: Extracting the enterprise information fields that need to be context corrected in the cross-source enterprise information semantic similarity analysis results, reconstructing the semantic matching path of the enterprise information, and outputting the enterprise information context optimization matching result;

[0046] S4: Counting the number of events associated with the enterprise within a specified time, and associating with the enterprise identifier in the enterprise information context optimization matching result, calculating the enterprise information event density within the target time;

[0047] S5: Comparing the enterprise information event density with the preset frequency control threshold, adjusting the enterprise information exposure interval time parameter corresponding to the enterprise information context optimization matching result, and obtaining the enterprise information exposure enhancement result;

[0048] The semantic vector information includes a basic information vector of an enterprise, a financial data vector, a news report vector, and a social media dynamic vector. The cross-source enterprise information semantic similarity analysis result includes a list of fields to be corrected, a semantic matching degree, and a correction field mark. The enterprise information context optimization matching result includes an optimized field relationship, semantic consistency marks, and a matching path adjustment record. The enterprise information event density includes the number of events and event distribution within a specified time. The enterprise information exposure enhancement result includes exposure frequency and exposure interval time.

[0049] Please refer to Figure 2 The semantic vector information of the enterprise information field is obtained by the following steps:

[0050] S111: Collecting enterprise structured data and unstructured data from multiple business sources and correlating them, wherein the structured data includes enterprise basic information and financial data, and the unstructured data includes news reports and social media dynamics, obtaining a cross-source enterprise information context data set;

[0051] For the collected structured and unstructured data of multiple business sources, first access the business information databases such as Qichacha and Tianyancha through the application program interface (API) to extract the structured data of a specific enterprise, for example, “A Technology Co., Ltd.” The specific fields include the unified social credit code ‘91310115MA1H888888’, registered capital ‘500 million RMB’, and establishment date ‘2015-03-01’. At the same time, access the Wind financial database to retrieve the latest financial report data of the company, i.e. the third quarter of 2024, to obtain the operating income ‘850 million RMB’ and net profit ‘72 million RMB’. Store these structured information in the ‘enterprise_structured_data’ table of the PostgreSQL database, and generate a unique primary key ID ‘A_TECH_001’ for “A Technology Co., Ltd.”. Then, use the Scrapy framework to write a web crawler to target the main financial news websites such as Sina Finance and East Money, and crawl the news reports published in the past six months that contain “A Technology Co., Ltd.” and “technology breakthrough” in the title or body, for example, obtain a news article titled “A Technology releases new photonic chip, leading industry change” and record its publication time ‘2024-08-20’ and body content. At the same time, use the official API of the microblog platform to obtain all the posts published by the official account “@A Technology Dynamics” in the past six months and the public microblog containing the topic tag ‘#A Technology New Product#’, and capture a post with the content “Our new generation of photonic chip is officially mass-produced today”. Store these crawled news body and social media text in Elasticsearch, and each data contains the source URL, publication timestamp, and content text. Finally, create a new association table ‘data_association’ in the PostgreSQL database, and map the structured data records of the enterprise with the IDs of all associated unstructured data documents in Elasticsearch through the enterprise primary key ID ‘A_TECH_001’, for example, establish an association record between ‘A_TECH_001’ and the news document ID ‘NEWS_SINA_20240820_015’ and the social media document ID ‘WEIBO_ATECH_20240905_003’, and align the timestamps of all data sources to form a cross-source enterprise information context dataset centered on the enterprise, integrating its information at different time points.

[0052] S112: Obtain enterprise industry classification parameters and enterprise information field label parameters, group enterprise information fields according to enterprise industry classification parameters, calculate the co-occurrence frequency between enterprise information fields, and construct an enterprise information context relationship graph according to the co-occurrence frequency;

[0053] After obtaining the enterprise industry classification parameter and the enterprise information field label parameter, first, according to the standard of “National Economic Industry Classification” (GB / T 4754-2017), the industry belonging of each enterprise in the data set is labeled, for example, it is found that “A Technology Co., Ltd.” belongs to “C39-Computer, Communication and Other Electronic Equipment Manufacturing Industry”, which is the enterprise industry classification parameter, at the same time, all information field labels appearing in the data set are sorted and extracted to form a field label parameter list, which contains items such as ‘registered capital’, ‘operating income’, ‘net profit’, ‘news report text’, ‘social media dynamic text’, etc., then, taking the industry classification parameter “C39” as the screening condition, all enterprises belonging to this industry are screened out from the data set, assuming that a total of 500 enterprises are screened out, and all the information fields owned by these enterprises are summarized, then the frequency of the common occurrence of these fields in pairs in the 500 enterprise samples is calculated, the specific calculation process is: traversing all field pairs, for example, (‘operating income’, ‘news report text’), checking how many of the 500 enterprises have both ‘operating income’ value and ‘news report text’ content in their data records, if 450 enterprises have both the two fields, then the co-occurrence frequency of the field pair is 450, and for the field pair (‘registered capital’, ‘social media dynamic text’), if only 150 enterprises have both, then the co-occurrence frequency is 150, after the co-occurrence frequency of all field pairs is calculated, a co-occurrence frequency threshold is set to construct a relationship, the setting of the threshold is based on the percentile of all calculated co-occurrence frequency values, and the specific setting is the 80th percentile value of all frequency values sorted from low to high, assuming that the range of all frequency values is between 10 and 490, the 80th percentile value is calculated to be 350, then the co-occurrence frequency threshold is set to 350, finally, according to the threshold, a graph is constructed, in which each information field label is a node, if the co-occurrence frequency between two fields is greater than 350, a connection edge is established between the two nodes, for example, the frequency of ‘operating income’ and ‘news report text’ is 450, which is greater than 350, so a connection is established between the two nodes, while the frequency of ‘registered capital’ and ‘social media dynamic text’ is 150, which is less than 350, so no connection is established, generating a graph composed of nodes and edges, which can represent the close relationship between information fields in a specific industry.

[0054] S113: According to the enterprise information context relationship graph, the cross-source enterprise information context data set is input into the Word2Vec model for semantic feature vectorization processing to obtain semantic vector information of the enterprise information field;

[0055] According to the enterprise information context relationship graph, the cross-source enterprise information context data set is input into the Word2Vec model for processing. This process first needs to convert the structured graph relationship into a serialized corpus. The specific method is to perform multiple random walks (RandomWalk) based on the graph. Each random walk starts from a random node (i.e., an information field label) in the graph and continuously moves on the graph. The moving rule is to jump from the current node to one of its connected neighbor nodes. This process continues until the preset walk length is reached. The walk length is set to 12. This value is set by referring to the average degree of the nodes in the graph and the diameter of the graph. It is set to a length that can cover most of the local connection relationships. For example, a random walk starting from the 'operating income' node may generate the sequence [ 'operating income', 'net profit', 'R&D investment', 'news report text', 'technology breakthrough', 'operating income' … ]. Repeat this process. Specifically, perform 200 random walks for each node in the graph to generate a large corpus composed of these field label sequences. These sequences are like sentences in natural language. Then, use the Word2Vec model in the Python Gensim library to train the corpus. In the model configuration, set the dimension of the semantic feature vector (vector_size) to 150. This value is set by considering the total number of field labels (assuming 200) and the desired vector precision. Generally, it is set to a multiple of the logarithm or square root of the total vocabulary. Here, a moderate value of 150 is taken. The context window size (window) is set to 5, which means that when predicting a field, the context relationship of the previous and next 5 fields in the sequence will be considered. The minimum word frequency (min_count) is set to 3, which means that fields with a frequency of less than 3 in the corpus will be ignored. After model training, each enterprise information field originally used as a text label, such as 'net profit', will be mapped and converted to a 150-dimensional floating-point array, i.e., [0.05, -0.23, 0.61,..., -0.11]. This array is the semantic vector information of the field.

[0056] Please refer to Figure 3 The steps for obtaining the cross-source enterprise information semantic similarity analysis result are as follows:

[0057] S211: According to the semantic vector information of the enterprise information field, the cosine similarity algorithm is used to calculate the semantic matching degree of the same type of enterprise information field in each business source, and the cross-source enterprise information semantic matching degree data is obtained.

[0058] After obtaining the semantic vector information of the enterprise information field, i.e. a 150-dimensional semantic feature vector corresponding to each field label, cosine similarity algorithm is used to quantify the matching degree between the same type of enterprise information fields with similar semantics in different business sources. For example, the semantic relevance between the structured field "operating income" from the authoritative financial database and the keyword "sales" extracted from the unstructured text in the financial news report is compared. The calculation process is based on the vector space model, which evaluates the consistency in the direction of two vectors by measuring the cosine value of the included angle between the two vectors. The closer the direction of the two vectors, the closer the cosine value to 1. If the direction of the two vectors is orthogonal, the cosine value is 0. If the direction of the two vectors is opposite, the cosine value is -1. In actual calculation, the function in the specific programming language library is called to perform the core calculation formula as follows: wherein, represents the cosine similarity calculation result of the vector and the vector , the value range is between -1 and 1, is the cosine function in the trigonometric function, is the included angle between the vector and the vector in the coordinate origin vertex in the dimensional space, and are the semantic vectors of two enterprise information fields to be compared, which are obtained from the Word2Vec model of S113, for example, may be a 150-dimensional vector of the structured data field "net profit", and may be a 150-dimensional vector of the keyword "company profit" extracted from the news report, represents the dot product of the vector and the vector , and respectively represent the Euclidean norm of the vector and the vector , also known as the module length, is the summation symbol, which means the results of the expressions after the symbol are added up, and the upper index and the lower index define the range of summation, wherein represents the dimension of the vector, the value of is 150, and the lower index is a counting variable, representing the th dimension of the vector, which takes a value from 1 to , respectively represent the vector and the vector In the first dimension, the specific component value, i.e. a floating point number, thus, the specific calculation process is to multiply the corresponding dimension components of the two vectors, and then add all the products together, while the specific calculation process is to square each component value of the vector , and then take the square root of the sum of all squares. By applying this calculation process to all pre-set pairs of the same type of fields, such as (“net profit”, “profit”), (“R&D investment”, “technology investment”), etc., a series of quantified matching degree values can be obtained, which together constitute the cross-source enterprise information semantic matching degree data.

[0059] Take a simplified 4-dimensional vector as an example. Suppose that after training the Word2Vec model, the semantic vectors of two information fields are obtained: field one is “operating income” from structured financial data, whose semantic vector is , and field two is “sales” from unstructured news reports, whose semantic vector is .

[0060] Calculate the dot product of the two vectors (the numerator part of the formula ): ;

[0061] Calculate the length of vector (the denominator part of the formula ): ;

[0062] Calculate the length of vector (the denominator part of the formula ): ;

[0063] Calculate the cosine similarity: ;

[0064] The calculation shows that the semantic similarity value of “operating income” and “sales” is about 0.9867.

[0065] By converting the semantic similarity of two enterprise information fields, such as “operating income” from financial reports and “sales” from news, into an objective and comparable specific numerical value, this process first regards each information field as a specific direction in a multi-dimensional meaning space, which represents the core meaning of the field.

[0066] S212: According to the cross-source enterprise information semantic matching degree data, identify the enterprise information fields with a semantic matching degree lower than the semantic matching degree threshold, and mark them as enterprise information fields that need to be corrected in context to obtain context correction marking data;

[0067] According to the cross-source enterprise information semantic matching degree data, this data contains a plurality of information field pairs and their corresponding cosine similarity values, such as ("net profit", "company profit", 0.991), ("registered capital", "paid-in capital", 0.852), and ("R&D expenses", "advertising", 0.314). Next, a clear semantic matching degree threshold will be set to identify enterprise information fields with insufficient semantic relevance. The setting process of the threshold is not arbitrary, but is based on the overall distribution of all similarity scores in the current data set. Specifically, first, extract all cosine similarity values from the data set and collect them into a numerical list. Then, sort all numerical values from low to high. Next, calculate the specific statistical position of the numerical distribution, i.e., calculate the first quartile. This position represents that 25% of the values in the data set are below this point. The reason for choosing the first quartile as the threshold is that it can objectively define the field pairs with the lowest matching degree scores in the data statistically. Assuming that the first quartile of the ten thousand similarity scores is 0.650, the semantic matching degree threshold is set to 0.650. After completing the threshold setting, each record in the cross-source enterprise information semantic matching degree data is traversed, and the similarity score in the record is compared with 0.650. For records with a similarity score lower than 0.650, such as the field pair ("R&D expenses", "advertising") with a similarity of 0.314, since 0.314 is lower than 0.650, both "R&D expenses" and "advertising" are identified and added to a newly created list, and a "need to correct" label is attached to each field. For records with a score not lower than 0.650, such as ("registered capital", "paid-in capital") with a score of 0.852, no operation is performed. By executing this identification and comparison process on all data records, a list is formed containing only the identified enterprise information fields with insufficient semantic matching degree and their attached labels.

[0068] S213: Integrate the context correction marking data with the cross-source enterprise information semantic matching degree data to generate a cross-source enterprise information semantic similarity analysis result;

[0069] After obtaining the context correction mark data and the cross-source enterprise information semantic matching degree data respectively, the two data are integrated. The context correction mark data is a list of all fields determined to be semantically mismatched, such as “R&D expenses” and “advertising placement”. The cross-source enterprise information semantic matching degree data is a more comprehensive record containing the original source, field name and semantic similarity score between all field pairs, such as the record (source: “financial statements”, field: “R&D expenses”; source: “news text”, field: “advertising placement”; similarity: 0.314). The specific implementation process of integration is to add a new data column to the cross-source enterprise information semantic matching degree data as the basis table, and name it “correction flag”. The initial value of this column is empty. Then, read the fields in the context correction mark data list one by one. For each field in the list, such as “R&D expenses”, the program will find all records containing “R&D expenses” in the cross-source enterprise information semantic matching degree data as the basis table. Once found, update the value of the “correction flag” column in the record to “yes”. Similarly, for the “advertising placement” field in the list, perform the same operation to update the “correction flag” column of the record to “yes”. For fields that never appear in the context correction mark data list, such as “net profit” and “company profit”, the “correction flag” column of their records will remain their initial empty value, or can be updated to “no” for clear distinction. This integration operation is equivalent to adding a qualitative judgment label based on the threshold to the original quantitative similarity analysis. After a series of data matching and information filling operations, a more detailed analysis summary table is formed. This summary table not only shows the semantic similarity values between any two information fields, but also clearly indicates which fields are to be manually reviewed or context corrected due to low similarity. This integrated summary table is the cross-source enterprise information semantic similarity analysis result.

[0070] Please refer to Figure 4 The steps for obtaining the enterprise information context optimization matching result are as follows:

[0071] S311: Obtain the enterprise information fields that need to be context corrected in the cross-source enterprise information semantic similarity analysis result, and call the original data of the enterprise information fields that need to be context corrected from the cross-source enterprise information context data set and establish a corresponding relationship to obtain the context correction field sequence;

[0072] First, the cross-source enterprise information semantic similarity analysis results are parsed, and all records with the "correction flag" set to "yes" are selected to obtain a list of enterprise information fields that need to be corrected in context, such as the list containing the "R&D expenses" and "advertising" fields of "A Technology Co., Ltd." Then, according to the field names in the list, the program will backtrack to the cross-source enterprise information context data set to accurately call the original data. Specifically, for the "R&D expenses" field, the program will locate the 2024 third quarter financial statements of "A Technology Co., Ltd." in the structured database and extract the corresponding original value '1.2 billion yuan', while recording the context of the data, including the report name, report period, and currency unit. Then, the same operation is performed on the "advertising" field in the list, and the original source is retrieved in the unstructured database, which is a news report titled "A Technology Market Trends" published on October 15, 2024, and the complete sentence containing the field is called "The company launched a one-month online and offline advertising campaign to promote its photon chip", and the source URL and publication timestamp of the news are recorded. Then, by binding the extracted field name, its corresponding original data content, and its source context, an explicit correspondence is established for each field that needs to be corrected. Finally, all "field-original data-context" correspondence items established for the company and other companies are integrated to form a structured sequence file, which is the context correction field sequence.

[0073] S312: According to the context correction field sequence, the transition probability between enterprise information fields is calculated according to the state transition rules of the conditional random field model, the arrangement order of enterprise information fields in the semantic matching path is adjusted, and the adjusted matching path is obtained.

[0074] According to the context correction field sequence generated by the previous process, a conditional random field model pre-trained on a large amount of information text of enterprises in the same industry is used to adjust the arrangement order of information fields in the sequence. The specific adjustment process is as follows: for any two possible adjacent fields in the sequence, the model calculates their "transition probability" through a two-step process to quantify the logical rationality of their connection. The first step is to calculate a comprehensive "association score", and the second step is to convert the score into a probability value between 0 and 1 through a function. The core calculation formula is as follows: Where the association score is calculated as follows: The introduction of each parameter in the formula is as follows, represents the probability, refers to the field of enterprise information currently being evaluated, such as "R&D expenses", while refers to its previous field in the sequence, such as "Net Profit", so, This overall expression represents the conditional probability that the next field is "R&D expenses" given that the previous field is "Net Profit", is the base of the natural logarithm, i.e. Euler's number, a mathematical constant approximately equal to 2.718, is a raw score calculated by the model for the relevance of these two fields, which integrates multiple dimensions of features, and the higher the value, the closer the association between the two fields, the weight , and are coefficients automatically determined by the model after learning a large amount of data in the training stage, used to balance the influence of different features on the total score, and their respective settings are as follows, the weight The setting basis of is the reliability of source consistency, as fields from the same structured report (such as financial statements) have a strong internal logical association, which is a very stable and reliable basis for judgment, so it is given the highest weight, for example, set to 0.6, the weight The setting basis of is the importance of text physical proximity, the closer the distance between fields in the original text, the greater the possibility of association, but this association is not as certain as source consistency, so it is given a lower weight, for example, set to 0.2, the weight The setting basis of is the auxiliary role of pure semantic similarity, semantic similarity can reflect general association, but may not be consistent with the specific context, so it is given a lower weight as a supplementary basis for judgment, for example, set to 0.2, is the source consistency feature value, which is obtained by judging whether the two fields come from the same original file, if the source is the same, the value is 1, otherwise it is 0, is the context proximity feature value, which is obtained by calculating the distance between the positions of the two fields in the original text, if they appear in the same sentence, the value is 1.0, if in adjacent sentences, the value is 0.5, if more than three sentences apart, the value is 0, is the semantic relevance feature value, which directly calls the cosine similarity value between the two fields, the value ranges from 0 to 1, after calculating the transition probabilities between all pairs of fields, the program will evaluate all possible field rearrangements, by multiplying the transition probabilities between all adjacent fields in each arrangement, a total sequence probability is obtained, the program will select the arrangement with the highest total sequence probability.

[0075] The fields "net profit" and "R&D expenses" are both extracted from the same financial statement. In the financial report, they are located in the same income statement section, but in different rows, and physically belong to adjacent sentences or paragraphs. The semantic similarity between them is calculated to be high.

[0076] Determining the feature values, source consistency : Since both fields are from the same financial report, the source is completely consistent, so Contextual proximity : Since they are located in adjacent paragraphs, according to the rules, the value is set to Semantic relevance : Assuming the cosine similarity value calculated earlier is .

[0077] Calculating the association score , substituting the above feature values and weights into the association score formula: .

[0078] Calculating the transition probability , substituting the calculated association score into the function: .

[0079] The probability of transitioning from "net profit" to "R&D expenses" is calculated to be about 0.7048. This is a high probability value, indicating that in the current context.

[0080] Calculating the transition probability of "net profit" to "advertising" sets the "net profit" field extracted from the financial statement, and the "advertising" field extracted from a news report. Since the sources are different, they are not physically adjacent. The semantic similarity between them is calculated to be low.

[0081] Determining the feature values, source consistency : Since the two fields come from different files (financial report vs. news report), the sources are not consistent, so .

[0082] Contextual proximity : Since the sources are different and not physically adjacent, so .

[0083] Semantic relevance : Assuming the cosine similarity value calculated earlier is .

[0084] Calculating the association score , substituting the above feature values and weights into the association score formula: .

[0085] Calculate transition probability , and substitute the calculated correlation score into the function: .

[0086] Result interpretation: The probability of transition from "net profit" to "advertising" is approximately 0.5157. This probability value is very close to 0.5, indicating that the model considers the connection between these two fields to be almost equivalent to random guessing, lacking a clear logical relationship.

[0087] By comparing these two calculation results (0.7048 vs 0.5157), the program can determine that the score of connecting "net profit" and "R&D expenses" in the same path is much higher than that of connecting "net profit" and "advertising". Therefore, when generating the adjusted matching path, the program will choose the former, while placing "advertising" in another path that better matches its context.

[0088] S313: Call the adjusted matching path and the associated field content in the cross-source enterprise information context dataset, reconstruct the semantic matching path of enterprise information and extract the corresponding enterprise identifier according to the optimal path output rule of conditional random field model, and generate the optimized matching result of enterprise information context;

[0089] After obtaining the adjusted matching path, the information will be reconstructed according to the field order list optimized by the context logic, for example, for "A Technology Co., Ltd.", one path has been formed as ["net profit", "R&D expenses"], and another path as ["new product launch", "advertising placement"]. First, the program calls each field in the path and returns to the cross-source enterprise information context data set to extract the original data content and its source directly associated with the field, for example, for the path ["net profit", "R&D expenses"], the program extracts the corresponding values '7200 million yuan' and '1.2 billion yuan' and attaches the source description such as "derived from the financial statements of the third quarter of 2024". Then, the path is processed according to the optimal path output rule of the conditional random field model. The specific implementation of this rule is to directly adopt the path with the highest total sequence probability calculated in the previous process as the effective connection. This total sequence probability is obtained by multiplying the transition probability values between all adjacent fields in the path. Selecting the path with the highest total probability is equivalent to finding the only "storyline" that best fits the logic and context in statistics among all possible connection methods. After confirming the optimal path, the program concatenates the fields in the path with their corresponding original data content in the optimized order to reconstruct a clear and coherent enterprise information semantic matching path. At the same time, the program extracts the enterprise unique identifier associated with the path from the original data record, i.e. 'A_TECH_001', and attaches this identifier to the front end of the reconstructed path. Finally, all the optimized information paths of the enterprise obtained through this reconstruction and attribution are collected together to form a structured and high-credibility data set, which is the enterprise information context optimization matching result.

[0090] Please refer to Figure 5 The steps for obtaining the enterprise information event density are as follows:

[0091] S411: Obtain time-labeled event data associated with the enterprise and associate it with the enterprise identifier in the enterprise information context optimization matching result to obtain the associated event data set.

[0092] First, a time-stamped event data associated with the enterprise is obtained, which is a record set containing a large number of events of different enterprises at different time points, and each record explicitly contains three core information: the specific date of the event, for example, "November 5, 2024"; the text description of the event content, for example, "sign a cooperation agreement with B company"; and the complete name of the enterprise associated with the event, for example, "A Technology Co., Ltd.", then the event data is associated with the optimized matching result of the enterprise information context generated in the previous process, and the optimized matching result assigns a unique and non-repeating enterprise identifier to each enterprise, such as "A Technology Co., Ltd.", the specific operation process of association is to read the records in the time-stamped event data one by one, for each record, the complete name of the enterprise is extracted, then the name is used as a query condition to find and obtain the corresponding unique enterprise identifier in the optimized matching result of the enterprise information context, finally, the unique enterprise identifier obtained by the query is combined with the original date and content description of the event record to form a new data entry, by performing the above finding and combining operation on all event records, a new data set is generated, in which each event is accurately bound to its corresponding unique enterprise identifier, which is the associated event data set.

[0093] S412: Based on the associated event data set, the number of events associated with each enterprise in a set time range is counted, and the weighted data of the number of events is obtained by combining the weighted value of each event;

[0094] Based on the associated event data set generated in the previous process, the event activity level of each enterprise in a specific time window is calculated, first, a specific statistical time range needs to be set, the setting of this range is based on the standard financial reporting period to ensure the synchronization and comparability of the analysis, for example, set to the fourth quarter of 2024, that is, from 2024-10-01 00:00 to 2024-12-31 24:00, next, a pre-prepared "event type weight table" needs to be referenced, which assigns different weighted values to events of different natures, the setting of which is based on the quantitative analysis results of the correlation between different event types in historical data and subsequent media report volume and market sentiment fluctuations, the higher the correlation, the higher the weighted value of the event, for example, the weighted value of "core technology breakthrough" is 1.5, the weighted value of "signing a major contract" is 1.2, the weighted value of "high-level change" is 1.0, and the weighted value of "publishing a routine news release" is 0.8, the calculation process is as follows: for each enterprise identifier, the program will first filter out all event records of the enterprise in the set time range, then the weighted values of these events are summed, the calculation formula is: wherein, represents the weight (Weight), and the subscript represents the sum, thus This whole represents the "event quantity weighted value" calculated for a single enterprise; is a summation symbol, indicating the accumulation of subsequent numerical values; is a counting variable, representing the i-th event of an enterprise within a time range, and indicates counting from the first event; and represents the total number of events of the enterprise within the set time range; represents weight, subscript indicates corresponding to the i-th event, thus represents the i-th event; corresponding weighted value found from the "event type weight table", and through the calculation of this formula, the event quantity weighted data is formed.

[0095] S413: Calculate the enterprise information event density in the corresponding time range according to the event quantity weighted data;

[0096] After obtaining the event quantity weighted data, the information event density of each enterprise within a specific time range is calculated, which contains the enterprise identifier and its total weighted event score within the fourth quarter of 2024. The calculation of event density requires a specific time span as a benchmark, which is based on the total number of days in the selected time range. For the fourth quarter of 2024, the total number of days is 92. The specific calculation operation is to take the total weighted event score of each enterprise, and then divide it by the total number of days in the time range, to obtain a standardized daily event density value. The calculation formula is: , wherein, represents density (Density), subscript represents events, thus This whole represents the "enterprise information event density" to be solved; represents the "event quantity weighted value" of a single enterprise; represents time (Time), subscript represents the number of days, thus represents the total number of days in the statistical time range (e.g. 92 days). This density value reflects the average daily information event influence of the enterprise in this time period. Finally, the identifiers of all enterprises and their calculated daily event density values are paired and combined to form a quantitative analysis result.

[0097] Please refer to Figure 6 , the steps for obtaining the enterprise information exposure enhancement result are as follows:

[0098] S511: According to the enterprise information event density, compare with the preset frequency control threshold value to determine whether the enterprise information exposure interval time parameter needs to be adjusted, and obtain exposure adjustment requirement data;

[0099] First, the enterprise information event density data is obtained, which provides each enterprise with its daily event influence value in a certain time range. Then, the density value is compared with a preset "frequency control threshold range". The threshold range is set based on the statistical distribution of the historical event density of all enterprises in the data set. The upper limit is set as the third quartile of the historical data, i.e. the seventy-fifth percentile point, and the lower limit is set as the first quartile, i.e. the twenty-fifth percentile point. For example, after statistical analysis, the range is set as 0.5 to 1.0. This range represents the normal interval of event activity. In the comparison process, for each enterprise, check whether its daily event density value falls within this normal interval. If the event density value of an enterprise is higher than 1.0, it is determined that the events are too dense and the information exposure interval time needs to be shortened. If the density value is lower than 0.5, it is determined that the events are too sparse and the exposure interval time needs to be extended. For enterprises with density values between 0.5 and 1.0, it is determined that the frequency is normal and no adjustment is needed. Finally, all enterprises that need to be adjusted are identified and their adjustment direction (shorten or extend) is recorded to form an exposure adjustment requirement data.

[0100] S512: Based on the exposure adjustment requirement data, analyze the event density change amplitude, and perform step-by-step adjustment when the change exceeds the set smoothing control threshold value. Adjust the exposure interval time to compress or extend the exposure frequency to obtain exposure interval adjustment data.

[0101] Based on the exposure adjustment demand data, the specific exposure interval adjustment range is determined, and the sudden change of exposure frequency caused by the sharp fluctuation of event density is prevented. Specifically, for each adjustment demand, the program first analyzes the change range of the event density, that is, the event density value of the current quarter is subtracted from the event density value of the last quarter, and then the difference is divided by the density value of the last quarter to obtain a change rate. Then, the change rate is compared with a preset "smoothing control threshold". The threshold is set to avoid the user feeling that the information push frequency changes too harshly, and the value is set to 50%. If the absolute value of the calculated change rate exceeds 50%, the step-by-step adjustment mechanism is triggered. For example, if the exposure interval of an enterprise needs to be shortened from 8 hours to 4 hours, a total of 4 hours is shortened, and the step-by-step adjustment will be divided into two times, which is adjusted to 6 hours this week and 4 hours next week to realize smooth transition. If the absolute value of the change rate does not exceed 50%, a one-time adjustment is performed to directly adjust the exposure interval to the right position. Through this calculation and judgment, a clear and new exposure interval time value will be generated for each adjustment demand, forming exposure interval adjustment data.

[0102] S513: Call the exposure interval adjustment data and the exposure parameters in the enterprise information context optimization matching result, so that the adjusted exposure interval will be automatically adjusted according to the event density value within the time window, and generate the enterprise information exposure enhancement result;

[0103] After obtaining the exposure interval adjustment data, the data and the previously generated enterprise information context optimization matching result are called to update the exposure parameters of each enterprise information. Specifically, the program will find the corresponding entry in the enterprise information context optimization matching result according to the enterprise identifier in the exposure interval adjustment data. These entries contain the original set exposure parameters. Then, the program will replace the original exposure parameter value with the new and specific exposure interval time value calculated for the enterprise in the exposure interval adjustment data. In this way, the adjusted exposure interval can be automatically compressed or extended according to the event density value within the recent time window. The whole process ensures that the information push frequency can dynamically adapt to the current information activity of each enterprise. After completing the parameter update for all enterprises that need to be adjusted, the obtained data set is the enterprise information exposure enhancement result.

[0104] The above merely describes the preferred embodiments of the present application, and is not intended to limit the present application in other forms. Any skilled person in the art can modify or change the disclosed technical content into equivalent embodiments with equivalent changes, and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application, without departing from the technical solution content of the present application, still falls within the protection scope of the present application.

Claims

1. A method for enhancing enterprise information exposure based on multi-source data fusion, characterized in that, Includes the following steps: S1: Collect structured and unstructured data from multiple business sources, and perform semantic feature vectorization processing to generate semantic vector information for enterprise information fields; S2: Based on the semantic vector information of the enterprise information fields, analyze the semantic matching degree of the same type of enterprise information fields in each business source, mark the enterprise information fields that need to be corrected in context according to the semantic matching degree, and generate cross-source enterprise information semantic similarity analysis results; S3: Extract the enterprise information fields that need to be context-corrected from the cross-source enterprise information semantic similarity analysis results, reconstruct the semantic matching path of enterprise information, and output the enterprise information context-optimized matching results; S4: Count the number of events associated with the enterprise within a specified time period, and associate them with the enterprise identifier in the enterprise information context optimization matching result to calculate the enterprise information event density within the target time period; S5: Compare the enterprise information event density with a preset frequency control threshold, adjust the enterprise information exposure interval time parameter corresponding to the enterprise information context optimization matching result, and obtain the enterprise information exposure enhancement result.

2. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 1, characterized in that, The semantic vector information includes enterprise basic information vectors, financial data vectors, news report vectors, and social media dynamic vectors. The cross-source enterprise information semantic similarity analysis results include a list of fields to be corrected, semantic matching degree, and corrected field markers. The enterprise information context optimization matching results include optimized field relationships, semantic consistency flags, and matching path adjustment records. The enterprise information event density includes the number of events and the event distribution within a specified time period. The enterprise information exposure enhancement results include exposure frequency and exposure interval time.

3. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 1, characterized in that, The specific steps for obtaining the semantic vector information of the enterprise information field are as follows: S111: Collect and correlate structured and unstructured data from multiple business sources. The structured data includes basic enterprise information and financial data, while the unstructured data includes news reports and social media activity, resulting in a cross-source enterprise information context dataset. S112: Obtain enterprise industry classification parameters and enterprise information field label parameters, group enterprise information fields according to enterprise industry classification parameters, calculate the co-occurrence frequency among enterprise information fields, and construct an enterprise information context relationship graph based on the co-occurrence frequency; S113: Based on the enterprise information context relationship graph, the cross-source enterprise information context dataset is input into the Word2Vec model for semantic feature vectorization processing to obtain the semantic vector information of the enterprise information field.

4. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 3, characterized in that, The specific steps for obtaining the semantic similarity analysis results of cross-source enterprise information are as follows: S211: Based on the semantic vector information of the enterprise information fields, the cosine similarity algorithm is used to calculate the semantic matching degree of the same type of enterprise information fields in each business source, and obtain cross-source enterprise information semantic matching degree data. S212: Based on the cross-source enterprise information semantic matching degree data, identify enterprise information fields whose semantic matching degree is lower than the semantic matching degree threshold, and mark them as enterprise information fields that need to be corrected in context, and obtain context correction marking data; S213: Integrate the context correction tag data with the cross-source enterprise information semantic matching degree data to generate cross-source enterprise information semantic similarity analysis results.

5. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 4, characterized in that, The specific steps for obtaining the enterprise information context optimization matching result are as follows: S311: Obtain the enterprise information fields that need to be context-corrected from the cross-source enterprise information semantic similarity analysis results, and call the original data of the enterprise information fields that need to be context-corrected from the cross-source enterprise information context dataset and establish a corresponding relationship to obtain the context-corrected field sequence; S312: Based on the context, correct the field sequence, combine the context content order of the enterprise information field in each source and the matching situation of the associated fields, calculate the transition probability between enterprise information fields according to the state transition rules of the conditional random field model, adjust the arrangement order of enterprise information fields in the semantic matching path, and obtain the adjusted matching path; S313: Call the adjusted matching path and the associated field content in the cross-source enterprise information context dataset, reconstruct the semantic matching path of enterprise information according to the optimal path output rule of the conditional random field model, extract the corresponding enterprise identifier, and generate the enterprise information context optimization matching result.

6. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 5, characterized in that, The specific steps for obtaining the enterprise information event density are as follows: S411: Obtain time-stamped event data associated with the enterprise, and associate it with the enterprise identifier in the enterprise information context optimization matching result to obtain the associated event dataset; S412: Based on the aforementioned associated event dataset, count the number of events associated with each enterprise within a set time range, and combine the weighted value of each event to obtain weighted event count data; S413: Calculate the enterprise information event density within the corresponding time range based on the weighted data of the event quantity.

7. The enterprise information exposure enhancement method based on multi-source data fusion according to claim 6, characterized in that, The specific steps for obtaining the enhanced enterprise information exposure results are as follows: S511: Based on the enterprise information event density, compare it with the preset frequency control threshold to determine whether the enterprise information exposure interval time parameter needs to be adjusted, and obtain exposure adjustment demand data; S512: Based on the exposure adjustment requirement data, analyze the magnitude of event density change. When the change exceeds the set smoothing control threshold, perform step-by-step adjustment by adjusting the exposure interval time to compress or extend the exposure frequency to obtain exposure interval adjustment data. S513: Call the exposure parameters in the exposure interval adjustment data and the enterprise information context optimization matching result, so that the adjusted exposure interval will be automatically adjusted according to the event density value within the time window, and generate enterprise information exposure enhancement result.

Citation Information

Patent Citations

  • Enterprise risk evaluation system and method based on open source data

    CN110020048A

  • Semantic-based enterprise research and development resource information modeling method

    CN113065343A