Enterprise information fusion method and system for multi-source heterogeneous data cross-checking
By using a multi-source heterogeneous data cross-validation method, the authority and timeliness of information sources are dynamically evaluated, solving the problems of information timeliness decay and insufficient cross-source correlation consistency in existing technologies, and improving the robustness and sensitivity of enterprise information evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHIYI SHUPU DATA SERVICE CO LTD
- Filing Date
- 2026-02-13
- Publication Date
- 2026-06-05
AI Technical Summary
Existing assessment technologies fail to dynamically reflect the decay of information timeliness and lack cross-source information correlation consistency verification, resulting in inaccurate assessment results and difficulty in adapting to sudden changes in the enterprise information environment.
A multi-source heterogeneous data cross-validation method is adopted. By obtaining a set of information items from multiple sources, source weights, timeliness attenuation, cross-validation, and dynamic attenuation coefficients are assigned to construct a comprehensive credibility index, thereby achieving dynamic integration of information source authority, content timeliness, and correlation consistency.
It enhances the dynamic adaptability and risk warning sensitivity of enterprise information credibility assessment, ensuring the robustness and accuracy of assessment results.
Smart Images

Figure CN122153781A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information fusion technology, and more specifically, to an enterprise information fusion method and system for cross-validation of multi-source heterogeneous data. Background Technology
[0002] Global corporate information comes from a variety of sources, including publicly available government data, news reports, financial reports, and industry association data.
[0003] Existing assessment techniques often employ single-source information verification or static scoring models to calculate fixed-weighted credibility scores. However, these techniques suffer from two prominent problems: first, they fail to dynamically reflect the decay of information over time, leading to outdated negative information having a long-term impact on assessment results; second, they lack cross-source information consistency verification mechanisms, resulting in insufficient ability to identify contradictory information from different channels and affecting assessment accuracy. Furthermore, fixed-weight models cannot adapt to sudden changes in the enterprise information environment, exhibiting delayed responses to major change events and failing to meet the needs of real-time risk assessment. Summary of the Invention
[0004] This invention provides a method and system for enterprise information fusion through cross-validation of multi-source heterogeneous data. This method addresses the technical challenges of rigid timeliness processing and lack of cross-source correlation analysis in existing assessment systems when dealing with multi-source heterogeneous information. It achieves the technical effect of dynamically integrating the authority of information sources, the timeliness of content, and the consistency of correlation into a credibility index. This significantly improves the dynamic adaptability and risk warning sensitivity of enterprise information credibility assessment while ensuring the robustness of assessment results.
[0005] To achieve the above objectives, the present invention provides a method for enterprise information fusion through cross-validation of multi-source heterogeneous data, comprising:
[0006] Acquire a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. According to the first preset source weighting rule at the current evaluation time, each information item in the information item set is weighted by source credibility to obtain a source weighting coefficient. According to the second preset timeliness decay rule at the current evaluation time, the timeliness of each information item in the information item set is quantified to obtain a timeliness decay coefficient. The source weighting coefficient and the timeliness decay coefficient are combined to obtain the basic credibility coefficient of each information item. The information entry set is processed based on the current evaluation time to obtain the set of information pairs to be verified, the cross-validation coefficient at the current evaluation time is determined, and the association consistency coefficient at the current evaluation time is obtained based on the distribution characteristics of the cross-validation coefficient. Based on the change in information content of the information item set between the current evaluation time and the previous evaluation time, a content fluctuation factor is obtained. Based on the dispersion of information source distribution of the information item set at the current evaluation time, a source stability factor is obtained. The content fluctuation factor and the source stability factor are fused to obtain the dynamic decay coefficient at the current evaluation time. Based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, the comprehensive credibility index of the target enterprise at the current evaluation time is obtained, and the multi-source heterogeneous enterprise information of multiple target enterprises is cross-validated and fused based on the comprehensive credibility index.
[0007] Furthermore, the method for determining the basic credibility coefficient includes: In the first preset source weighting rule, the source level value, historical accuracy and authentication status value corresponding to the information source identifier are obtained; The source level value, historical accuracy, and authentication status value are normalized and then weighted and fused to obtain the source weight coefficient. In the second preset time decay rule, the time interval between the current evaluation time and the timestamp is calculated to determine the time decay coefficient; The basic credibility coefficient is obtained by multiplying the source weight coefficient and the time-related decay coefficient.
[0008] Furthermore, the method for determining the source grade value includes: Information sources are categorized into three levels: first-tier official sources, second-tier cooperative sources, and third-tier publicly available sources. A first preset benchmark value is assigned to the first-level official source, a second preset benchmark value is assigned to the second-level cooperative source, and a third preset benchmark value is assigned to the third-level public source. The first preset benchmark value is greater than the second preset benchmark value, and the second preset benchmark value is greater than the third preset benchmark value. Based on the level to which the information source identifier belongs, the corresponding benchmark value is matched as the source level value.
[0009] Furthermore, in the second preset time-decrease rule, calculating the time interval between the current evaluation time and the timestamp, and determining the time-decrease coefficient, includes: At the current evaluation time, based on the difference between the time interval value corresponding to each information item in the information item set and the interval belonging to the first preset time threshold and the second preset time threshold, the information item set is divided into time-sensitive clusters, corresponding to obtain short-term effective information clusters, transitional decay information clusters and long-term invalid information clusters. Based on the degree of concentration of the time interval values within each information cluster, short-term concentration factor, transitional dispersion factor and long-term sparsity factor are determined respectively. The basic attenuation rate is determined by integrating the relative distance relationship between the preset time threshold and the current evaluation time; The short-time concentration factor, the transitional dispersion factor, and the long-time sparsity factor are respectively differentially coupled with the basic attenuation rate to obtain the short-time attenuation response value, the transitional attenuation response value, and the long-time attenuation response value. Based on the information cluster category to which each information entry belongs, the corresponding attenuation response value is used as the time-degradation coefficient for that information entry.
[0010] Furthermore, when processing the set of information entries based on the current evaluation time to obtain the set of information pairs to be verified and determining the cross-validation coefficients at the current evaluation time, the process includes: Based on the third preset association constraint at the current evaluation time, two information items whose information content belongs to the same information category and whose information source identifiers are different are selected to form the information pair to be verified. The two information contents in the information pair to be verified are structured and parsed to extract the numerical data field and the text description field respectively; Calculate the relative deviation rate between the numerical data fields, and calculate the semantic similarity between the text description fields; The cross-validation coefficients are obtained by fusing the relative deviation rate and the semantic similarity.
[0011] Furthermore, when calculating the relative deviation rate between the numerical data fields and the semantic similarity between the textual description fields, the following steps are included: Calculate the ratio of the absolute difference between two values to the maximum value, and use the ratio as the relative deviation rate; Input the text description field into a pre-trained semantic matching model to obtain a semantic vector; Calculate the cosine similarity between two semantic vectors, and use the cosine similarity as the semantic similarity.
[0012] Further, when obtaining a content fluctuation factor based on the change in information content of the information item set between the current evaluation time and the previous evaluation time, and obtaining a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and fusing the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time, the process includes: At the current evaluation moment, the total number of information entries in the information entry set is counted and used as the first statistical value; At the previous assessment point, the total number of statistical information items was used as the second statistical value; Calculate the absolute value of the difference between the first statistical value and the second statistical value, and normalize the absolute value of the difference to use as the content fluctuation factor; At the current evaluation moment, the number of categories of different information source identifiers in the information item set is counted, and the ratio of the number of categories to the first statistical value is used as the source distribution entropy value; The source distribution entropy value is normalized and used as the source stability factor; The dynamic attenuation coefficient is obtained by weighting and summing the content fluctuation factor and the source stability factor.
[0013] Furthermore, when obtaining the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, the process includes: Along the time axis, the basic confidence coefficients of the current evaluation time and several previous evaluation times are arranged in chronological order to construct a basic confidence response sequence; The correlation consistency coefficient is mapped to a time-series correlation weight value, and the basic credibility response sequence is weighted and fused according to the time-series correlation weight value to obtain a first credibility response value. Filter out a set of historical evaluation times that are within a preset neighborhood of the dynamic decay coefficient at the current evaluation time, and obtain the historical correlation mapping pair between the dynamic decay coefficient and the comprehensive credibility index at each time in the set of historical evaluation times; Linear regression analysis was performed on the historical correlation mapping pairs, and the fitting slope was used as the historical trend factor, and the fitting intercept was used as the benchmark offset factor. The second confidence response value is obtained by combining the first confidence response value, the historical trend factor, and the benchmark offset factor; Using the dynamic decay coefficient as an exponential smoothing parameter, the second credibility response value is smoothed and converged to obtain the comprehensive credibility index at the current evaluation time.
[0014] Furthermore, when performing cross-validation and fusion of multi-source heterogeneous enterprise information for multiple target enterprises based on the comprehensive credibility index, the process includes: A preset comprehensive credibility index threshold is established, and multi-source heterogeneous enterprise information of target enterprises with comprehensive credibility indices greater than the preset threshold is cross-validated and fused.
[0015] To achieve the above objectives, the present invention also provides an enterprise information fusion system for cross-validation of multi-source heterogeneous data, comprising: The information acquisition module is used to acquire a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. The credibility assessment module is used to assign source credibility weight to each information item in the information item set according to the first preset source weight rule at the current assessment time to obtain the source weight coefficient, and to quantify the timeliness of each information item in the information item set according to the second preset timeliness decay rule at the current assessment time to obtain the timeliness decay coefficient. The module then merges the source weight coefficient and the timeliness decay coefficient to obtain the basic credibility coefficient of each information item. The coefficient association module is used to process the set of information entries based on the current evaluation time to obtain a set of information pairs to be verified, determine the cross-validation coefficients at the current evaluation time, and obtain the association consistency coefficients at the current evaluation time based on the distribution characteristics of the cross-validation coefficients. The information dispersion module is used to obtain a content fluctuation factor based on the change range of information content of the information item set between the current evaluation time and the previous evaluation time, obtain a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and fuse the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time. The information fusion module is used to obtain the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient and the dynamic decay coefficient, and to perform cross-validation and fusion of multi-source heterogeneous enterprise information of multiple target enterprises based on the comprehensive credibility index.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention discloses a method and system for cross-validating enterprise information from multiple heterogeneous sources. The method involves acquiring a set of multi-source information entries from multiple target enterprises; obtaining source weight coefficients according to a first preset source weight rule, and obtaining time-decrease coefficients according to a second preset time-decrease rule, thus obtaining a basic credibility coefficient; determining cross-validation coefficients and correlation consistency coefficients; obtaining content fluctuation factors and source stability factors based on the magnitude of information content changes, thus obtaining a dynamic decay coefficient; and obtaining a comprehensive credibility index based on the basic credibility coefficient, correlation consistency coefficient, and dynamic decay coefficient. This method performs cross-validation and fusion of multi-source heterogeneous enterprise information from multiple target enterprises, determining the credibility index based on the authority of the information source, the timeliness of the content, and the correlation consistency, ensuring the robustness of the evaluation results and improving the dynamic adaptability and risk warning sensitivity of enterprise information credibility assessment. Attached Figure Description
[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating an enterprise information fusion method based on cross-validation of multi-source heterogeneous data in an embodiment of the present invention is shown. Figure 2 The diagram shows a structural schematic of an enterprise information fusion system for cross-validation of multi-source heterogeneous data in an embodiment of the present invention. Detailed Implementation
[0018] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0019] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0020] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0021] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0022] The following is a description of preferred embodiments of the present invention in conjunction with the accompanying drawings.
[0023] like Figure 1As shown, an embodiment of the present invention discloses a method for enterprise information fusion through cross-validation of multi-source heterogeneous data, comprising: S110: Obtain a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. S120: According to the first preset source weight rule at the current evaluation time, each information item in the information item set is weighted by source credibility to obtain a source weight coefficient. According to the second preset timeliness decay rule at the current evaluation time, the timeliness of each information item in the information item set is quantified to obtain a timeliness decay coefficient. The source weight coefficient and the timeliness decay coefficient are combined to obtain the basic credibility coefficient of each information item. S130: Process the set of information entries based on the current evaluation time to obtain a set of information pairs to be verified, determine the cross-validation coefficient at the current evaluation time, and obtain the association consistency coefficient at the current evaluation time based on the distribution characteristics of the cross-validation coefficient. S140: Based on the change range of information content of the information item set at the current evaluation time and the previous evaluation time, obtain the content fluctuation factor; based on the dispersion of information source distribution of the information item set at the current evaluation time, obtain the source stability factor; and fuse the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time. S150: Based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, obtain the comprehensive credibility index of the target enterprise at the current evaluation time, and perform cross-validation and fusion of multi-source heterogeneous enterprise information of multiple target enterprises based on the comprehensive credibility index.
[0024] In some embodiments of this application, the method for determining the basic credibility coefficient includes: In the first preset source weighting rule, the source level value, historical accuracy and authentication status value corresponding to the information source identifier are obtained; The source level value, historical accuracy, and authentication status value are normalized and then weighted and fused to obtain the source weight coefficient. In the second preset time decay rule, the time interval between the current evaluation time and the timestamp is calculated to determine the time decay coefficient; The basic credibility coefficient is obtained by multiplying the source weight coefficient and the time-related decay coefficient.
[0025] In this embodiment, the preset time window can be set to 7 days, 30 days, or 90 days, and can be flexibly adjusted according to the evaluation accuracy requirements. The multi-source information entry set refers to enterprise-related information collected from different channels such as business registration platforms, company websites, news reports, social media, and regulatory announcements. Information content includes specific text or numerical data such as changes in registered capital, changes in senior management, adjustments to business scope, and administrative penalty records. The information source identifier is used to distinguish different source types such as official channels, partner institutions, and public media, and is stored in string encoding format. The timestamp is accurate to the hour, recording the specific time point of information release. Historical accuracy is calculated by statistically analyzing the accuracy of information released by the source over the past year; the percentage of accurate information out of the total number of information is the historical accuracy. For example, if 85 out of 100 pieces of information released by a source are accurate, the historical accuracy is 0.85. The authentication status value is determined based on whether the source is officially certified; certified sources have a value of 1.0, and uncertified sources have a value of 0.5. Normalization maps the three parameters to the interval between 0 and 1, using a minimum-maximum normalization method. In the weighted fusion process, the source rating value is weighted at 0.4, the historical accuracy value at 0.5, and the authentication status value at 0.1, fully reflecting the importance of historical performance. The time interval is calculated in hours; the current evaluation time is subtracted from the timestamp to obtain the number of hours.
[0026] The beneficial effects of the above technical solution are: by constructing a multi-dimensional evaluation system that considers source authority, historical performance, and certification status, and by combining the time decay effect, the initial credibility of each piece of information is quantified, providing a reliable data foundation for subsequent comprehensive evaluation.
[0027] In some embodiments of this application, the method for determining the source grade value includes: Information sources are categorized into three levels: first-tier official sources, second-tier cooperative sources, and third-tier publicly available sources. A first preset benchmark value is assigned to the first-level official source, a second preset benchmark value is assigned to the second-level cooperative source, and a third preset benchmark value is assigned to the third-level public source. The first preset benchmark value is greater than the second preset benchmark value, and the second preset benchmark value is greater than the third preset benchmark value. Based on the level to which the information source identifier belongs, the corresponding benchmark value is matched as the source level value.
[0028] In this embodiment, the first-level official sources specifically include official platforms operated by regulatory agencies such as the Enterprise Credit Information Disclosure System, China Judgments Online, and information disclosure platforms designated by the China Securities Regulatory Commission, with a first preset benchmark value set at 0.9. The second-level cooperative sources include commercial banks, credit rating agencies, and industry association member systems that have signed data cooperation agreements with enterprises, with a second preset benchmark value set at 0.6. The third-level public sources include commercial and financial media and social media platforms, with a third preset benchmark value set at 0.3. The classification is based on three dimensions: the authority of information release, the strictness of data review mechanisms, and the ability to assume legal responsibility. Information source identifiers are matched using a pre-established source coding library, with each source corresponding to a unique level label. For example, the ".gov.cn" domain suffix is automatically identified as the first level, "credit" as a cooperative institution identifier as the second level, and "media" as a public media identifier as the third level. The difference between benchmark values is set to 0.3 to ensure significant differentiation between different levels. The matching process uses a lookup table method to establish a mapping relationship table between source identifiers and benchmark values for quick retrieval of corresponding values. The source level value is configured once during system initialization and can be dynamically adjusted based on source performance.
[0029] The beneficial effects of the above technical solution are: through standardized grading and benchmark value assignment mechanisms, objective quantitative evaluation of information sources can be achieved, avoiding the subjectivity of human judgment and improving the consistency and operability of the evaluation system.
[0030] In some embodiments of this application, when calculating the time interval between the current evaluation time and the timestamp to determine the time decay coefficient in the second preset time decay rule, the following steps are included: At the current evaluation time, based on the difference between the time interval value corresponding to each information item in the information item set and the interval belonging to the first preset time threshold and the second preset time threshold, the information item set is divided into time-sensitive clusters, corresponding to obtain short-term effective information clusters, transitional decay information clusters and long-term invalid information clusters. Based on the degree of concentration of the time interval values within each information cluster, short-term concentration factor, transitional dispersion factor and long-term sparsity factor are determined respectively. The basic attenuation rate is determined by integrating the relative distance relationship between the preset time threshold and the current evaluation time; The short-time concentration factor, the transitional dispersion factor, and the long-time sparsity factor are respectively differentially coupled with the basic attenuation rate to obtain the short-time attenuation response value, the transitional attenuation response value, and the long-time attenuation response value. Based on the information cluster category to which each information entry belongs, the corresponding attenuation response value is used as the time-degradation coefficient for that information entry.
[0031] In this embodiment, the first preset time threshold is set to 24 hours, and the second preset time threshold is set to 168 hours (7 days). Information entries with time intervals less than 24 hours are classified as short-term valid information clusters, those between 24 and 168 hours are classified as transitionally decaying information clusters, and those greater than 168 hours are classified as long-term invalid information clusters. The degree of distribution concentration is obtained by calculating the standard deviation of the time interval values within each cluster. Clusters with a standard deviation less than 5 hours are considered highly concentrated, with a short-term concentration factor of 0.9; those with a standard deviation between 5 and 20 hours are considered moderately dispersed, with a transitional dispersion factor of 0.5; and those with a standard deviation greater than 20 hours are considered highly sparse, with a long-term sparsity factor of 0.2. The relative distance between the preset time threshold and the current evaluation time is determined by calculating the number of days between the threshold time point and the current time. The closer the distance, the greater the basic decay rate. The initial value of the basic decay rate is set to 0.05. Differential coupling employs a multiplication mechanism: the short-term concentrated factor is multiplied by the base attenuation rate to obtain a short-term attenuation response value of 0.045; the transitional dispersed factor is multiplied by the base attenuation rate to obtain a transitional attenuation response value of 0.025; and the long-term sparse factor is multiplied by the base attenuation rate to obtain a long-term attenuation response value of 0.01. Each information entry is automatically assigned to a corresponding information cluster based on its time interval value, and the attenuation response value of that cluster is directly used as the time-dependent attenuation coefficient.
[0032] The beneficial effects of the above technical solution are: by dividing time into three levels and using a clustering mechanism, it achieves refined attenuation control of information timeliness, avoids excessive uniform processing of information with different time spans by a single attenuation model, and improves the accuracy of timeliness assessment.
[0033] In some embodiments of this application, when processing the set of information entries based on the current evaluation time to obtain a set of information pairs to be verified and determining the cross-validation coefficients at the current evaluation time, the process includes: Based on the third preset association constraint at the current evaluation time, two information items whose information content belongs to the same information category and whose information source identifiers are different are selected to form the information pair to be verified. The two information contents in the information pair to be verified are structured and parsed to extract the numerical data field and the text description field respectively; Calculate the relative deviation rate between the numerical data fields, and calculate the semantic similarity between the text description fields; The cross-validation coefficients are obtained by fusing the relative deviation rate and the semantic similarity.
[0034] In this embodiment, the third preset association constraint is set as a pairing rule for information of the same category but different sources. The information categories include five major categories: enterprise registration information, senior management information, administrative penalties, judicial proceedings, and operational anomalies. Structured parsing uses natural language processing technology to identify numerical entities and descriptive content in the text. The rule for extracting numerical data fields is to identify combinations of continuous numbers and currency units, while the rule for extracting textual descriptive fields is to extract the remaining non-numerical descriptive statements. Fusion uses a weighted average method, with a relative deviation rate weight of 0.4 and a semantic similarity weight of 0.6.
[0035] In this embodiment, the association consistency coefficient is obtained by averaging all cross-validation coefficients and then multiplying it by a normalization factor for the number of information pairs.
[0036] The beneficial effects of the above technical solution are: through the dual verification mechanism of numerical deviation and semantic similarity, it can achieve a refined consistency assessment of cross-source information, effectively identify information contradictions and content similarities, and improve the accuracy of cross-validation.
[0037] In some embodiments of this application, calculating the relative deviation rate between the numerical data fields and the semantic similarity between the textual description fields includes: Calculate the ratio of the absolute difference between two values to the maximum value, and use the ratio as the relative deviation rate; Input the text description field into a pre-trained semantic matching model to obtain a semantic vector; Calculate the cosine similarity between two semantic vectors, and use the cosine similarity as the semantic similarity.
[0038] In this embodiment, the pre-trained semantic matching model is the Sentence-BERT model, which is trained on a large number of sentence pairs and can generate high-quality semantic representations. The semantic vector dimension is a 768-dimensional floating-point array. Cosine similarity is calculated using the standard formula of vector dot product divided by the product of the magnitudes. When two text description fields are exactly the same, the cosine similarity is 1.0; when the semantics are completely opposite, the similarity is close to 0.
[0039] The beneficial effects of the above technical solution are: by distinguishing between valid and invalid numerical values through a processing mechanism, combined with advanced semantic vector technology, accurate calculation of deviation rate and similarity can be achieved, thereby improving the accuracy of the cross-validation process.
[0040] In some embodiments of this application, when obtaining a content fluctuation factor based on the change in information content of the information item set between the current evaluation time and the previous evaluation time, obtaining a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and fusing the content fluctuation factor and the source stability factor to obtain the dynamic attenuation coefficient at the current evaluation time, the method includes: At the current evaluation moment, the total number of information entries in the information entry set is counted and used as the first statistical value; At the previous assessment point, the total number of statistical information items was used as the second statistical value; Calculate the absolute value of the difference between the first statistical value and the second statistical value, and normalize the absolute value of the difference to use as the content fluctuation factor; At the current evaluation moment, the number of categories of different information source identifiers in the information item set is counted, and the ratio of the number of categories to the first statistical value is used as the source distribution entropy value; The source distribution entropy value is normalized and used as the source stability factor; The dynamic attenuation coefficient is obtained by weighting and summing the content fluctuation factor and the source stability factor.
[0041] In this embodiment, the first statistical value is obtained by traversing the set of information entries and counting. For example, if 85 pieces of information are collected at the current moment, the first statistical value is 85. The second statistical value is read from the system cache, which contains the statistical result from the previous moment, for example, 80 pieces of information at the previous moment. Normalization divides the difference by the first statistical value; 5 divided by 85 yields 0.0588, which is used as the content fluctuation factor. In the calculation of the source distribution entropy value, the category quantity is counted after deduplication using a hash set. Assuming that 85 pieces of information come from 12 different sources, the ratio 12 divided by 85 equals 0.1412. Normalization maps the source distribution entropy value to the interval between 0 and 1, using the method of dividing by the maximum possible number of sources. The maximum number of sources is set to 50, so 0.1412 divided by 50 yields 0.0028, which is used as the source stability factor. The weighted summation is calculated as follows: content volatility factor weight 0.6, source stability factor weight 0.4, and dynamic decay coefficient equals 0.0588 multiplied by 0.6 plus 0.0028 multiplied by 0.4, which equals 0.0364.
[0042] The beneficial effects of the above technical solution are: by quantifying the volatility of information quantity and the dispersion of source distribution, a dynamic decay mechanism is constructed to reflect the stability and reliability changes of enterprise information status in real time, thereby enhancing the dynamic adaptability of the evaluation system.
[0043] In some embodiments of this application, when obtaining the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, the following steps are included: Along the time axis, the basic confidence coefficients of the current evaluation time and several previous evaluation times are arranged in chronological order to construct a basic confidence response sequence; The correlation consistency coefficient is mapped to a time-series correlation weight value, and the basic credibility response sequence is weighted and fused according to the time-series correlation weight value to obtain a first credibility response value. Filter out a set of historical evaluation times that are within a preset neighborhood of the dynamic decay coefficient at the current evaluation time, and obtain the historical correlation mapping pair between the dynamic decay coefficient and the comprehensive credibility index at each time in the set of historical evaluation times; Linear regression analysis was performed on the historical correlation mapping pairs, and the fitting slope was used as the historical trend factor, and the fitting intercept was used as the benchmark offset factor. The second confidence response value is obtained by combining the first confidence response value, the historical trend factor, and the benchmark offset factor; Using the dynamic decay coefficient as an exponential smoothing parameter, the second credibility response value is smoothed and converged to obtain the comprehensive credibility index at the current evaluation time.
[0044] In this embodiment, the number of evaluation times is set to 5, constructing a basic credibility response sequence containing 6 time points, such as [0.65, 0.68, 0.70, 0.72, 0.75, 0.78]. The correlation consistency coefficient is mapped to the time-series correlation weight value using a linear mapping, with a coefficient of 0.8 mapping to a weight value of 0.8. The weighted average of the sequence elements is calculated using weighted fusion to obtain the first credibility response value of 0.73. The preset neighborhood range is set to the range of ±20% of the dynamic decay coefficient, for example, if the current dynamic decay coefficient is 0.0364, the neighborhood range is 0.0291 to 0.0437. The historical evaluation time set is filtered to include all times within the past 30 days where the dynamic decay coefficient falls within this range, assuming 12 historical times are selected. The historical correlation mapping records the dynamic decay coefficient and the corresponding comprehensive credibility index of each of these 12 times, forming 12 sets of data points. Linear regression analysis uses the least squares method to fit a straight line. The historical trend factor of the fitted slope reflects the influence of the decay coefficient on the credibility index; for example, it is calculated to be 0.5. The baseline offset factor of the fitted intercept reflects the basic credibility level; for example, it is calculated to be 0.6. The fusion calculation multiplies the first credibility response value (0.73) by the historical trend factor (0.5) to obtain 0.365, and adds the baseline offset factor (0.6) to obtain the second credibility response value (0.965). The exponential smoothing parameter uses the dynamic decay coefficient (0.0364) as a smoothing coefficient to smooth the second credibility response value (0.965); the calculation formula is 0.965. 0.0364 + 0.73 (1-0.0364) yields 0.73, where... The sign is a multiplication sign, which serves as the final overall credibility index.
[0045] The beneficial effects of the above technical solution are: by introducing historical trend analysis and index smoothing mechanism, the time-series correlation mining and fluctuation smoothing of the credibility index can be realized, which not only reflects historical patterns but also maintains the sensitivity of current assessment, thereby improving the robustness and predictability of the comprehensive credibility index.
[0046] In some embodiments of this application, when performing cross-validation and fusion of multi-source heterogeneous enterprise information of multiple target enterprises based on the comprehensive credibility index, the following is included: A preset comprehensive credibility index threshold is established, and multi-source heterogeneous enterprise information of target enterprises with comprehensive credibility indices greater than the preset threshold is cross-validated and fused.
[0047] In this embodiment, the comprehensive credibility index threshold is preferably 0.7.
[0048] The beneficial effects of the above technical solution are: to perform consistency analysis and calculation on enterprise information from different sources, and to generate a high-precision enterprise database.
[0049] To further illustrate the technical concept of this invention, the technical solution of this invention will now be described in conjunction with specific application scenarios.
[0050] Correspondingly, such as Figure 2 As shown, this application also provides an enterprise information fusion system for cross-validation of multi-source heterogeneous data, including: The information acquisition module is used to acquire a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. The credibility assessment module is used to assign source credibility weight to each information item in the information item set according to the first preset source weight rule at the current assessment time to obtain the source weight coefficient, and to quantify the timeliness of each information item in the information item set according to the second preset timeliness decay rule at the current assessment time to obtain the timeliness decay coefficient. The module then merges the source weight coefficient and the timeliness decay coefficient to obtain the basic credibility coefficient of each information item. The coefficient association module is used to process the set of information entries based on the current evaluation time to obtain a set of information pairs to be verified, determine the cross-validation coefficients at the current evaluation time, and obtain the association consistency coefficients at the current evaluation time based on the distribution characteristics of the cross-validation coefficients. The information dispersion module is used to obtain a content fluctuation factor based on the change range of information content of the information item set between the current evaluation time and the previous evaluation time, obtain a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and fuse the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time. The information fusion module is used to obtain the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient and the dynamic decay coefficient, and to perform cross-validation and fusion of multi-source heterogeneous enterprise information of multiple target enterprises based on the comprehensive credibility index.
[0051] In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0052] Although the invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the embodiments disclosed in this invention can be combined with each other in any way. The fact that not all of these combinations are described in this specification is merely for the sake of brevity and resource conservation.
[0053] It will be understood by those skilled in the art that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for enterprise information fusion through cross-validation of multi-source heterogeneous data, characterized in that, include: Acquire a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. According to the first preset source weighting rule at the current evaluation time, each information item in the information item set is weighted by source credibility to obtain a source weighting coefficient. According to the second preset timeliness decay rule at the current evaluation time, the timeliness of each information item in the information item set is quantified to obtain a timeliness decay coefficient. The source weighting coefficient and the timeliness decay coefficient are combined to obtain the basic credibility coefficient of each information item. The information entry set is processed based on the current evaluation time to obtain the set of information pairs to be verified, the cross-validation coefficient at the current evaluation time is determined, and the association consistency coefficient at the current evaluation time is obtained based on the distribution characteristics of the cross-validation coefficient. Based on the change in information content of the information item set between the current evaluation time and the previous evaluation time, a content fluctuation factor is obtained. Based on the dispersion of information source distribution of the information item set at the current evaluation time, a source stability factor is obtained. The content fluctuation factor and the source stability factor are fused to obtain the dynamic decay coefficient at the current evaluation time. Based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, the comprehensive credibility index of the target enterprise at the current evaluation time is obtained, and the multi-source heterogeneous enterprise information of multiple target enterprises is cross-validated and fused based on the comprehensive credibility index.
2. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 1, characterized in that, The method for determining the basic credibility coefficient includes: In the first preset source weighting rule, the source level value, historical accuracy and authentication status value corresponding to the information source identifier are obtained; The source level value, historical accuracy, and authentication status value are normalized and then weighted and fused to obtain the source weight coefficient. In the second preset time decay rule, the time interval between the current evaluation time and the timestamp is calculated to determine the time decay coefficient; The basic credibility coefficient is obtained by multiplying the source weight coefficient and the time-related decay coefficient.
3. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 2, characterized in that, The method for determining the source grade value includes: Information sources are categorized into three levels: first-tier official sources, second-tier cooperative sources, and third-tier publicly available sources. A first preset benchmark value is assigned to the first-level official source, a second preset benchmark value is assigned to the second-level cooperative source, and a third preset benchmark value is assigned to the third-level public source. The first preset benchmark value is greater than the second preset benchmark value, and the second preset benchmark value is greater than the third preset benchmark value. Based on the level to which the information source identifier belongs, the corresponding benchmark value is matched as the source level value.
4. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 2, characterized in that, In the second preset time-degradation rule, calculating the time interval between the current evaluation time and the timestamp, and determining the time-degradation coefficient, includes: At the current evaluation time, based on the difference between the time interval value corresponding to each information item in the information item set and the interval belonging to the first preset time threshold and the second preset time threshold, the information item set is divided into time-sensitive clusters, corresponding to obtain short-term effective information clusters, transitional decay information clusters and long-term invalid information clusters. Based on the degree of concentration of the time interval values within each information cluster, short-term concentration factor, transitional dispersion factor and long-term sparsity factor are determined respectively. The basic attenuation rate is determined by integrating the relative distance relationship between the preset time threshold and the current evaluation time; The short-time concentration factor, the transitional dispersion factor, and the long-time sparsity factor are respectively differentially coupled with the basic attenuation rate to obtain the short-time attenuation response value, the transitional attenuation response value, and the long-time attenuation response value. Based on the information cluster category to which each information entry belongs, the corresponding attenuation response value is used as the time-degradation coefficient for that information entry.
5. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 1, characterized in that, When processing the set of information entries based on the current evaluation time to obtain the set of information pairs to be verified, and determining the cross-validation coefficients at the current evaluation time, the process includes: Based on the third preset association constraint at the current evaluation time, two information items whose information content belongs to the same information category and whose information source identifiers are different are selected to form the information pair to be verified. The two information contents in the information pair to be verified are structured and parsed to extract the numerical data field and the text description field respectively; Calculate the relative deviation rate between the numerical data fields, and calculate the semantic similarity between the text description fields; The cross-validation coefficients are obtained by fusing the relative deviation rate and the semantic similarity.
6. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 5, characterized in that, When calculating the relative deviation rate between the numerical data fields and the semantic similarity between the textual description fields, the following steps are included: Calculate the ratio of the absolute difference between two values to the maximum value, and use the ratio as the relative deviation rate; Input the text description field into a pre-trained semantic matching model to obtain a semantic vector; Calculate the cosine similarity between two semantic vectors, and use the cosine similarity as the semantic similarity.
7. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 1, characterized in that, When obtaining a content fluctuation factor based on the change in information content of the information item set between the current evaluation time and the previous evaluation time, and obtaining a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and then fusing the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time, the process includes: At the current evaluation moment, the total number of information entries in the information entry set is counted and used as the first statistical value; At the previous assessment point, the total number of statistical information items was used as the second statistical value; Calculate the absolute value of the difference between the first statistical value and the second statistical value, and normalize the absolute value of the difference to use as the content fluctuation factor; At the current evaluation moment, the number of categories of different information source identifiers in the information item set is counted, and the ratio of the number of categories to the first statistical value is used as the source distribution entropy value; The source distribution entropy value is normalized and used as the source stability factor; The dynamic attenuation coefficient is obtained by weighting and summing the content fluctuation factor and the source stability factor.
8. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 1, characterized in that, When obtaining the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient, and the dynamic decay coefficient, the following steps are included: Along the time axis, the basic confidence coefficients of the current evaluation time and several previous evaluation times are arranged in chronological order to construct a basic confidence response sequence; The correlation consistency coefficient is mapped to a time-series correlation weight value, and the basic credibility response sequence is weighted and fused according to the time-series correlation weight value to obtain a first credibility response value. Filter out a set of historical evaluation times that are within a preset neighborhood of the dynamic decay coefficient at the current evaluation time, and obtain the historical correlation mapping pair between the dynamic decay coefficient and the comprehensive credibility index at each time in the set of historical evaluation times; Linear regression analysis was performed on the historical correlation mapping pairs, and the fitting slope was used as the historical trend factor, and the fitting intercept was used as the benchmark offset factor. The second confidence response value is obtained by combining the first confidence response value, the historical trend factor, and the benchmark offset factor; Using the dynamic decay coefficient as an exponential smoothing parameter, the second credibility response value is smoothed and converged to obtain the comprehensive credibility index at the current evaluation time.
9. The enterprise information fusion method based on multi-source heterogeneous data cross-validation according to claim 1, characterized in that, When performing cross-validation and fusion of multi-source heterogeneous enterprise information for multiple target enterprises based on the comprehensive credibility index, the following is included: A preset comprehensive credibility index threshold is established, and multi-source heterogeneous enterprise information of target enterprises with comprehensive credibility indices greater than the preset threshold is cross-validated and fused.
10. A multi-source heterogeneous data cross-validation enterprise information fusion system, applied to the multi-source heterogeneous data cross-validation enterprise information fusion method as described in any one of claims 1-9, characterized in that, include: The information acquisition module is used to acquire a set of multi-source information entries from multiple target companies within a preset time window. Each information entry includes information content, information source identifier, and timestamp. The credibility assessment module is used to assign source credibility weight to each information item in the information item set according to the first preset source weight rule at the current assessment time to obtain the source weight coefficient, and to quantify the timeliness of each information item in the information item set according to the second preset timeliness decay rule at the current assessment time to obtain the timeliness decay coefficient. The module then merges the source weight coefficient and the timeliness decay coefficient to obtain the basic credibility coefficient of each information item. The coefficient association module is used to process the set of information entries based on the current evaluation time to obtain a set of information pairs to be verified, determine the cross-validation coefficients at the current evaluation time, and obtain the association consistency coefficients at the current evaluation time based on the distribution characteristics of the cross-validation coefficients. The information dispersion module is used to obtain a content fluctuation factor based on the change range of information content of the information item set between the current evaluation time and the previous evaluation time, obtain a source stability factor based on the dispersion of information source distribution of the information item set at the current evaluation time, and fuse the content fluctuation factor and the source stability factor to obtain the dynamic decay coefficient at the current evaluation time. The information fusion module is used to obtain the comprehensive credibility index of the target enterprise at the current evaluation time based on the basic credibility coefficient, the correlation consistency coefficient and the dynamic decay coefficient, and to perform cross-validation and fusion of multi-source heterogeneous enterprise information of multiple target enterprises based on the comprehensive credibility index.