Medium and small-sized enterprise diagnosis method and system based on multi-source data fusion and AI
By assimilating and segmenting multi-source data from SMEs, reconstructing state vectors and risk feature vectors using a diagnostic analysis engine, dynamically adjusting weight distribution, and generating concurrent diagnostic subtasks, the problem of data integration in the operational health and risk assessment of SMEs is solved, achieving efficient and adaptive diagnostic analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-06
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for assessing the operational health and risks of SMEs rely on analysis from single or limited data sources, resulting in narrow data dimensions and rigid analysis processes. This makes it difficult to comprehensively and dynamically depict the true state of the enterprise. Furthermore, traditional diagnostic models cannot effectively integrate multi-source heterogeneous data, leading to a one-sided diagnostic perspective, low efficiency, and insufficient consistency in conclusions.
By acquiring multi-source data from SMEs, the data is assimilated and then divided into multiple data slices. Data purification and information enhancement are performed on each slice. The diagnostic analysis engine is used to reconstruct the state vector and risk feature vector, dynamically adjust the weight distribution, generate and execute diagnostic subtasks concurrently, and verify and synthesize diagnostic reports in real time.
It enables efficient, adaptive, and coordinated multi-dimensional diagnostic analysis for SMEs, ensuring the internal consistency and efficiency of diagnostic results. It breaks through the bottleneck of traditional methods, provides more granular and targeted diagnostic inputs, and overcomes the problems of information loss and blurred focus.
Smart Images

Figure CN121808644A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of enterprise intelligent diagnostic technology, specifically to a diagnostic method and system for small and medium-sized enterprises based on multi-source data fusion and AI. Background Technology
[0002] Currently, the assessment of operational health and risks for SMEs generally relies on the analysis of single or limited data sources, such as audit conclusions of financial statements, qualitative judgments from market research reports, or isolated data from internal business systems. These methods typically employ linear or static models, resulting in narrow data dimensions and rigid analysis processes. The massive, multi-source, and heterogeneous data generated both internally and externally is difficult to effectively integrate and collaboratively utilize, leading to a one-sided diagnostic perspective and an inability to comprehensively and dynamically depict the true state of the enterprise.
[0003] Existing technical solutions, when processing multi-source data, often involve simple merging or selective use, lacking a refined data governance mechanism based on business logic, time attributes, and security levels. This results in critical information and anomaly signals being buried in data noise. Furthermore, traditional diagnostic models are typically holistic, sequential analysis processes, unable to dynamically adjust the analysis focus based on the inherent relationships within the data, and struggling to handle complex multi-dimensional diagnostic tasks in parallel. This leads to low analysis efficiency, and intervening intermediate conclusions from different analysis modules may conflict, resulting in insufficient comprehensiveness and consistency in the final report. The purpose of this invention is to solve the challenges of refined information extraction and fusion under multi-source heterogeneous data, and to achieve efficient, adaptive, and coordinated concurrent diagnostic analysis. Summary of the Invention
[0004] The purpose of this invention is to provide a diagnostic method and system for small and medium-sized enterprises based on multi-source data fusion and AI, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, this invention provides a diagnostic method for small and medium-sized enterprises based on multi-source data fusion and AI, the method comprising:
[0006] Acquire the business operation data, industry trend data, market feedback data, and compliance audit data required for the diagnosis of SMEs, and perform format assimilation processing to form a preliminary fusion dataset;
[0007] The primary fusion dataset is segmented according to business domain, time period, and data sensitivity to obtain multiple data slice units. Data purification and information enhancement operations are performed on each data slice unit to mark core information nodes and abnormal information nodes.
[0008] All marked data slices are input into the diagnostic analysis engine. Based on the core information node, the diagnostic analysis engine aggregates related nodes to form feature clusters, extracts representative feature values of the feature clusters and arranges and fills them to obtain a state vector, and evaluates the abnormal intensity and transmission risk of various abnormal information nodes based on the abnormal information nodes, and generates a risk feature vector by combining them according to a preset risk vector structure.
[0009] The state vector and the risk feature vector are cross-correlated and compared. The weight distribution of the state vector and the risk feature vector is dynamically adjusted according to the preset diagnostic rule network to generate a comprehensive diagnostic framework with multiple interrelated diagnostic subtasks. Computational resources and data streams are allocated to each diagnostic subtask to start concurrent execution.
[0010] The system receives intermediate diagnostic results from various concurrently executed diagnostic subtasks in real time, performs consistency verification and conflict resolution on the intermediate diagnostic results, and synthesizes and assembles the processed intermediate diagnostic results to generate a structured diagnostic report containing multiple diagnostic dimensions and specific diagnostic items.
[0011] Preferably, the format assimilation process to form a primary fused dataset includes:
[0012] Identify the native data format and storage structure corresponding to the business operation data, the industry trend data, the market feedback data, and the compliance audit data;
[0013] For the aforementioned business operation data, the time-series transaction records are converted into standard event sequences under a unified spatiotemporal coordinate system;
[0014] For the industry trend data, extract the trend description text and numerical indicators, and convert the trend description text into a numerical trend vector.
[0015] Based on the market feedback data, sentiment polarity analysis and topic clustering are performed on the unstructured comments and rating information to generate a quantitative feedback rating matrix;
[0016] For the aforementioned compliance audit data, we analyze its internal legal references and compliance checkpoints to construct a structured compliance rule map;
[0017] The standard event sequence, the numerical trend vector, the quantified feedback scoring matrix, and the structured compliance rule graph are aligned and stitched together through a preset data fusion channel to form the primary fusion dataset.
[0018] Preferably, the initial fused dataset is segmented according to business domain, time period, and data sensitivity dimensions to obtain multiple data slice units, including:
[0019] Identify all business area labels related to SME operations from the primary fusion dataset;
[0020] Based on the business domain labels, the initial fusion dataset is preliminarily divided into multiple business domain data blocks;
[0021] For each business area data block, the continuous operational timeline is divided into multiple discrete time windows according to a preset time granularity;
[0022] Within each time window, the data is classified into sensitivity levels based on the categories and density of personally identifiable information, trade secret information, and financially sensitive information contained in the data;
[0023] Based on the sensitivity classification results, the data of each business domain data block within each time window is further divided into independent data sub-blocks with different sensitivity levels;
[0024] Each independent data sub-block is encapsulated into a data slice unit, and each data slice unit is attached with metadata tags containing business domain, time window, and sensitivity level.
[0025] Preferably, the step of performing data cleansing and information enhancement operations on each data slice unit, and marking core information nodes and abnormal information nodes, includes:
[0026] Missing values are detected in the data slice unit, and based on the business domain characteristics indicated by the metadata tags of the data slice unit, an appropriate data imputation strategy is selected to impute the missing values.
[0027] Noise detection is performed on the filled data slice unit, and a noise reduction algorithm matching the sensitivity level of the data slice unit is used to filter out random noise;
[0028] The information density of the data slices after noise reduction is evaluated, and regions whose information entropy exceeds a preset threshold are identified as high information density regions.
[0029] The preset threshold is a quantitative reference value pre-set during the information density assessment process;
[0030] In the high information density area, frequently occurring patterns, numerical points that significantly deviate from the historical average, and records that conform to the preset key event characteristics are extracted, and the extraction results are marked as the core information nodes;
[0031] In areas other than the high information density region, detect whether there are data points that violate business logic, segments that fluctuate continuously and rapidly in the numerical sequence, or records that match entries that are explicitly prohibited in the compliance rule map, and mark the detection results as the abnormal information nodes.
[0032] Preferably, the diagnostic analysis engine, based on the core information node, aggregates related nodes to form feature clusters, extracts representative feature values from the feature clusters, and arranges and fills them to obtain a state vector, including:
[0033] Collect the core information nodes from all data slice units and group them according to the business domain and time window to which the core information nodes belong;
[0034] Within each combination of business domain and time window, the semantic and numerical relationships between the core information nodes are analyzed, and multiple closely related core information nodes are aggregated into a state feature cluster.
[0035] Representative feature values are extracted from each state feature cluster, including but not limited to: average value, peak value, rate of change, and frequency;
[0036] The feature values of all state feature clusters under different business domains and different time windows are arranged and filled according to a preset global state vector template to form the state vector. The global state vector template defines the position, dimension and normalization method of each feature value in the vector.
[0037] The steps for building the diagnostic analysis engine include:
[0038] Obtain a historical diagnostic case library, which contains complete operational datasets of multiple historical small and medium-sized enterprises and their final confirmed diagnostic conclusions;
[0039] Extract the occurrence patterns of the core information nodes and the abnormal information nodes from the historical diagnostic case database, and their correspondence with the diagnostic conclusions to form a diagnostic knowledge graph;
[0040] Based on the diagnostic knowledge graph, a neural network model with multi-layer perception capability is constructed. The input layer of the neural network model is designed to receive the dimension of the state vector and the risk feature vector.
[0041] An attention mechanism is embedded in the hidden layer of the neural network model to automatically learn the influence weights of different state features in the state vector and different risk dimensions in the risk feature vector on the diagnostic conclusion during training.
[0042] The neural network model is trained under supervision using the historical diagnostic case library until the consistency between the diagnostic conclusions output by the model and the historical real conclusions reaches a preset threshold.
[0043] The trained neural network model is coupled with the diagnostic knowledge graph and encapsulated as the diagnostic analysis engine.
[0044] Preferably, the step of evaluating the anomaly intensity and transmission risk of various anomaly information nodes based on the anomaly information nodes, and generating a risk feature vector by combining them according to a preset risk vector structure, includes:
[0045] Collect the abnormal information nodes from all data slice units, and classify the abnormal information nodes into categories including data quality abnormalities, business logic abnormalities, compliance abnormalities, and trend abnormalities.
[0046] For each type of anomalous information node, its anomalous strength is evaluated. The anomalous strength is calculated based on the degree to which the anomalous information node deviates from the normal baseline and its spatiotemporal aggregation density.
[0047] For each type of abnormal information node, its transmission risk is assessed. The transmission risk is estimated based on the upstream and downstream associations of the abnormal information node in the business process and its co-occurrence relationship with other abnormal categories.
[0048] Each anomaly category and its corresponding anomaly intensity assessment value and transmission risk assessment value are mapped to a multidimensional risk sub-vector;
[0049] All risk subvectors of anomaly categories are sequentially connected according to a preset risk vector structure to generate the risk feature vector.
[0050] Preferably, the step of cross-correlation comparison between the state vector and the risk feature vector includes:
[0051] Establish an association mapping table between each state feature in the state vector and each risk dimension in the risk feature vector. The association mapping table describes the probability that a state change may induce or aggravate a specific risk.
[0052] Based on the association mapping table, calculate the potential influence of the current operating state represented by the state vector on each risk dimension in the risk feature vector;
[0053] Simultaneously, the potential erosion effect of the existing risks represented by the risk feature vector on each state feature in the state vector is calculated;
[0054] Perform matrix operations on the potential influence and the potential erosion effect to generate a correlation strength matrix between state and risk;
[0055] By analyzing the correlation strength matrix, we can identify the state and risk combinations whose correlation strength exceeds the warning threshold, and use these state and risk combinations as key interaction points to focus on in the diagnostic analysis.
[0056] Preferably, the step of dynamically adjusting the weight distribution of the state vector and the risk feature vector according to a preset diagnostic rule network includes:
[0057] The preset diagnostic rule network contains multiple diagnostic rule nodes, each corresponding to a specific diagnostic knowledge or experience logic;
[0058] The state vector, the risk feature vector, and the correlation strength matrix are input into the preset diagnostic rule network;
[0059] Each diagnostic rule node, based on its internal logic, judges the input state, risk, and correlation strength, and outputs weight adjustment suggestions for specific state features in the state vector and specific risk dimensions in the risk feature vector.
[0060] The system aggregates weight adjustment suggestions from all diagnostic rule nodes and coordinates and integrates potentially conflicting adjustment suggestions through a weight arbitration mechanism.
[0061] Based on the weight adjustment suggestions after coordination and fusion, the weight coefficients of each state feature in the state vector and the weight coefficients of each risk dimension in the risk feature vector are updated in real time.
[0062] Preferably, the real-time reception of intermediate diagnostic results from each concurrently executed diagnostic subtask, and the performance of consistency verification and conflict resolution on the intermediate diagnostic results, includes:
[0063] Establish a shared intermediate diagnostic results storage area to receive and temporarily store the intermediate diagnostic results output by all diagnostic subtasks;
[0064] For each intermediate diagnostic result, attach the identifier of the source diagnostic subtask, the timestamp of its generation, and the version information of the input data it is based on;
[0065] Set up a consistency checker, which continuously scans the intermediate storage area of the diagnostic results to check whether there are numerical contradictions, logical conflicts or contradictory conclusions in the intermediate diagnostic results from different diagnostic subtasks for the same diagnostic item.
[0066] When a conflict is detected, the consistency checker sorts the intermediate diagnostic results of the conflict according to the preset confidence level of the source diagnostic subtask, the time of generation, and the integrity of the input data on which the intermediate diagnostic results are based.
[0067] Based on the reliability ranking, select the best intermediate diagnostic result, or trigger a negotiation protocol between diagnostic subtasks to generate a negotiated and consistent result to resolve conflicts.
[0068] Preferably, the present invention also includes a diagnostic system for small and medium-sized enterprises based on multi-source data fusion and AI. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the diagnostic method for small and medium-sized enterprises based on multi-source data fusion and AI as described above.
[0069] Compared with the prior art, the beneficial effects of the present invention are:
[0070] The system collaboratively segments the fused primary dataset based on three dimensions: business domain, time period, and data sensitivity, generating data slices with clearly defined semantic boundaries. This multi-dimensional segmentation mechanism enables subsequent data cleansing and information enhancement operations to be more targeted, employing differentiated processing strategies based on the business scenario, time stage, and sensitivity level of each slice. Building upon this cleansing, the system accurately identifies and marks core and anomalous information nodes within each slice, transforming the originally chaotic data stream into a structured information set anchored by key nodes. This provides finer-grained and more targeted input for deep analysis, overcoming the shortcomings of traditional methods that suffer from information loss or blurred focus when processing coarse-grained data.
[0071] The diagnostic analysis engine utilizes identified core and abnormal nodes to reconstruct a state vector representing the overall operational status and a risk feature vector revealing potential problems. By cross-referencing and comparing these two sets of vectors and dynamically adjusting their weight distribution based on a pre-defined network of diagnostic rules, the system can focus on the most critical diagnostic dimensions in real time. Based on this, a comprehensive framework consisting of multiple interconnected diagnostic subtasks is automatically generated, and appropriate computing resources and data flows are allocated to each subtask to drive their concurrent execution. This mechanism achieves flexible allocation of analytical resources and parallel decoupling of tasks. After receiving intermediate results from each subtask in real time, the system performs consistency checks and conflict resolution, ultimately assembling the coordinated results into a unified, structured diagnostic report. This enables complex enterprise diagnostic processes to operate in a modular and efficient manner, ensuring internal consistency of the final conclusions and overcoming the efficiency and coordination bottlenecks of traditional sequential execution and rigid model diagnostic modes. Attached Figure Description
[0072] Figure 1 This is a schematic diagram illustrating the working principle of the SME diagnostic method based on multi-source data fusion and AI described in this invention.
[0073] Figure 2 A flowchart for splitting the primary fusion dataset;
[0074] Figure 3 A flowchart for generating risk feature vectors;
[0075] Figure 4 A two-bar chart comparing the weight adjustments for risk dimensions;
[0076] Figure 5 This is a comparison chart showing the sequential and parallel execution of diagnostic subtasks. Detailed Implementation
[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0078] Please see Figure 1 This invention provides a diagnostic method for SMEs based on multi-source data fusion and AI. The method includes: acquiring business operation data, industry trend data, market feedback data, and compliance audit data required for SME diagnosis. This multi-source data needs to undergo format assimilation processing to form a preliminary fusion dataset. The preliminary fusion dataset is then segmented according to business domain, time period, and data sensitivity dimensions to obtain multiple data slice units. Data purification and information enhancement operations are performed on each data slice unit, marking core information nodes and abnormal information nodes. All marked data slice units are input into a diagnostic analysis engine. This engine reconstructs a state vector representing the operational status of the SME based on the core information nodes and generates a risk feature vector based on the abnormal information nodes. The state vector and risk feature vector are cross-referenced and compared. The weight distribution of the state vector and risk feature vector is dynamically adjusted according to a preset diagnostic rule network, generating a comprehensive diagnostic framework with multiple interrelated diagnostic sub-tasks. Computational resources and data streams are allocated to each diagnostic sub-task for concurrent execution. It receives intermediate diagnostic results from various concurrently executed diagnostic subtasks in real time, performs consistency verification and conflict resolution on the intermediate diagnostic results, and synthesizes and assembles the processed intermediate diagnostic results to generate a structured diagnostic report containing multiple diagnostic dimensions and specific diagnostic items.
[0079] In one embodiment of the present invention, see [reference] Figure 2This process identifies the native data formats and storage structures of business operation data, industry trend data, market feedback data, and compliance audit data. For business operation data, time-series transaction records are converted into standard event sequences within a unified spatiotemporal coordinate system. For industry trend data, trend descriptions and numerical indicators are extracted, and the trend descriptions are transformed into numerical trend vectors. For market feedback data, sentiment polarity analysis and topic clustering are performed on unstructured comments and ratings to generate a quantitative feedback rating matrix. For compliance audit data, legal citations and compliance checkpoints are analyzed to construct a structured compliance rule graph. The standard event sequences, numerical trend vectors, quantitative feedback rating matrix, and structured compliance rule graph are aligned and stitched together through a pre-defined data fusion channel to form a preliminary fusion dataset.
[0080] From the initial fusion dataset, all business domain labels relevant to SME operations are identified. Based on these labels, the initial fusion dataset is initially divided into multiple business domain data blocks. For each business domain data block, the continuous operational timeline is divided into multiple discrete time windows according to a preset time granularity. Within each time window, the data is classified into sensitivity levels based on the category and density of personally identifiable information, trade secret information, and financially sensitive information contained within the data. Based on the sensitivity classification results, the data within each business domain data block in each time window is further segmented into independent data sub-blocks with different sensitivity levels. Each independent data sub-block is encapsulated as a data slice unit, and each data slice unit is attached with metadata labels containing the business domain, time window, and sensitivity level.
[0081] In practice, multi-source data from a small or medium-sized enterprise (SME) engaged in retail business was acquired and processed to form a preliminary fused dataset, which was then structurally segmented. Business operation data included internal sales records and inventory change logs; industry trend data came from publicly available industry analysis reports and market statistics briefings; market feedback data covered customer text reviews and star ratings from major e-commerce platforms; and compliance audit data included tax declaration records and annual business license documents. The original data formats and storage structures of these data were identified: sales records were timestamped transaction tables in a relational database; industry analysis reports were PDF documents; customer text reviews were JSON-formatted API return data; and tax declaration records were structured XML files.
[0082] For time-series transaction records in business operation data, each record, such as "2023-10-01 10:05:30, Store A, Transaction No. 1001, Sales Amount 500 Yuan", is converted into a standard event sequence under a unified spatiotemporal coordinate system. The converted form is "Event (Timestamp: 2023-10-01T10:05:30+08:00, Spatial Coordinates: Store A Geographic Code, Event Type: Sales Transaction, Attribute Set: {Transaction Identifier: 1001, Amount: 500, Currency Unit: CNY})". For industry trend data, trend description text within the report is extracted, such as "Consumer willingness in the fourth quarter showed a moderate recovery trend", and numerical indicators, such as "The industry average growth rate is expected to be 5.2%". The description text "moderate recovery trend" is transformed into a numerical trend vector through a pre-trained natural language processing model. This vector can be represented as [0.7, 0.2, -0.1], where each dimension represents the strength of growth, stability, and decline, respectively. Based on market feedback data, sentiment polarity analysis was performed on unstructured reviews such as "The product quality is very good, but the logistics speed is slow," resulting in positive sentiment scores for "product quality" and negative sentiment scores for "logistics speed." Simultaneously, the massive amount of reviews was clustered into themes such as "product quality," "service attitude," and "logistics experience," ultimately generating a quantitative feedback rating matrix. The rows of the matrix represent theme clusters, the columns represent time windows, and the values in the cells represent the average sentiment score of that theme during that time period.
[0083] The standard event sequences, numerical trend vectors, quantified feedback scoring matrices, and structured compliance rule graphs obtained after the above transformations are aligned and stitched together through a pre-configured data fusion channel. This pre-configured data fusion channel is a pre-configured data integration mechanism used to collaboratively align and structurally stitch together multi-source heterogeneous data after standardization processing. The data fusion channel aligns the four types of data along the time dimension based on a unified time benchmark and associates them along the entity dimension based on the enterprise's unique identifier, ultimately forming a primary fused dataset containing multi-dimensional fields. Each record in this dataset is associated with a time point, an enterprise entity, and fused features from different sources.
[0084] For each business area data block, the continuous operational timeline is divided into multiple discrete time windows according to a preset "monthly" time granularity. Within each time window, the data is classified into sensitivity levels based on the category and density of personally identifiable information, trade secret information, and financially sensitive information contained within the data. For example, the data in the "Sales" data block within the "November 2023" time window might contain detailed customer personal information and discount strategies, while the data in the "Finance" data block within the same time window might contain detailed profit and loss statements. The sensitivity classification operation is calculated using an evaluation function; optionally, an example of the evaluation function is as follows:
[0085] L s =α·ρ pii +β·ρ trade +γ·ρ finance
[0086] Where: L s ρ represents the sensitivity level value of this data block. pii The density of personally identifiable information, ρ trade ρ represents the density of trade secret information. finance The density represents financially sensitive information, and α, β, and γ are the weight coefficients for the corresponding information categories. Based on the calculated L... s The value is used to classify data blocks into sensitivity levels such as "public", "internal", and "confidential".
[0087] In one embodiment of the present invention, missing value detection is performed on data slice units, and based on the business domain characteristics indicated by the metadata tags of the data slice units, an appropriate data imputation strategy is selected to impute the missing values. Noise detection is then performed on the imputed data slice units, and a denoising algorithm matching the sensitivity level of the data slice units is used to filter out random noise. Information density assessment is performed on the denoised data slice units, and regions with information entropy exceeding a preset threshold are identified as high information density regions. Within the high information density regions, frequently occurring patterns, numerical points significantly deviating from historical averages, and records conforming to preset key event characteristics are extracted, and the extraction results are marked as core information nodes. In other regions outside the high information density regions, the presence of data points violating business logic, segments with continuous and rapid fluctuations in the numerical sequence, or records matching items explicitly prohibited in the compliance rule map is detected, and the detection results are marked as abnormal information nodes.
[0088] In practice, data cleansing and information enhancement operations are performed on the encapsulated data slice units, and information nodes are marked. Taking a data slice unit with metadata tags of {Business Domain: "Sales", Time Window: "2023-11", Sensitivity Level: "Confidential"} as an example, this data slice unit contains daily customer transaction details for November 2023. Missing value detection is performed on the data slice unit. The system scans fields such as "Transaction Amount" and "Customer Age Group," identifying multiple consecutive null values in the "Customer Age Group" field in the data from November 15th. Based on the business domain characteristic indicated by the metadata tag of the data slice unit being "Sales," and considering that sales data typically exhibits weekly fluctuations, a data imputation strategy based on temporal proximity and periodic similarity is chosen. Specifically, the age distribution of the corresponding customer groups on November 8th and November 22nd is used to probability-sample and impute the missing values from November 15th.
[0089] Noise detection was performed on the imputed data slices. Several extremely high-value records that significantly deviated from the normal daily consumption range but were not marked as returns were detected in the "single transaction amount" sequence. A denoising algorithm matching the "confidential" sensitivity level of the data slice was used. For "confidential" data involving personal transaction details, noise addition techniques meeting differential privacy requirements were used to perturb the original high values instead of directly deleting them, thus smoothing out abnormal fluctuations while protecting individual privacy. Optionally, for data slices with a sensitivity level of "internal," moving average filtering or median filtering algorithms were used to directly filter out random noise. Information density was evaluated on the denoised data slices. The information entropy of different data regions was calculated, and regions with information entropy exceeding a preset threshold were identified as high information density regions. The preset threshold is a quantitative reference value pre-set during the information density evaluation process.
[0090] In high information density areas, frequently occurring patterns, numerical points significantly deviating from historical averages, and records matching preset key event characteristics are extracted. For example, in the "weekend nighttime" area, the frequent pattern of "snack and beverage purchase frequency is 3 times that of weekday evenings" is extracted; in the "last three days of the month" area, the significant deviation point of "customer C001's single-day consumption amount reaches 5 times its monthly average consumption amount" is extracted; and records matching the preset key event characteristic of "single transaction amount greater than 5,000 yuan and using a corporate payment account" are extracted. These extracted results are marked as core information nodes, with each core information node accompanied by its spatiotemporal context and quantitative description. In other areas outside the high information density areas identified by the information density assessment, the presence of data points that violate business logic, segments with continuous and rapid fluctuations in numerical sequences, or records matching items explicitly prohibited in the compliance rule map is detected. For example, a record showing "transaction time earlier than store opening time" is detected, which is a data point that violates business logic; a segment showing "sales volume of 'stationery' category fluctuated drastically between zero and peak values multiple times between 10:00 AM and 10:30 AM" is detected, which is a continuous and rapidly fluctuating segment; a record showing "sales of restricted goods to minors" is detected, which matches the entry "prohibition of selling alcohol and tobacco products to minors" in the compliance rule map. These detection results are marked as anomalous information nodes, with each anomalous information node accompanied by its anomalous type and the specific rule or pattern violated.
[0091] In some embodiments, information density assessment is performed through a quantization calculation process. Optionally, for a local data region containing discrete data records, its information density assessment value is calculated using the following formula:
[0092]
[0093] Wherein: H I This represents the information density assessment value of this local data region, where n represents the number of unique record types within the region, and p... i Let represent the probability of the i-th record type appearing in this region, and b be the base of the logarithm. The calculated H... I The value is compared with the preset threshold for that region. If H I If the value exceeds a preset threshold, the region is determined to be a high information density region. It can be understood that by performing the purification, enhancement, and marking operations of the above system on each data slice unit, the original numerical sequence is transformed into a series of core information nodes and abnormal information nodes with clear semantic directions.
[0094] In one embodiment of the present invention, see [reference] Figure 3The process involves collecting core information nodes from all data slice units and grouping them according to their respective business domains and time windows. Within each business domain and time window combination, the semantic and numerical relationships between core information nodes are analyzed, aggregating closely related core information nodes into a state feature cluster. Representative feature values, including average, peak, rate of change, and frequency, are extracted from each state feature cluster. The feature values of all state feature clusters under different business domains and time windows are arranged and filled according to a preset global state vector template to form a state vector. The global state vector template defines the position, dimensions, and normalization method of each feature value in the vector.
[0095] The construction steps of the diagnostic analysis engine include: acquiring a historical diagnostic case library, which contains complete operational datasets of multiple historical SMEs and their final confirmed diagnostic conclusions; extracting the occurrence patterns of core information nodes and abnormal information nodes from the historical diagnostic case library and their correspondence with the diagnostic conclusions to form a diagnostic knowledge graph; building a neural network model with multi-layer perception capabilities based on the diagnostic knowledge graph, with the input layer of the neural network model designed to receive the dimensions of the state vector and risk feature vector; embedding an attention mechanism in the hidden layers of the neural network model to automatically learn the influence weights of different state features in the state vector and different risk dimensions in the risk feature vector on the diagnostic conclusions during training; supervising the training of the neural network model using the historical diagnostic case library until the consistency between the model's output diagnostic conclusions and historical real conclusions reaches a preset threshold; and coupling the trained neural network model with the diagnostic knowledge graph to encapsulate it as a diagnostic analysis engine.
[0096] Anomaly information nodes are collected from all data slice units and categorized into data quality anomalies, business logic anomalies, compliance anomalies, and trend anomalies. For each category, the anomaly strength is assessed based on the degree to which the anomaly deviates from the normal baseline and its spatiotemporal clustering density. The propagation risk of each category is also assessed, based on its upstream and downstream connections in the business process and its co-occurrence with other anomaly categories. Each anomaly category and its corresponding anomaly strength and propagation risk assessment values are mapped to a multi-dimensional risk sub-vector. All risk sub-vectors for all anomaly categories are sequentially concatenated according to a predefined risk vector structure to generate a risk feature vector.
[0097] In practice, core information nodes from all data slice units are collected and grouped according to their business domain and time window. For example, all core information nodes with the business domain "Sales" and the time window "2023-11" are grouped together. This group might include nodes such as "Snack and beverage purchase frequency on weekend nights is 3 times that of weekday nights," "Customer C001's daily spending is 5 times the monthly average," and "Multiple corporate payment transactions exceeding 5,000 yuan." Within the "Sales" and "2023-11" group, the semantic and numerical relationships between core information nodes are analyzed. Both "Customer C001's daily spending is 5 times the monthly average" and "Multiple corporate payment transactions exceeding 5,000 yuan" semantically point to large-value transactions and numerically are significantly higher than daily transaction levels. These two closely related core information nodes are aggregated into a state feature cluster called "Monthly Large-Value Transaction Activity." Representative feature values are extracted from the state feature cluster of "Monthly Large Transaction Activity". These feature values include the average transaction amount, peak transaction amount, rate of change of transaction amount compared to the same period of the previous month, and frequency of large transactions within the cluster. The feature values of all state feature clusters under different business areas and time windows are arranged and filled according to a preset global state vector template to form a state vector. For example, the global state vector template predefines positions 1 to 5 to store the average, peak, rate of change, and frequency of "Sales - 2023-11-Monthly Large Transaction Activity" and the feature value of "Inventory - 2023-11-Turnover Efficiency". Each feature value has a specified dimension and a unified normalization method. The state vector is a multi-dimensional array formed by filling all feature values into the corresponding positions according to this template.
[0098] The construction steps of the diagnostic analysis engine include: acquiring a historical diagnostic case library, which contains complete operational datasets of multiple historical SMEs and their final expert-confirmed diagnostic conclusions; extracting the occurrence patterns of core information nodes and abnormal information nodes from the historical diagnostic case library and their correspondence with the diagnostic conclusions, for example, extracting "when the core information node 'accounts receivable turnover rate' shows a continuous downward trend accompanied by the abnormal information node 'surge in bad debt provision,' the diagnostic conclusion usually includes 'high liquidity risk'"; and forming a diagnostic knowledge graph depicting the complex relationship between node patterns and diagnostic conclusions based on a large number of such correspondences. Based on the diagnostic knowledge graph, a neural network model with multi-layer perception capability is constructed. The input layer of the neural network model is designed to receive the total number of dimensions of the state vector and risk feature vector. An attention mechanism is embedded in the hidden layer of the neural network model. During training, the attention mechanism automatically learns the influence weights of different state features such as "monthly large transaction activity" in the state vector and different risk dimensions such as "compliance risk" in the risk feature vector on the final diagnostic conclusion. The neural network model is trained under supervision using a historical diagnostic case library. The model parameters are continuously adjusted during training until the diagnostic conclusions output by the model based on the input state vector and risk feature vector achieve a predetermined threshold of consistency with historical real-world conclusions. The trained neural network model is then coupled with a diagnostic knowledge graph and encapsulated as a diagnostic analysis engine. This engine is capable of receiving node information, reconstructing vectors, and performing inference analysis.
[0099] Anomaly information nodes are collected from all data slice units and categorized. For example, nodes showing "transaction time earlier than store opening time" are categorized as business logic anomalies; nodes showing "sales of restricted goods to minors" are categorized as compliance anomalies; nodes showing "sales volume fluctuating drastically between zero and peak values" are categorized as data quality anomalies; and nodes showing "market share of core product lines declining continuously for the past three months" are categorized as trend anomalies. For each category of anomaly information nodes, their anomaly strength is assessed. Anomaly strength is calculated based on the degree to which the anomaly information node deviates from the normal baseline and its spatiotemporal clustering density. Optionally, for business logic anomalies, the degree of deviation is the absolute value of the time deviation, and the clustering density is the number of times such anomalies occur per unit time. For each type of abnormal information node, assess its transmission risk. The transmission risk is estimated based on the upstream and downstream connections of the abnormal information node in the business process and its co-occurrence relationship with other abnormal categories. For example, the compliance abnormal node "selling restricted goods to minors" may be associated with the risk of "facing regulatory penalties" downstream. Moreover, if it co-occurs with the data quality abnormal node "sales record tampering", the transmission risk will increase.
[0100] Each anomaly category and its corresponding anomaly intensity assessment value and transmission risk assessment value are mapped to a multi-dimensional risk sub-vector. For example, the risk sub-vector for a compliance anomaly category is represented as [category code, anomaly intensity value λ, transmission risk value φ], where λ and φ are calculated scalar values. All anomaly category risk sub-vectors are sequentially concatenated according to a preset risk vector structure. For example, the structure specifies that the first part of the risk feature vector is a data quality anomaly sub-vector, the second part is a business logic anomaly sub-vector, and so on, combining to generate the final risk feature vector. In some embodiments, anomaly intensity assessment can be implemented through a calculation process. It can be understood that, for example, for the anomaly intensity λ of a certain anomaly category within a specific time window, its calculation can consider the magnitude of the anomaly value's deviation from the baseline and the degree of concentration of its distribution. Optionally, an exemplary calculation formula is as follows:
[0101]
[0102] Where: λ represents the anomaly intensity assessment value of this category within the specified spatiotemporal range, N represents the number of anomaly information nodes of this category within the range, and v k D represents the key metric value of the k-th anomaly information node, where μ and σ represent the mean and standard deviation of this metric value under the normal benchmark, respectively. c This represents the density measure of the aggregation of these anomalous nodes in the spatiotemporal coordinates.
[0103] In one embodiment of the present invention, a correlation mapping table is established between each state feature in the state vector and each risk dimension in the risk feature vector. This correlation mapping table describes the probability that a state change may induce or exacerbate a specific risk. Based on the correlation mapping table, the potential influence of the current operating state represented by the state vector on each risk dimension in the risk feature vector is calculated. Simultaneously, the potential erosion effect of existing risks represented by the risk feature vector on each state feature in the state vector is calculated. A matrix operation is performed on the potential influence and potential erosion effect to generate a correlation strength matrix between state and risk. The correlation strength matrix is analyzed to identify state-risk combinations whose correlation strength exceeds a warning threshold. These state-risk combinations are then identified as key interaction points requiring focus in the diagnostic analysis.
[0104] The pre-defined diagnostic rule network contains multiple diagnostic rule nodes, each corresponding to a diagnostic knowledge or experience logic. The state vector, risk feature vector, and correlation strength matrix are input into the pre-defined diagnostic rule network. Each diagnostic rule node, based on its internal logic, judges the input state, risk, and correlation strength, and outputs weight adjustment suggestions for specific state features in the state vector and specific risk dimensions in the risk feature vector. The weight adjustment suggestions output by all diagnostic rule nodes are collected, and a weight arbitration mechanism coordinates and merges potentially conflicting suggestions. Based on the coordinated and merged weight adjustment suggestions, the weight coefficients of each state feature in the state vector and the weight coefficients of each risk dimension in the risk feature vector are updated in real time.
[0105] In practical implementation, the state vector and risk feature vector are cross-referenced and compared to establish a correlation mapping table between each state feature in the state vector and each risk dimension in the risk feature vector. This mapping table describes the probability that a change in state may induce or exacerbate a specific risk. Taking a small and medium-sized retail enterprise as an example, the state vector includes state features such as "monthly large transaction activity," "inventory turnover efficiency," and "customer satisfaction index," while the risk feature vector includes risk dimensions such as "compliance risk," "liquidity risk," and "market reputation risk." Through analysis of historical data and expert rules, the correlation mapping table quantifies that, for example, an abnormal increase in "monthly large transaction activity" may induce "compliance risk" with a probability of 0.85, while a decrease in "inventory turnover efficiency" may exacerbate "liquidity risk" with a probability of 0.72. See Table 1 for a simplified correlation mapping table.
[0106] Table 1: Mapping Table of Relationship between State Characteristics and Risk Dimensions
[0107] State characteristics Risk Dimensions Association probability Monthly large transaction activity Compliance risks 0.85 Inventory turnover efficiency Liquidity risk 0.72 Customer Satisfaction Index Market reputation risk 0.68 Marketing expenses as a percentage Operational efficiency risks 0.60
[0108] Based on the association mapping table, the potential influence of the current operational state represented by the state vector on each risk dimension in the risk feature vector is calculated. Simultaneously, the potential erosion effect of existing risks represented by the risk feature vector on each state feature in the state vector is calculated. For example, the potential influence of the current "Monthly Large Transaction Activity" feature value of 0.9 (after normalization) on the "Compliance Risk" dimension is calculated, and the potential erosion effect of the current "Compliance Risk" dimension value of 0.8 on the "Monthly Large Transaction Activity" state feature is calculated. A matrix operation is performed on the potential influence and potential erosion effect to generate a state-risk association strength matrix. Analyzing the association strength matrix identifies state-risk combinations with association strength exceeding the warning threshold. These state-risk combinations are identified as key interaction points requiring focus in diagnostic analysis. For example, the association strength value of the combination of "Monthly Large Transaction Activity" and "Compliance Risk" is identified as 0.93, exceeding the preset warning threshold of 0.9.
[0109] In some embodiments, the pre-defined diagnostic rule network includes multiple diagnostic rule nodes, each corresponding to a diagnostic knowledge or experience logic. For example, a diagnostic rule node might encode the rule "If the compliance risk is high and large transactions are cash transactions, then the anti-money laundering risk weight should be increased." The state vector, risk feature vector, and correlation strength matrix are input into the pre-defined diagnostic rule network. Each diagnostic rule node, based on its internal logic, judges the input state, risk, and correlation strength, and outputs weight adjustment suggestions for specific state features in the state vector and specific risk dimensions in the risk feature vector. For example, the aforementioned node might output suggestions such as "increase the weight of the compliance risk dimension by 0.1" and "increase the weight of the large transaction-related features in the state vector by 0.05." All weight adjustment suggestions output by the diagnostic rule nodes are aggregated, and a weight arbitration mechanism coordinates and merges potentially conflicting adjustment suggestions. The weight arbitration mechanism makes decisions based on factors such as node priority and the historical validity of the suggestions. Based on the weight adjustment suggestions after coordination and integration, the weight coefficients of each state feature in the state vector and the weight coefficients of each risk dimension in the risk feature vector are updated in real time. The updated weight coefficients will directly affect the attention paid to each feature by subsequent diagnostic subtasks.
[0110] Optionally, the potential influence Γ of the i-th state feature in the state vector on the j-th risk dimension in the risk feature vector. ij It can be calculated using the following formula:
[0111] Γ ij =P ij ·V i ·W i
[0112] Wherein: Γ ijP represents the potential influence of state feature i on risk dimension j. ij V represents the association probability between state feature i obtained from the association mapping table and risk dimension j. i W represents the current normalized eigenvalue of state feature i. i This represents the current weight coefficient of state feature i in the state vector. It can be understood that by introducing the association mapping table and weight coefficients, cross-association comparison not only considers the inherent relationship between state and risk, but also incorporates dynamic importance assessment based on the current context. This enables the diagnostic analysis engine to more accurately focus on the most influential risk-state interactions and provides a quantitative basis for the dynamic weight adjustment of the preset diagnostic rule network.
[0113] See Figure 4 This is a double-bar chart comparing the weight adjustments of risk dimensions. It shows the changes in the weight coefficients of different risk dimensions in SME diagnosis between "before adjustment" and "after adjustment," and is a core analysis chart in the pre-set diagnostic rule network stage. The adjusted weights of all risk dimensions are not lower than before adjustment, with some dimensions (such as "compliance risk" and "supply chain risk") showing significant weight increases. The adjusted weight of "liquidity risk" reaches approximately 0.20, making it the highest priority risk dimension. This type of chart serves the weight optimization scenario in SME diagnosis, demonstrating through weight changes that rule adjustments have increased the priority of high-risk dimensions. High-weight dimensions such as "liquidity risk" and "compliance risk" are the core objects of subsequent diagnostic analysis.
[0114] In one embodiment of the invention, a shared intermediate diagnostic result storage area is established to receive and temporarily store intermediate diagnostic results output by all diagnostic subtasks. Each intermediate diagnostic result is appended with the identifier of the source diagnostic subtask, a timestamp of its generation, and the version information of the input data it is based on. A consistency checker is set up to continuously scan the intermediate diagnostic result storage area, checking whether intermediate diagnostic results from different diagnostic subtasks for the same diagnostic entry contain numerical contradictions, logical conflicts, or contradictory conclusions. When a conflict is detected, the consistency checker sorts the conflicting intermediate diagnostic results by their confidence level based on the preset confidence level of the source diagnostic subtask, the time of generation, and the completeness of the input data. Based on the confidence ranking, an optimal intermediate diagnostic result is selected, or a negotiation protocol between diagnostic subtasks is triggered to generate a negotiated consensus result to resolve the conflict.
[0115] In practical implementation, a shared intermediate diagnostic result storage area is established to receive and temporarily store the intermediate diagnostic results output by all diagnostic subtasks. Logically, this intermediate diagnostic result storage area can be a distributed database table with version management functionality. Each intermediate diagnostic result is appended with the identifier of the source diagnostic subtask, the timestamp of its generation, and the version information of the input data it is based on. For example, an intermediate diagnostic result regarding "cash flow risk level" would have the appended information as {Source: Cash Flow Risk Diagnostic Subtask, Timestamp: 2023-12-01 14:30:25.123, Data Version: State Vector v2.1_Risk Feature Vector v1.7}. Another intermediate diagnostic result regarding the same diagnostic item might have the appended information as {Source: Market Risk Diagnostic Subtask, Timestamp: 2023-12-01 14:30:25.456, Data Version: State Vector v2.1_Risk Feature Vector v1.7}.
[0116] A consistency checker is set up to continuously scan the intermediate storage area of diagnostic results, checking whether there are numerical contradictions, logical conflicts, or contradictory conclusions in the intermediate diagnostic results from different diagnostic subtasks for the same diagnostic item. For example, the consistency checker finds that the cash flow risk diagnostic subtask outputs an intermediate diagnostic result of "high risk" for "capital turnover risk level," while the market risk diagnostic subtask outputs an intermediate diagnostic result of "medium risk" for the same "capital turnover risk level." These two results contradict each other in terms of numerical rating. When a conflict is detected, the consistency checker sorts the conflicting intermediate diagnostic results by their credibility based on the preset credibility level of the source diagnostic subtask, the time of generation, and the completeness of the input data. The credibility level of the source diagnostic subtask is set based on historical accuracy during system initialization. For example, the cash flow risk diagnostic subtask has a "high" credibility level because of its high model maturity, while the market risk diagnostic subtask has a "medium" credibility level.
[0117] In some embodiments, the credibility ranking is achieved through a quantification calculation process, where the consistency checker calculates a comprehensive credibility score for each intermediate diagnostic result involved in conflict resolution. Optionally, the comprehensive credibility score S... c The calculation formula is as follows:
[0118] S c =R·τ(t)·η(d)
[0119] Wherein: S cThe overall credibility score represents the intermediate diagnostic result. R represents the numerical weight corresponding to the preset credibility level of the source diagnostic subtask. τ(t) is a time decay function based on the timestamp t, ensuring that the updated result receives a higher timeliness weight when the input data versions are the same. η(d) is an integrity factor function based on the percentage of completeness d of the input data; a higher η(d) value indicates higher input data completeness. The overall credibility score S is calculated as follows: c The intermediate diagnostic results of the conflict are sorted in descending order.
[0120] Based on the reliability ranking, select the optimal intermediate diagnostic result, or trigger a negotiation protocol between diagnostic subtasks to generate a negotiated and consistent result. If the intermediate diagnostic result ranked first has a comprehensive reliability score S... c If the intermediate diagnostic result is significantly higher than the second-ranked result, for example, exceeding a preset threshold difference, then the intermediate diagnostic result ranked first is directly selected as the valid result for that diagnostic item. If the overall reliability score S of the top few intermediate diagnostic results is... c If the difference is small and does not reach the threshold for direct selection, a negotiation agreement between diagnostic subtasks is triggered. The negotiation agreement requires the relevant diagnostic subtasks to exchange some intermediate inference data or reassess shared key features. For example, the cash flow risk diagnostic subtask and the market risk diagnostic subtask may verify the calculation process of the shared feature "accounts receivable turnover ratio" and adjust their respective diagnostic logics based on the verification results, ultimately outputting a consistent result agreed upon by both parties, such as negotiating and determining the "capital turnover risk level" as "medium-high risk".
[0121] See Figure 5 This is a comparison chart of the serial and parallel execution of diagnostic subtasks, used to illustrate the time differences between "serial execution" and "parallel execution" modes for different subtasks in SME diagnostics. It is a core chart for diagnostic task scheduling efficiency analysis. The "parallel" execution time for all subtasks is significantly longer than the "serial" execution time. Among them, the serial execution time for "Risk Diagnosis" is close to 9 seconds, making it the longest-running subtask; the serial execution time for "Operational Diagnosis" is only about 1.7 seconds, making it the most efficient subtask. This type of chart serves diagnostic task scheduling optimization scenarios, clarifying the advantages of the serial mode in terms of single-task time, while the parallel mode is generally suitable for scenarios where multiple tasks are processed simultaneously; long-running subtasks such as "Risk Diagnosis" can be considered as key areas for subsequent performance optimization.
[0122] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0123] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A diagnostic method for SMEs based on multi-source data fusion and AI, characterized in that: include: Acquire the business operation data, industry trend data, market feedback data, and compliance audit data required for the diagnosis of SMEs, and perform format assimilation processing to form a preliminary fusion dataset; The primary fusion dataset is segmented according to business domain, time period, and data sensitivity to obtain multiple data slice units. Data purification and information enhancement operations are performed on each data slice unit to mark core information nodes and abnormal information nodes. All marked data slices are input into the diagnostic analysis engine. Based on the core information node, the diagnostic analysis engine aggregates related nodes to form feature clusters, extracts representative feature values of the feature clusters and arranges and fills them to obtain a state vector, and evaluates the abnormal intensity and transmission risk of various abnormal information nodes based on the abnormal information nodes, and generates a risk feature vector by combining them according to a preset risk vector structure. The state vector and the risk feature vector are cross-correlated and compared. The weight distribution of the state vector and the risk feature vector is dynamically adjusted according to the preset diagnostic rule network to generate a comprehensive diagnostic framework with multiple interrelated diagnostic subtasks. Computational resources and data streams are allocated to each diagnostic subtask to start concurrent execution. The system receives intermediate diagnostic results from various concurrently executed diagnostic subtasks in real time, performs consistency verification and conflict resolution on the intermediate diagnostic results, and synthesizes and assembles the processed intermediate diagnostic results to generate a structured diagnostic report containing multiple diagnostic dimensions and specific diagnostic items.
2. The SME diagnostic method based on multi-source data fusion and AI according to claim 1, characterized in that, The process of format assimilation to form a primary fused dataset includes: Identify the native data format and storage structure corresponding to the business operation data, the industry trend data, the market feedback data, and the compliance audit data; For the aforementioned business operation data, the time-series transaction records are converted into standard event sequences under a unified spatiotemporal coordinate system; For the industry trend data, extract the trend description text and numerical indicators, and convert the trend description text into a numerical trend vector. Based on the market feedback data, sentiment polarity analysis and topic clustering are performed on the unstructured comments and rating information to generate a quantitative feedback rating matrix; For the aforementioned compliance audit data, we analyze its internal legal references and compliance checkpoints to construct a structured compliance rule map; The standard event sequence, the numerical trend vector, the quantified feedback scoring matrix, and the structured compliance rule graph are aligned and stitched together through a preset data fusion channel to form the primary fusion dataset.
3. The SME diagnostic method based on multi-source data fusion and AI according to claim 2, characterized in that, The initial fused dataset is segmented based on business domain, time period, and data sensitivity dimensions to obtain multiple data slice units, including: Identify all business area labels related to SME operations from the primary fusion dataset; Based on the business domain labels, the initial fusion dataset is preliminarily divided into multiple business domain data blocks; For each business area data block, the continuous operational timeline is divided into multiple discrete time windows according to a preset time granularity; Within each time window, the data is classified into sensitivity levels based on the categories and density of personally identifiable information, trade secret information, and financially sensitive information contained in the data; Based on the sensitivity classification results, the data of each business domain data block within each time window is further divided into independent data sub-blocks with different sensitivity levels; Each independent data sub-block is encapsulated into a data slice unit, and each data slice unit is attached with metadata tags containing business domain, time window, and sensitivity level.
4. The SME diagnostic method based on multi-source data fusion and AI according to claim 3, characterized in that, The step of performing data cleansing and information enhancement operations on each data slice unit, and marking core information nodes and abnormal information nodes, includes: Missing values are detected in the data slice unit, and based on the business domain characteristics indicated by the metadata tags of the data slice unit, an appropriate data imputation strategy is selected to impute the missing values. Noise detection is performed on the filled data slice units, and random noise is filtered out using a noise reduction algorithm that matches the sensitivity level of the data slice units. The information density of the data slices after noise reduction is evaluated, and regions whose information entropy exceeds a preset threshold are identified as high information density regions. The preset threshold is a quantitative reference value pre-set during the information density assessment process; In the high information density area, frequently occurring patterns, numerical points that significantly deviate from the historical average, and records that conform to the preset key event characteristics are extracted, and the extraction results are marked as the core information nodes; In areas other than the high information density region, detect whether there are data points that violate business logic, segments that fluctuate continuously and rapidly in the numerical sequence, or records that match entries that are explicitly prohibited in the compliance rule map, and mark the detection results as the abnormal information nodes.
5. The SME diagnostic method based on multi-source data fusion and AI according to claim 4, characterized in that, The diagnostic analysis engine, based on the core information nodes, aggregates related nodes to form feature clusters, extracts representative feature values from the feature clusters, and arranges and fills them to obtain a state vector, including: Collect the core information nodes from all data slice units and group them according to the business domain and time window to which the core information nodes belong; Within each combination of business domain and time window, the semantic and numerical relationships between the core information nodes are analyzed, and multiple closely related core information nodes are aggregated into a state feature cluster. Representative feature values are extracted from each state feature cluster, including but not limited to: average value, peak value, rate of change, and frequency; The feature values of all state feature clusters under different business domains and different time windows are arranged and filled according to a preset global state vector template to form the state vector. The global state vector template defines the position, dimension and normalization method of each feature value in the vector. The steps for building the diagnostic analysis engine include: Obtain a historical diagnostic case library, which contains complete operational datasets of multiple historical small and medium-sized enterprises and their final confirmed diagnostic conclusions; Extract the occurrence patterns of the core information nodes and the abnormal information nodes from the historical diagnostic case database, and their correspondence with the diagnostic conclusions to form a diagnostic knowledge graph; Based on the diagnostic knowledge graph, a neural network model with multi-layer perception capability is constructed. The input layer of the neural network model is designed to receive the dimension of the state vector and the risk feature vector. An attention mechanism is embedded in the hidden layer of the neural network model to automatically learn the influence weights of different state features in the state vector and different risk dimensions in the risk feature vector on the diagnostic conclusion during training. The neural network model is trained under supervision using the historical diagnostic case library until the consistency between the diagnostic conclusions output by the model and the historical real conclusions reaches a preset threshold. The trained neural network model is coupled with the diagnostic knowledge graph and encapsulated as the diagnostic analysis engine.
6. The SME diagnostic method based on multi-source data fusion and AI according to claim 4, characterized in that, The process involves evaluating the anomaly intensity and transmission risk of various anomaly information nodes based on the anomaly information nodes, and generating a risk feature vector by combining them according to a preset risk vector structure, including: Collect the abnormal information nodes from all data slice units, and classify the abnormal information nodes into categories including data quality abnormalities, business logic abnormalities, compliance abnormalities, and trend abnormalities. For each type of anomalous information node, its anomalous strength is evaluated. The anomalous strength is calculated based on the degree to which the anomalous information node deviates from the normal baseline and its spatiotemporal aggregation density. For each type of abnormal information node, its transmission risk is assessed. The transmission risk is estimated based on the upstream and downstream associations of the abnormal information node in the business process and its co-occurrence relationship with other abnormal categories. Each anomaly category and its corresponding anomaly intensity assessment value and transmission risk assessment value are mapped to a multidimensional risk sub-vector; All risk subvectors of anomaly categories are sequentially connected according to a preset risk vector structure to generate the risk feature vector.
7. The SME diagnostic method based on multi-source data fusion and AI according to claim 6, characterized in that, The cross-correlation comparison of the state vector and the risk feature vector includes: Establish an association mapping table between each state feature in the state vector and each risk dimension in the risk feature vector. The association mapping table describes the probability that a state change may induce or aggravate a specific risk. Based on the association mapping table, calculate the potential influence of the current operating state represented by the state vector on each risk dimension in the risk feature vector; Simultaneously, the potential erosion effect of the existing risks represented by the risk feature vector on each state feature in the state vector is calculated; Perform matrix operations on the potential influence and the potential erosion effect to generate a correlation strength matrix between state and risk; By analyzing the correlation strength matrix, we can identify the state and risk combinations whose correlation strength exceeds the warning threshold, and use these state and risk combinations as key interaction points to focus on in the diagnostic analysis.
8. The SME diagnostic method based on multi-source data fusion and AI according to claim 7, characterized in that, The step of dynamically adjusting the weight distribution of the state vector and the risk feature vector according to a preset diagnostic rule network includes: The preset diagnostic rule network contains multiple diagnostic rule nodes, each corresponding to a specific diagnostic knowledge or experience logic; The state vector, the risk feature vector, and the correlation strength matrix are input into the preset diagnostic rule network; Each diagnostic rule node, based on its internal logic, judges the input state, risk, and correlation strength, and outputs weight adjustment suggestions for specific state features in the state vector and specific risk dimensions in the risk feature vector. The system aggregates weight adjustment suggestions from all diagnostic rule nodes and coordinates and integrates potentially conflicting adjustment suggestions through a weight arbitration mechanism. Based on the weight adjustment suggestions after coordination and fusion, the weight coefficients of each state feature in the state vector and the weight coefficients of each risk dimension in the risk feature vector are updated in real time.
9. The SME diagnostic method based on multi-source data fusion and AI according to claim 8, characterized in that, The real-time reception of intermediate diagnostic results from each concurrently executed diagnostic subtask, and the performance of consistency verification and conflict resolution on the intermediate diagnostic results, includes: Establish a shared intermediate diagnostic results storage area to receive and temporarily store the intermediate diagnostic results output by all diagnostic subtasks; For each intermediate diagnostic result, attach the identifier of the source diagnostic subtask, the timestamp of its generation, and the version information of the input data it is based on; Set up a consistency checker, which continuously scans the intermediate storage area of the diagnostic results to check whether there are numerical contradictions, logical conflicts or contradictory conclusions in the intermediate diagnostic results from different diagnostic subtasks for the same diagnostic item. When a conflict is detected, the consistency checker sorts the intermediate diagnostic results of the conflict according to the preset confidence level of the source diagnostic subtask, the time of generation, and the integrity of the input data on which the intermediate diagnostic results are based. Based on the reliability ranking, select the best intermediate diagnostic result, or trigger a negotiation protocol between diagnostic subtasks to generate a negotiated and consistent result to resolve conflicts.
10. A diagnostic system for SMEs based on multi-source data fusion and AI, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the SME diagnostic method based on multi-source data fusion and AI as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Business risk prediction method and device, computer equipment and storage medium
CN114662570A
Enterprise intelligent diagnosis method, system and equipment based on large model and medium
CN120412984A
Multi-dimensional data integration and dynamic risk assessment method for enterprise purchase anomaly detection
CN120725468A
Knowledge graph construction method and system for enterprise dynamic risk
CN121599070A