Computer data intelligent analysis system based on artificial intelligence

By constructing a local window set to identify differential fluctuations in multi-dimensional information and dynamically adjusting semantic category matching, the problem of insufficient adaptability of high-dimensional dynamically changing data in the existing technology is solved, and accurate identification and stable classification of data trajectories are achieved, and the accuracy and consistency of data analysis are improved.

CN120541567APending Publication Date: 2025-08-26JINAN HOTZ INFORMATION TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510605746.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When processing high-dimensional dynamically changing data, the existing computer data intelligent analysis system based on artificial intelligence has weak structural adaptability, insufficient semantic evolution tracking, and rigid label response mechanism, resulting in high sensitivity to abnormal fluctuations in classification results, and lack of dynamic modeling of data trajectory continuity and local perturbation amplitude, making it difficult to adapt to the time-by-time offset of semantic boundaries in business data.

Method used

By constructing a local window set to identify differential fluctuations in multi-dimensional information, iteratively analyze the bias impact, combining attribution correction, tension correction and label frequency analysis, the semantic category matching method is dynamically adjusted, data behavior and classification stability are optimized, the ability to identify and compensate disturbance points is enhanced, and the accuracy of semantic trajectory offset response is improved.

Benefits of technology

It realizes accurate identification and stable classification of high-dimensional dynamically changing data, improves the dynamic modeling ability of data trajectory, enhances the consistency of label expression and the accuracy of semantic matching, and builds a closed-loop linkage between data behavior and classification stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541567A_ABST
    Figure CN120541567A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data mining, in particular to an intelligent computer data analysis system based on artificial intelligence, which comprises a data deviation identification module, an attribution correction module, a tension correction module, a label frequency analysis module and a semantic deviation adjustment module. According to the method, by constructing the local window set and analyzing the difference fluctuation between the dimensions, the key dimension of the continuous deviation feature is accurately recognized, the situation that local anomaly disturbs the re-weighting of the high deviation dimension in classification judgment and affiliation evaluation is avoided, the stability of classification under label missing or affiliation fuzziness is improved, and the classification accuracy is improved. The dynamic modeling of the data track enhances the recognition and compensation capability of disturbance points and improves the classification continuity, the time sequence monitoring and fluctuation adjustment mechanism of the tag frequency enhances the consistency of tag expression, the semantic track offset response optimizes the semantic matching accuracy of category attribution, and the classification accuracy is improved. And constructing closed-loop linkage among data behaviors, classification stability and semantic adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data mining, and in particular to an artificial intelligence-based computer data intelligent analysis system. Background Art

[0002] The field of data mining technology encompasses computational methods and systems that analyze large amounts of structured or unstructured data to extract underlying patterns and useful information. Its core content lies in modeling and analyzing data sets through statistical learning, machine learning, pattern recognition, and other means to identify representative, relevant, or predictive patterns from massive amounts of data. This technical field, encompassing data preprocessing, feature extraction, association rule mining, cluster analysis, classification and recognition, is widely used in fields such as business intelligence, medical diagnosis, financial risk control, and industrial manufacturing. Data mining relies on efficient computing resources and complex mathematical modeling processes, requiring efficient collaboration between algorithm design and data management to effectively summarize and analyze multi-source heterogeneous data.

[0003] Among them, the computer data intelligent analysis system based on artificial intelligence refers to a system device that uses artificial neural networks, deep learning models and adaptive classification mechanisms to perform feature modeling and analysis on various types of business data generated in massive computer systems. The subject of this patent mainly addresses the dimensional complexity and missing data labels problems existing in the data acquisition stage. It extracts high-dimensional features through convolutional neural networks, models time series data characteristics through long and short-term memory networks, and completes target classification and labeling in combination with decision tree structures. During the implementation process, the system uses embedded data fusion coding to unify raw data from different sources into computable inputs, and performs semantic understanding and structural mapping of the data based on the supervised learning model training process, thereby building a data mining logic system that adapts to different business scenarios.

[0004] Existing technologies often rely on static feature modeling and label association rules for classification. These mechanisms lag in responding to label ambiguity and temporal fluctuations, leading to an imbalance between the expressive strength and semantic attribution of local high-frequency labels. When dealing with dimensional perturbations and data tension fluctuations, they lack dynamic modeling of data trajectory continuity and local perturbation amplitude, making it impossible to accurately define perturbation nodes and compensation paths, resulting in overly sensitive classification results to abnormal fluctuations. During the attribution determination phase, static classification boundaries are often used, lacking a mechanism to reassess classification confidence under bias, leading to unstable attribution decisions in the presence of missing labels or a concentration of marginal samples. Semantic category matching methods are often based on fixed rules and lack modeling and response to semantic drift, making them difficult to adapt to the time-to-time shifts in semantic boundaries within business data. For example, in financial risk identification scenarios, the same behavioral label may correspond to different risk levels in different market cycles. Without the ability to dynamically adjust the matching method, the classification model is prone to misjudgment or delayed response, impacting overall decision-making effectiveness. These shortcomings indicate that existing systems have many limitations when dealing with high-dimensional dynamically changing data, such as weak structural adaptability, insufficient semantic evolution tracking, and rigid label response mechanism. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an artificial intelligence-based computer data intelligent analysis system.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: a computer data intelligent analysis system based on artificial intelligence includes: The data bias identification module obtains multi-dimensional information of data entities, calculates dimensional difference fluctuation indicators based on the distribution characteristics within the window, determines the fluctuation direction of multiple windows, iteratively analyzes the potential impact on the classification results, and generates bias impact results; The attribution correction module calls the bias impact result, extracts the bias dimension, re-evaluates and corrects the bias in combination with the classification behavior, and iteratively generates corrected attribution information; The tension correction module constructs a change trajectory of the data point in the multidimensional space based on the corrected attribution information, identifies the tension mutation within the time step, marks it as a disturbance point if it exceeds a threshold, and generates a tension compensation result through compensation correction; The tag frequency analysis module calls the tension compensation result, uses a time sliding window to count the tag frequency, filters out tags with frequency fluctuations, and generates a tag frequency fluctuation result; The semantic shift adjustment module analyzes the changes in the center trajectory of the semantic category within the time period based on the tag frequency fluctuation results, adjusts the matching method according to the trajectory mutation, and generates computer data intelligent analysis results.

[0007] As a further solution of the present invention, the bias impact results include trend consistency identification, difference fluctuation index set, and local dimension bias label; the corrected attribution information includes corrected classification label, attribution confidence parameter, and dimension rescore distribution; the tension compensation results include disturbance point records, tension mutation amplitude set, and compensation correction factor; the label frequency fluctuation results include high-frequency label sequence, time period fluctuation characteristics, and label adjustment priority; the computer data intelligent analysis results include semantic center trajectory set, category offset node, and matching rule adjustment item.

[0008] As a further solution of the present invention, the data bias identification module includes: The local window construction submodule obtains multidimensional information of each data entity based on its structural characteristics and temporal properties. It then constructs multiple data window sets based on fixed window widths and sliding intervals, numbers and identifies the data within each window, and synchronizes the distribution intervals of differentiated data entities within the same dimension to obtain a window distribution alignment indicator. The dimension fluctuation calculation submodule calls the window distribution alignment index to calculate the distribution difference of each dimension in all windows. Based on the numerical changes of the dimensions in multiple windows, it calculates the change difference and average offset rate between adjacent windows, and extracts the change direction vector of each dimension to obtain the dimension fluctuation difference value. The bias trend determination submodule compares and determines the change trend direction of each dimension in the window based on the dimension fluctuation difference value. If the direction is consistent in multiple windows, the dimension is marked as a bias dimension, using the formula: ; Obtain the bias fluctuation measurement value of the dimension by calculation, filter out the dimensions that are greater than the bias judgment threshold, obtain the bias impact index, and generate the bias impact result; in, Representative The biased volatility measure of the dimension, and Respectively represent The dimension in The difference between the positive and negative changes in the window, For the Dimension The change direction vector in the window, is the change amplitude corresponding to the direction, For the The overall mean change of the dimension, is the moving average of all dimensions.

[0009] As a further solution of the present invention, the attribution correction module includes: The dimension screening submodule calls the bias impact indicator, screens dimensions whose fluctuation amplitude is greater than the attribution offset threshold, extracts key dimension information that affects data attribution judgment, and obtains an attribution bias dimension set; The bias resolution submodule determines the difference between the data change trend of the dimension and the category boundary position based on the attribute bias dimension set, and calculates the attribute deviation degree of each dimension to the classification result using the formula: ; The calculation obtains the attribution offset adjustment value of the bias dimension, adjusts and corrects the corresponding values ​​of all bias dimensions, and generates an attribution correction data set; in, Indicates the The attribute offset adjustment value of the bias dimension, For the The dimension in The weight of the category in the sample, is the boundary distance value in the corresponding sample, is the fluctuation interference term, For the Dimension The original classification value of the sample, For the Sample target category value, For the The cumulative value of the disturbance in each dimension, is the trend suppression factor under dimension; The attribution re-judgment submodule re-establishes the distance mapping matrix between samples and categories based on the attribution correction data set, selects and classifies the corrected samples according to their attribution at the center point of multiple categories, and obtains the corrected attribution information.

[0010] As a further solution of the present invention, the tension correction module includes: The trajectory construction submodule extracts the multidimensional state values ​​of the data points in continuous time steps based on the corrected attribution information, calculates the change rate and direction combination in each time step, and splices them in time sequence to form a multidimensional vector path to obtain the dimension change trajectory value; The disturbance identification submodule calculates the magnitude of the dimension change of the data points between consecutive time steps based on the dimension change trajectory value, and compares it with the set tension disturbance threshold to determine whether it constitutes a disturbance state. The formula is: ; The disturbance intensity coefficient of each time step is obtained by calculation. If the intensity coefficient is greater than the tension disturbance threshold, the corresponding position is marked as the disturbance occurrence point, and the disturbance position identification set is obtained; in, For the The perturbation intensity coefficient of the data point, is the reference amplitude of normal tension change, For the The data point is The coefficient of the direction of change of the time step, is the change amplitude value of the point, is the cumulative value of fluctuation at a time point, For the The stability factor of the time step, is the tension projection value of the category to which the data point belongs, The tension effect mapping value corresponding to time; The disturbance compensation submodule extracts the corresponding change path segment and the previous time step value of the disturbance according to the disturbance position identification set, calculates the disturbance direction and amplitude error and performs inverse adjustment, applies the compensation vector to the corresponding time step value, and obtains the tension compensation result.

[0011] As a further solution of the present invention, the tag frequency analysis module includes: The frequency statistics submodule collects the tag data in the time window according to the tension compensation result, calculates the number of occurrences and the frequency ratio of each tag in the window, and constructs the tag frequency sequence in the time dimension to obtain the tag time series frequency value; The fluctuation screening submodule calls the tag time series frequency value, calculates the frequency difference and fluctuation gradient of the tag between adjacent time windows, and identifies high-fluctuation tags using the formula: ; The frequency fluctuation coefficient of each tag is obtained by calculation, and tags with fluctuation coefficients higher than the tag fluctuation determination threshold are screened to obtain a high-fluctuation tag set; in, Representation Label The frequency fluctuation coefficient, Representation Label In the The frequency of the time window, is the slope of frequency change, is the cumulative offset value, is the label distribution density within the window, is the average frequency of the tag over the entire period, For the The median frequency of the time window, For label The normalized stability factor of The label adjustment submodule extracts the frequency sequence of the corresponding time period based on the high-volatility label set, calculates the change trend and direction ratio of the time point, determines the label frequency change tendency, reconstructs the weight coverage range and performs label mapping adjustment to obtain the label frequency fluctuation result.

[0012] As a further solution of the present invention, the semantic offset adjustment module includes: The trajectory extraction submodule extracts the center point coordinates of each semantic category within a continuous time period based on the label frequency fluctuation results, calculates the spatial change path between time steps, constructs a multi-period category trajectory sequence in chronological order, and obtains the category center trajectory value; The offset recognition submodule calculates the center point movement amplitude and direction transfer angle in a continuous time period based on the category center trajectory value, segments the mutation path and sets the offset judgment boundary using the formula: ; Obtain the semantic shift degree value of each category by operation, determine whether a sudden change occurs, and obtain the semantic shift category identification set; in, Indicates the The semantic deviation value of the category, For the Category in The displacement value of the center point of the time period, is the slope of its direction change, is the angle difference between the front and back vectors of the point, is the path disturbance amplitude, is the category aggregation stabilization factor, For the The amount of fluctuation mapping for the time period, is the evolutionary stability of the semantic category itself, For category The offset absorption threshold, For the Deviation of the center point of the time period; The matching correction submodule extracts the frequency path and center point trajectory of the offset category in the corresponding time period based on the semantic offset category identification set, reconstructs the spatial distance matrix between categories, readjusts the category attribution method and marks the mapping position, and obtains computer data intelligent analysis results.

[0013] Compared with the prior art, the advantages and positive effects of the present invention are: In the present invention, by constructing a local window set and analyzing the difference fluctuations between dimensions, key dimensions with persistent bias characteristics are accurately identified, local anomalies are avoided from interfering with classification judgments, and high-bias dimensions are re-weighted in attribution evaluation to improve the stability of classification in the absence of labels or ambiguous attribution. The dynamic modeling of data trajectories enhances the ability to identify and compensate for disturbance points and improves classification continuity. The time series monitoring and fluctuation adjustment mechanism of label frequency enhances the consistency of label expression. The semantic trajectory offset response optimizes the semantic matching accuracy of category attribution, and a closed-loop linkage between data behavior, classification stability and semantic adaptation is established. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a system flow chart of the present invention; Figure 2 This is a flow chart of the data bias identification module of the present invention; Figure 3 This is a flow chart of the attribution correction module of the present invention; Figure 4 This is a flow chart of the tension correction module of the present invention; Figure 5 This is a flow chart of the tag frequency analysis module of the present invention; Figure 6 This is a flow chart of the semantic offset adjustment module of the present invention. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0016] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0017] See also Figure 1 , the computer data intelligent analysis system based on artificial intelligence includes: The data bias identification module obtains multi-dimensional information for each data entity, constructs a set of local windows, calculates the difference fluctuation index between each dimension based on the distribution of data within the window, evaluates the change trend, and determines the fluctuation direction of multiple windows. If the difference trend within multiple windows is consistent, it marks the dimension as having a biased influence. It then iteratively analyzes the potential impact on classification and generates a biased influence result. The attribution correction module calls the bias impact results, extracts the bias dimensions, and corrects the bias of data classification based on the dimensions. By re-evaluating the dimension changes, it eliminates the impact of bias on data attribution, it iteratively determines the corrected attribution category, and generates corrected attribution information. The tension correction module constructs the change trajectory of data points between dimensions based on the corrected attribution information, identifies the tension change of data points within a time step, and marks the disturbance point if the change amplitude exceeds the set threshold. It then iteratively analyzes the impact of the disturbance point and corrects the disturbance through compensation to generate the tension compensation result. The tag frequency analysis module uses the tension compensation results to count the tag frequencies using a time sliding window. It then filters out tags with frequency fluctuations based on the frequency change trend. It iteratively analyzes the fluctuation patterns of tags within differentiated time periods, prioritizes adjustments for fluctuating tags, and generates tag frequency fluctuation results. The semantic shift adjustment module analyzes the shift of semantic categories within a time period based on the results of label frequency fluctuations. It determines whether semantic shift has occurred based on the changing trajectory of the category center point in each time period, adjusts the category matching method based on the trajectory mutation, and generates computer data intelligent analysis results.

[0018] The bias impact results include trend consistency identification, difference fluctuation indicator set, and local dimension bias label. The corrected attribution information includes corrected classification label, attribution confidence parameter, and dimension rescore distribution. The tension compensation results include disturbance point records, tension mutation amplitude set, and compensation correction factor. The label frequency fluctuation results include high-frequency label sequence, time period fluctuation characteristics, and label adjustment priority. The computer data intelligent analysis results include semantic center trajectory set, category offset node, and matching rule adjustment item.

[0019] See also Figure 2 , the data bias identification module includes: The local window construction submodule obtains multidimensional information of each data entity based on its structural characteristics and temporal properties. It then constructs multiple data window sets based on fixed window widths and sliding intervals, numbers and identifies the data within each window, and synchronizes the distribution intervals of differentiated data entities within the same dimension to obtain a window distribution alignment indicator. In the initial stages of data processing, to ensure localized and detailed data analysis, windows are first constructed based on the multidimensional information of each data entity. During this process, each window is set to contain a fixed number of data points—for example, a window width of 100 data points, with a sliding interval of 10 data points—to gradually cover the entire dataset. Under this setting, if a data entity contains 1,000 data points, 91 windows will be formed (the first window contains data points 1 to 100, the second window from 11 to 110, and so on). This sliding window approach helps capture local characteristics and volatility in the data. Furthermore, when constructing windows, the data within each window must be numbered and identified for subsequent analysis. This approach, for example, can be used in financial market analysis to capture patterns in stock price or currency fluctuations within a specific time window, enabling further trend analysis or anomaly detection, ultimately yielding a window distribution alignment metric.

[0020] The dimension fluctuation calculation submodule uses the window distribution alignment index to calculate the distribution difference of each dimension across all windows. Based on the numerical changes of the dimensions within multiple windows, it calculates the change difference and average offset rate between adjacent windows, and extracts the change direction vector of each dimension to obtain the dimension fluctuation difference value. Calculate the distribution difference of each dimension across all windows. First, calculate the numerical change of each dimension within each window by comparing the data difference of the same dimension between adjacent windows. For example, if the average value of a dimension is 100 and 110 in two consecutive windows, the change difference of the dimension is 10. By counting this difference across all windows, the average offset rate of the dimension can be estimated. In addition, it is necessary to extract the change direction vector of each dimension, which involves calculating the direction and intensity of the data change within each window. For example, if the data mainly shows growth, the change direction vector is positive; if the data mainly shows decline, the change direction vector is negative. In this way, for example, in consumer behavior analysis, the sales trend of a certain product during a promotion period can be observed, thereby helping the marketing team evaluate the promotion effect and ultimately calculate the dimension fluctuation difference value.

[0021] The bias trend determination submodule compares and determines the change trend direction of each dimension in the window based on the dimension fluctuation difference value. If the direction is consistent in multiple windows, the dimension is marked as a biased dimension using the formula: ; Obtain the bias fluctuation measurement value of the dimension by calculation, filter out the dimensions that are greater than the bias judgment threshold, obtain the bias impact index, and generate the bias impact result; in, Representative The biased volatility measure of the dimension, and Respectively represent The dimension in The difference between the positive and negative changes in the window, For the Dimension The change direction vector in the window, is the change amplitude corresponding to the direction, For the The overall mean change of the dimension, is the moving average of all dimensions; First, if the directions in multiple windows are consistent, that is, the change direction vectors of all windows are positive or negative, then the dimension is marked as a biased dimension. In a specific embodiment, the following formula is used for calculation: ; For example, suppose there is a dimension with positive differences in three windows [4.2, 3.8, 4.0], reverse difference is [3.0,3.1,2.9], the change direction vector is [0.6, 0.7, 0.5], the range of change is [2.0, 1.8, 1.9], the mean of this dimension is 3.9, the average mean of all dimensions is 3.63. First, calculate the numerator, which is the sum of the absolute values ​​of the forward and reverse differences of all windows: ; Next, calculate the denominator, which includes the square root of the sum of the squares of the product of the change direction vector and the change amplitude of all windows and the absolute value of the mean difference: ; ; Finally, calculate the bias volatility measure : ; if If the value is greater than a set threshold (e.g., 0.5), the dimension is labeled as a bias-influencing dimension, indicating that the dimension exhibits a consistent biased change trend across multiple windows. This analysis can clarify the stability and bias of each dimension, providing an important basis for further data analysis.

[0022] See also Figure 3 , the attribution correction module includes: The dimension screening submodule calls the bias impact indicator to screen dimensions whose fluctuation amplitude is greater than the attribution offset threshold, extracts key dimension information that affects data attribution judgment, and obtains the attribution bias dimension set; The first step is to extract key dimensions from the calculated bias impact indicators. Consider a dataset covering multiple business segments, such as a customer satisfaction survey, which includes multiple evaluation dimensions (e.g., service speed, product quality, and customer support). This process first identifies those dimensions with significant bias impact. Specifically, by comparing the bias impact indicators for each dimension, those dimensions whose values ​​exceed a set threshold (e.g., 0.8) are selected. For example, if the bias impact indicator for service speed is 0.85, it exceeds the threshold of 0.8 and is therefore selected as a key bias dimension. These dimensions are then labeled and collected into a set of attributed bias dimensions for subsequent correction. Ultimately, this operation generates a set of attributed bias dimensions, which includes all key bias dimensions that require adjustment, such as service speed and customer support, if their bias indices also exceed the threshold.

[0023] The bias resolution submodule determines the difference between the data change trend of the dimension and the category boundary position based on the attribute bias dimension set, and calculates the attribute deviation degree of each dimension to the classification result using the formula: ; The calculation obtains the attribution offset adjustment value of the bias dimension, adjusts and corrects the corresponding values ​​of all bias dimensions, and generates an attribution correction data set; in, Indicates the The attribute offset adjustment value of the bias dimension, For the The dimension in The weight of the category in the sample, is the boundary distance value in the corresponding sample, is the fluctuation interference term, For the Dimension The original classification value of the sample, For the Sample target category value, For the The cumulative value of the disturbance in each dimension, is the trend suppression factor under dimension; First, calculate the attribution offset adjustment value for each dimension using the following formula: ; Taking the dimension "service speed" as an example, let Rxγ (category weight), Zxγ (boundary distance), and Yxγ (fluctuation interference term) be 0.65, 1.8, and 0.5, respectively. The calculation process is as follows: 1. Calculate the square root of the sum of the squares of the distance and interference term: 2. Calculate the numerator: 3. Assuming the threshold Tx and the attribute difference Hxη are 1.0 and 0.9, calculate the first term in the denominator: 4. Add the adjustment value of the disturbance accumulation value Ψx and the trend suppression factor ζx: 5. Calculate the attribution offset adjustment value Ξx:

[0024] This value indicates that the attribution judgment for the dimension "Service Speed" is significantly biased and needs to be adjusted. In this way, an attribution bias adjustment value is calculated for each key bias dimension and the attribution classification of each dimension is adjusted accordingly to reduce the impact of bias.

[0025] The attribution re-classification submodule re-establishes the distance mapping matrix between samples and categories based on the attribution correction data set, selects and classifies the corrected samples according to their attribution at the center point of multiple categories, and obtains the corrected attribution information; First, a distance mapping matrix is ​​established between the corrected dataset and the center points of each category. For example, if the data for service speed and customer support dimensions were corrected in the previous step in a customer feedback analysis of a telecommunications service center, each sample in these dimensions needs to be reclassified. By comparing the relative position of each sample with respect to the center points of each category, the most appropriate category for it is re-determined. For example, a sample that was originally misclassified as low satisfaction due to bias may be closer to the center of a high satisfaction category after correction. Ultimately, this process generates corrected classification information, ensuring the accuracy of data classification and minimizing bias, thereby improving the reliability and effectiveness of the overall data analysis.

[0026] See also Figure 4 , the tension correction module includes: The trajectory construction submodule extracts the multidimensional state values ​​of the data points in continuous time steps based on the corrected attribution information, calculates the change rate and direction combination in each time step, and splices them in chronological order to form a multidimensional vector path to obtain the dimension change trajectory value; According to the corrected attribution information, the multidimensional attribute sequence of the data points in the continuous time steps needs to be extracted one by one. The corrected attribution information usually represents the corrected attribution result of each sample point in the standard category space. For example, the dimensional attributes of the data point D1 in the time series T1 to T5 are temperature T, humidity H and air pressure P, respectively, where the value of T is [22.1, 22.3, 22.7, 23.2, 23.6], H is [55, 57, 58, 60, 63], and P is [101.1, 101.0, 100.8, 100.6, 100.3]. For D1, its three-dimensional state vector for each time step is (t, h, p). In order to construct the multidimensional change trajectory, the time steps t and t need to be calculated. +1, and record its direction, that is, the Δv between step 1 and step 2 is (22.3-22.1, 57-55, 101.0-101.1) = (0.2, 2, -0.1), recorded as positive direction, positive direction, and negative direction to form a direction sequence. At the same time, the moduli of all Δv are normalized to obtain the unit direction vector, forming the change path of the data point in the entire time interval; in practical applications, if the change rate standard of each dimension is set to maximum value normalization, the maximum temperature change is 1.5°C, the humidity is 10%, and the air pressure is 1.2hPa. Then, after normalizing Δv, the direction intensity vector is obtained, and then the change trajectory under all time steps is accumulated to obtain the dimension change trajectory value.

[0027] The disturbance identification submodule calculates the magnitude of the dimension change of the data points between consecutive time steps based on the dimension change trajectory value, and compares it with the set tension disturbance threshold to determine whether it constitutes a disturbance state. The formula is: ; The disturbance intensity coefficient of each time step is obtained by calculation. If the intensity coefficient is greater than the tension disturbance threshold, the corresponding position is marked as the disturbance occurrence point, and the disturbance position identification set is obtained; in, For the The perturbation intensity coefficient of the data point, is the reference amplitude of normal tension change, For the The data point is The coefficient of the direction of change of the time step, is the change amplitude value of the point, is the cumulative value of fluctuation at a time point, For the The stability factor of the time step, is the tension projection value of the category to which the data point belongs, The tension effect mapping value corresponding to time; The core of disturbance identification lies in calculating the disturbance intensity coefficient Θ^μ, which is used to measure whether a data point has tension disturbance within a certain period of time. Taking data point D2 as an example, its monitoring records between time steps ρ=1 to ρ=3 are selected. Its direction coefficients Υ are 0.78, 0.81, and 0.75 respectively, the variation amplitudes Ω are 0.29, 0.32, and 0.28, the fluctuation values ​​Λ are 0.15, 0.13, and 0.16, the stability factor Φ is uniformly set to 1.2, the projection value τ is 0.48, the influence mapping value χ is 0.44, and the reference amplitude δ is set to 0.33. Now put them into the formula for calculation: ; Here are the steps: (1) Calculate the first group of items: ; ; (2) Group 2 items: ; (3) Group 3 items: ; Summation term: ; This value is greater than the set disturbance threshold of 1.2, indicating that D2 is disturbed in this time period, so its corresponding time step is included in the disturbance position identification set.

[0028] The disturbance compensation submodule extracts the corresponding change path segment and the previous time step value of the disturbance according to the disturbance position identification set, calculates the disturbance direction and amplitude error and performs reverse adjustment, applies the compensation vector to the corresponding time step value, and obtains the tension compensation result; Obtain the original state vector of the disturbance data point at the determined time step and its state at the previous time step one by one, determine the vector angle difference between the direction vector of the disturbance position and the direction vector of the position before the disturbance, and calculate the compensation offset according to the tension error trend. For example, for the disturbance point D3, it is determined to be a disturbance at time step 4, the original state vector is (23.4, 61, 100.5), the state at the previous time step is (22.9, 60, 100.7), and the change direction is (+0.5, +1, -0.2). Since the amplitude deviates from the standard in three dimensions, the deviation is significant. The quasi-interval setting (maximum temperature change 0.4, maximum humidity change 0.8, maximum pressure change 0.15) indicates that the disturbance has exceeded the limit. According to the proportional correction coefficient, the reverse vector compensation percentage of the exceeded part is set to 60%, that is, the correction value is (-0.06, -0.12, +0.03). The new state after compensation is (23.34, 60.88, 100.53). The tension is then re-compared with the belonging center vector, and the stability of the change trend is recorded. If the new trend falls into the stable interval, the compensation is valid, and the tension compensation result is finally formed.

[0029] See also Figure 5 , the label frequency analysis module includes: The frequency statistics submodule collects the tag data within the time window based on the tension compensation results, calculates the number of occurrences and the frequency ratio of each tag within the window, and constructs the tag frequency sequence in the time dimension to obtain the tag time series frequency value; Collect the tag data of each tag in each time window, and count the number of times the tag appears in each time period. For example, the number of times a monitored object is marked with tag A in time window 1 to time window 4 is 12, 18, 14 and 20 times respectively. In practical applications, the frequency statistics of each tag need to be extracted from the original log records or the record logs of the monitoring system. For example, if tag B appears 6 times in 5 seconds and 8 times in the next 5 seconds in the system record, the frequencies marked in the two time windows are 6 and 8 times respectively. Then convert the frequency data into relative frequency percentage. For example, if the total number of tags in window 1 is 60 times, the relative frequency of tag A is 12 / 60=0.2. If the frequency of tag B is 6 times, the relative frequency is 6 / 60=0.1. Then construct a frequency sequence in the time dimension for subsequent trend judgment and fluctuation detection. The frequency sequence must have a consistent time span definition. For example, collecting data within 30 seconds at a sliding interval of 5 seconds can form 6 time spans. Time window, the data of each window is regarded as a point in the frequency sequence. It is further necessary to establish a matrix form of frequency record for different label frequency data for subsequent call. For example, the i-th row in the frequency matrix corresponds to label i, and the j-th column is the frequency value in time window j. This matrix form provides direct data support for the subsequent calculation of fluctuation trend. The time window length and sliding interval must be kept consistent throughout the process. For example, when the window width is set to 5 seconds and the sliding step is 1 second, 56 windows will be obtained for a sequence with a total duration of 60 seconds. The frequency statistical operation must be completed independently in each window to ensure that the results can be smoothly compared. In the implementation scenario, for example, when collecting personnel behavior labels in a video surveillance scenario, the label "running" is recorded 7 times in window 1 and 10 times in window 2. In the window sequence, it is a frequency value sequence [7, 10, ...]. The corresponding label "falling" record value is [0, 1, ...], forming label frequency time series data and generating label time series frequency values.

[0030] The fluctuation screening submodule calls the tag time series frequency value, calculates the frequency difference and fluctuation gradient of the tag between adjacent time windows, and identifies high-fluctuation tags using the formula: ; The frequency fluctuation coefficient of each tag is obtained by calculation, and tags with fluctuation coefficients higher than the tag fluctuation determination threshold are screened to obtain a high-fluctuation tag set; in, Representation Label The frequency fluctuation coefficient, Representation Label In the The frequency of the time window, is the slope of frequency change, is the cumulative offset value, is the label distribution density within the window, is the average frequency of the tag over the entire period, For the The median frequency of the time window, For label The normalized stability factor of Call the tag time series frequency value and process the frequency change of each tag between adjacent time windows to calculate the frequency fluctuation coefficient. First, define the parameters involved in the formula. Representation Label The frequency fluctuation coefficient, For the label The frequency in the time window is obtained through the above statistics. For example, if the frequency of label A in the four time windows is 12, 18, 14, and 20, then its difference sequence is |18-12|=6, |14-18|=4, |20-14|=6, and the sum of the differences is 16. The frequency change slope It can be expressed as the ratio of the frequency difference between two consecutive windows, that is, (18-12) / 1=6, (14-18) / 1=-4, (20-14) / 1=6, then the corresponding square values ​​are 36, 16, 36, and the cumulative offset value is It can be set to the average of the sum of the squares of the three slopes, that is, (36+16+36) / 3=29.33, and the distribution density Set to 1.2 (indicating the density of label distribution in each window), the average label frequency =16, median label frequency =16, stability factor =0.8, normalized stability factor =0.95, substitute into the formula: ; ; ; The result shows that the frequency fluctuation coefficient of label A is 0.5679. If the fluctuation judgment threshold is set to 0.3, the coefficient exceeds the threshold, indicating that label A is a high-fluctuation label and should be screened into the high-fluctuation label set. The benefit of the formula is that it dynamically evaluates the persistence and suddenness of label fluctuation changes through the composite structure of slope, offset, label density and window stability factor, forming an operational judgment standard.

[0031] The label adjustment submodule extracts the frequency series of the corresponding time period based on the high-volatility label set, calculates the change trend and direction ratio of the time point, determines the label frequency change tendency, reconstructs the weight coverage range, and adjusts the label mapping to obtain the label frequency fluctuation results; Extract the frequency series data in the corresponding time period. Take label D as an example. The frequency values ​​of its four time windows are 15, 13, 16, and 12. Calculate the change trend and direction ratio of each time point. The trend direction can be judged by the difference sign. For example, the change from time window 1 to 2 is 13-15=-2, which is a downward trend. From 2 to 3, it is 16-13=+3, which is an upward trend, forming a direction sequence [−,+]. Further, each change trend is converted into a numerical quantization process, corresponding to +1 for upward and −1 for downward, forming a trend vector. Perform short-term fluctuation analysis on the trend vector and extract the number of consecutive opposite changes. For example, the direction vector [+,−,+,−] indicates frequent fluctuations. The number is 3. The number of frequent fluctuations is compared with the maximum possible value to form the fluctuation ratio. For example, the maximum is 3, the actual value is 3, and the ratio is 1.0. If it is greater than the threshold value 0.75, it is marked as a label that needs to be adjusted. The linear convergence method can be used to reconstruct the label weight. For example, if the current label D has a frequency of 16 times in the third window, it becomes 16 / 60≈0.267 after normalization. If the fluctuation value in the previous step is too large, the adjustment factor can be set to 0.8, and the frequency is compressed to 16×0.8=12.8, remapped to 13, and finally the adjusted frequency sequence is [15,13,13,12]. The corresponding row of the label in the frequency matrix is ​​updated to form a new label time series frequency data, and the label frequency fluctuation result is generated.

[0032] See also Figure 6 ,The semantic offset adjustment module includes: The trajectory extraction submodule extracts the center point coordinates of each semantic category within a continuous time period based on the label frequency fluctuation results, calculates the spatial change path between time steps, constructs a multi-period category trajectory sequence in chronological order, and obtains the category center trajectory value; Based on the label frequency fluctuation results, we obtain label records for different time periods. For each semantic category, we extract its position coordinates according to the sample data in each time step. Based on the value of each dimension of the data point belonging to this label in the multidimensional feature space, we calculate the spatial position of the category center point in each time period. For example, in two-dimensional space, if there are three samples under a semantic category in time period 1, and their feature values ​​are (1.2, 2.3), (1.0, 2.1), and (1.1, 2.2), then its center point is (1.1, 2.2). Similarly, for samples (1.3, 2.5), (1.5, 2.6), and (1.4, 2.4) in time period 2, we obtain the center point (1.4, 2.5). These are then connected in chronological order to form a trajectory. Subsequently, the trajectory of all category center point sequences is spliced ​​to form the movement path of the semantic category in the time dimension. The above process is repeated for each semantic label to obtain the trajectory set of each label changing over time, namely the category center trajectory value.

[0033] The offset recognition submodule calculates the center point movement amplitude and direction transfer angle in a continuous time period based on the category center trajectory value, segments the mutation path and sets the offset judgment boundary using the formula: ; Obtain the semantic shift degree value of each category by operation, determine whether a sudden change occurs, and obtain the semantic shift category identification set; in, Indicates the The semantic deviation value of the category, For the Category in The displacement value of the center point of the time period, is the slope of its direction change, is the angle difference between the front and back vectors of the point, is the path disturbance amplitude, is the category aggregation stabilization factor, For the The amount of fluctuation mapping for the time period, is the evolutionary stability of the semantic category itself, For category The offset absorption threshold, For the Deviation of the center point of the time period; Based on the category center trajectory values ​​obtained above, we extract the spatial displacement, directional slope, and path perturbation amplitude of the center point of each category between consecutive time periods, and calculate its semantic drift degree. Taking category A as an example, its center point displacement values ​​in the four time periods are 0.6, 0.7, 0.5, and 0.4, respectively; the directional change slope is 0.9, 1.0, 0.8, and 0.7; the directional angle difference is 0.6, 0.5, 0.7, and 0.6; the perturbation amplitude is 0.2, 0.3, 0.25, and 0.2; the aggregation stability factor is 0.9; the time period fluctuation mapping amount is 0.3, 0.4, 0.5, and 0.6; the category drift absorption threshold is 1.2; and the time period center point deviation is 0.3, 0.2, 0.25, and 0.2.

[0034] The semantic deviation value of category A is calculated according to the following formula: ; Calculate step by step: 1) Paragraph 1: ; ; 2) Paragraph 2: ; ; 3) Paragraph 3: ; ; 4) Paragraph 4: ; ; Calculate the sum of the numerators: 1.064 + 1.381 + 0.949 + 0.821 = 4.215; Denominator: ; Final offset value: ; The result shows that the semantic deviation degree value of category A is greater than the judgment threshold of 1.2, so category A belongs to the semantic deviation category.

[0035] The benefit of the formula is that it integrates multiple dynamic features into one by introducing factors such as directional angle difference, disturbance amplitude, aggregation stability factor and fluctuation mapping amount, effectively capturing the mutation behavior existing in the category movement trajectory.

[0036] The matching correction submodule extracts the frequency path and center point trajectory of the offset category in the corresponding time period based on the semantic offset category identification set, reconstructs the spatial distance matrix between categories, readjusts the category attribution method and marks the mapping position, and obtains the computer data intelligent analysis results; According to the offset category numbers recorded in the semantic offset category identification set, the frequency series and spatial trajectory data of these categories in each time period are extracted. Based on the frequency series, the activity curve of the category in each time period can be constructed. Then, by comparing the change path of its spatial center trajectory with the trajectory direction of other categories in the same time period, the distance matrix between adjacent categories is constructed. For example, the distance between the center point of category A and category B in time period 1 is 0.5 and in time period 2 is 0.6. The Euclidean distance series of the four time periods are calculated accordingly. By setting a distance boundary, such as 0.4 as the attribution update critical value, it is determined whether matching adjustment is required. If the distance between the two categories in a certain time period is less than 0.4, the label matching attribution of the time point is adjusted from category A to B, or the category sample is set as a fuzzy attribution mark. The category label distribution is further reconstructed based on the distance change trend to generate computer data intelligent analysis results.

[0037] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. An artificial intelligence-based computer data intelligent analysis system, characterized by: The system comprises: The data bias identification module obtains multi-dimensional information of data entities, calculates dimensional difference fluctuation indicators based on the distribution characteristics within the window, determines the fluctuation direction of multiple windows, iteratively analyzes the potential impact on the classification results, and generates bias impact results; The attribution correction module calls the bias impact result, extracts the bias dimension, re-evaluates and corrects the bias in combination with the classification behavior, and iteratively generates corrected attribution information; The tension correction module constructs a change trajectory of the data point in the multidimensional space based on the corrected attribution information, identifies the tension mutation within the time step, marks it as a disturbance point if it exceeds a threshold, and generates a tension compensation result through compensation correction; The tag frequency analysis module calls the tension compensation result, uses a time sliding window to count the tag frequency, filters out tags with frequency fluctuations, and generates a tag frequency fluctuation result; The semantic shift adjustment module analyzes the changes in the center trajectory of the semantic category within the time period based on the tag frequency fluctuation results, adjusts the matching method according to the trajectory mutation, and generates computer data intelligent analysis results.

2. The computer data intelligent analysis system based on artificial intelligence according to claim 1, characterized in that: The bias impact results include trend consistency identification, difference fluctuation index set, and local dimension bias label; the corrected attribution information includes modified classification label, attribution confidence parameter, and dimension rescore distribution; the tension compensation results include disturbance point record, tension mutation amplitude set, and compensation correction factor; the label frequency fluctuation results include high-frequency label sequence, time period fluctuation characteristics, and label adjustment priority; the computer data intelligent analysis results include semantic center trajectory set, category offset node, and matching rule adjustment item.

3. The computer data intelligent analysis system based on artificial intelligence according to claim 1, characterized in that: The data bias identification module includes: The local window construction submodule obtains multidimensional information of each data entity based on its structural characteristics and temporal properties. It then constructs multiple data window sets based on fixed window widths and sliding intervals, numbers and identifies the data within each window, and synchronizes the distribution intervals of differentiated data entities within the same dimension to obtain a window distribution alignment indicator. The dimension fluctuation calculation submodule calls the window distribution alignment index to calculate the distribution difference of each dimension in all windows. Based on the numerical changes of the dimensions in multiple windows, it calculates the change difference and average offset rate between adjacent windows, and extracts the change direction vector of each dimension to obtain the dimension fluctuation difference value. The bias trend determination submodule compares and determines the change trend direction of each dimension in the window based on the dimension fluctuation difference value. If the direction is consistent in multiple windows, the dimension is marked as a bias dimension, using the formula: ; Obtain the bias fluctuation measurement value of the dimension by calculation, filter out the dimensions that are greater than the bias judgment threshold, obtain the bias impact index, and generate the bias impact result; in, Representative The biased volatility measure of the dimension, and Respectively represent The dimension in The difference between the positive and negative changes in the window, For the Dimension The change direction vector in the window, is the change amplitude corresponding to the direction, For the The overall mean change of the dimension, is the moving average of all dimensions.

4. The computer data intelligent analysis system based on artificial intelligence according to claim 1, characterized in that: The attribution correction module includes: The dimension screening submodule calls the bias impact indicator, screens dimensions whose fluctuation amplitude is greater than the attribution offset threshold, extracts key dimension information that affects data attribution judgment, and obtains an attribution bias dimension set; The bias resolution submodule determines the difference between the data change trend of the dimension and the category boundary position based on the attribute bias dimension set, and calculates the attribute deviation degree of each dimension to the classification result using the formula: ; The calculation obtains the attribution offset adjustment value of the bias dimension, adjusts and corrects the corresponding values ​​of all bias dimensions, and generates an attribution correction data set; in, Indicates the The attribute offset adjustment value of the bias dimension, For the The dimension in The weight of the category in the sample, is the boundary distance value in the corresponding sample, is the fluctuation interference term, For the Dimension The original classification value of the sample, For the Sample target category value, For the The cumulative value of the disturbance in each dimension, is the trend suppression factor under dimension; The attribution re-judgment submodule re-establishes the distance mapping matrix between samples and categories based on the attribution correction data set, selects and classifies the corrected samples according to their attribution at the center point of multiple categories, and obtains the corrected attribution information.

5. The computer data intelligent analysis system based on artificial intelligence according to claim 1 is characterized in that: The tension correction module includes: The trajectory construction submodule extracts the multidimensional state values ​​of the data points in continuous time steps based on the corrected attribution information, calculates the change rate and direction combination in each time step, and splices them in time sequence to form a multidimensional vector path to obtain the dimension change trajectory value; The disturbance identification submodule calculates the magnitude of the dimension change of the data points between consecutive time steps based on the dimension change trajectory value, and compares it with the set tension disturbance threshold to determine whether it constitutes a disturbance state. The formula is: ; The disturbance intensity coefficient of each time step is obtained by calculation. If the intensity coefficient is greater than the tension disturbance threshold, the corresponding position is marked as the disturbance occurrence point, and the disturbance position identification set is obtained; in, For the The perturbation intensity coefficient of the data point, is the reference amplitude of normal tension change, For the The data point is The coefficient of the direction of change of the time step, is the change amplitude value of the point, is the cumulative value of fluctuation at a time point, For the The stability factor of the time step, is the tension projection value of the category to which the data point belongs, The tension effect mapping value corresponding to time; The disturbance compensation submodule extracts the corresponding change path segment and the previous time step value of the disturbance according to the disturbance position identification set, calculates the disturbance direction and amplitude error and performs inverse adjustment, applies the compensation vector to the corresponding time step value, and obtains the tension compensation result.

6. The computer data intelligent analysis system based on artificial intelligence according to claim 1, characterized in that: The tag frequency analysis module includes: The frequency statistics submodule collects the tag data in the time window according to the tension compensation result, calculates the number of occurrences and the frequency ratio of each tag in the window, and constructs the tag frequency sequence in the time dimension to obtain the tag time series frequency value; The fluctuation screening submodule calls the tag time series frequency value, calculates the frequency difference and fluctuation gradient of the tag between adjacent time windows, and identifies high-fluctuation tags using the formula: ; The frequency fluctuation coefficient of each tag is obtained by calculation, and tags with fluctuation coefficients higher than the tag fluctuation determination threshold are screened to obtain a high-fluctuation tag set; in, Representation Label The frequency fluctuation coefficient, Representation Label In the The frequency of the time window, is the slope of frequency change, is the cumulative offset value, is the label distribution density within the window, is the average frequency of the tag over the entire period, For the The median frequency of the time window, For label The normalized stability factor of The label adjustment submodule extracts the frequency sequence of the corresponding time period based on the high-volatility label set, calculates the change trend and direction ratio of the time point, determines the label frequency change tendency, reconstructs the weight coverage range and performs label mapping adjustment to obtain the label frequency fluctuation result.

7. The computer data intelligent analysis system based on artificial intelligence according to claim 1 is characterized in that: The semantic offset adjustment module includes: The trajectory extraction submodule extracts the center point coordinates of each semantic category within a continuous time period based on the label frequency fluctuation results, calculates the spatial change path between time steps, constructs a multi-period category trajectory sequence in chronological order, and obtains the category center trajectory value; The offset recognition submodule calculates the center point movement amplitude and direction transfer angle in a continuous time period based on the category center trajectory value, segments the mutation path and sets the offset judgment boundary using the formula: ; Obtain the semantic shift degree value of each category by operation, determine whether a sudden change occurs, and obtain the semantic shift category identification set; in, Indicates the The semantic deviation value of the category, For the Category in The displacement value of the center point of the time period, is the slope of its direction change, is the angle difference between the front and back vectors of the point, is the path disturbance amplitude, is the category aggregation stabilization factor, For the The amount of fluctuation mapping for the time period, is the evolutionary stability of the semantic category itself, For category The offset absorption threshold, For the Deviation of the center point of the time period; The matching correction submodule extracts the frequency path and center point trajectory of the offset category in the corresponding time period based on the semantic offset category identification set, reconstructs the spatial distance matrix between categories, readjusts the category attribution method and marks the mapping position, and obtains computer data intelligent analysis results.

Citation Information

Cited By

  • Fuzzy test method and system based on program control flow

    CN120686799A

  • A fuzz testing method and system based on program control flow

    CN120686799B

  • New energy equipment intelligent operation and maintenance management system based on Internet of Things

    CN120779758A

  • Method and system for identifying operation mode of large-scale new energy grid-connected power system

    CN121211137A

  • Computer data intelligent matching analysis system based on deep learning

    CN121327404A