A data processing method based on multidimensional data
By cleaning and preprocessing historical identification data, calculating validity and reliability, analyzing the incidence of bias, and formulating management evaluation criteria, the problem of monitoring and managing abnormal data identification biases was solved, and efficient data identification results were achieved.
Patent Information
- Application Number
- CN202411904590.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-21
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-03-21
AI Technical Summary
Existing technologies are unable to efficiently monitor and manage deviations and anomalies in data identification, making it difficult to improve the monitoring effectiveness of data identification.
By cleaning and preprocessing historical identification data, grouping and labeling, calculating the validity and reliability of anomaly identification data, comparing reliability with critical thresholds, analyzing the incidence of deviations, formulating management evaluation criteria, and generating management instructions at different levels.
It enables timely and efficient management of data identification deviations, ensures the normal operation of data processing methods, and improves the monitoring effect of data identification.
Smart Images

Figure CN119829914B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis technology, and in particular to a data processing method based on multidimensional data. Background Technology
[0002] Artificial intelligence refers to the process of using artificial intelligence technology and algorithms to analyze, judge, and classify input data. It can recognize and understand various forms of data such as images, voice, text, and video, and extract useful information and patterns from them.
[0003] However, current technologies are unable to proactively and efficiently monitor and analyze deviations and anomalies in identification data, and it is difficult to implement targeted management to improve the monitoring effectiveness of data identification. Therefore, this invention proposes a data processing method based on multidimensional data. This method processes and filters acquired historical identification data to obtain anomaly identification data, calculates and analyzes the validity and reliability of the anomaly identification data to obtain deviation data, and evaluates and manages the deviation data to improve the timely and efficient monitoring and handling of anomalies. Summary of the Invention
[0004] The purpose of this invention is to solve the problems in the background art by proposing a data processing method based on multidimensional data.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] A data processing method based on multidimensional data includes:
[0007] Obtain historical identification data from the target database, clean the obtained historical identification data, and preprocess and filter the cleaned data to obtain abnormal identification data;
[0008] The anomaly identification data are grouped and labeled, and the validity and reliability of each anomaly identification data are calculated to obtain the corresponding validity and reliability.
[0009] Deviation data is obtained by comparing and analyzing the reliability of anomaly identification data with a preset reliability threshold.
[0010] Calculate and analyze the deviation rate of the deviation data, and dynamically formulate deviation data management and evaluation criteria based on the analysis results.
[0011] Preferably, the steps of cleaning the acquired historical identification data and preprocessing the cleaned data include:
[0012] Automatically identify and correct data quality issues using data quality tools, and standardize data formats;
[0013] Historical recognition data is acquired based on the historical recognition data acquisition module and grouped according to different types, including natural language data group, voice data group and image data group.
[0014] Each data set of type is labeled Hj, j = 1, 2, 3.
[0015] Preferably, the recognition rate of the statistical data set Hj is used; the recognition rate of the data set is labeled as SLj;
[0016] Set the standard value of the data group recognition rate Bj; and calculate the recognition index ZS corresponding to the data group set using the formula ZS=(SLj / Bj)×100%.
[0017] When plotting the recognition index curve based on the calculated recognition index value, and analyzing the trend of recognition index changes in historical recognition data;
[0018] Obtain the identification index change curve corresponding to any set of identification data; determine the abnormal threshold value through the identification index change curve.
[0019] Several calculation periods are determined based on the abnormal threshold value. The average of the abnormal threshold values of all data groups is taken based on the calculation period. Data below the average value is marked as abnormal identification data.
[0020] Preferably, the anomaly identification data is grouped according to the corresponding data type and labeled Ai, i = 1, 2, 3, based on several calculation cycles, and the identification rate Si corresponding to the anomaly identification data is recorded by group;
[0021] For each anomaly identification data point, the validity is calculated to obtain the validity D of each anomaly identification data point; based on the validity of each anomaly identification data point, the reliability is further calculated to obtain the reliability XD of the anomaly identification data point.
[0022] Preferably, when calculating the validity of each anomaly identification data, the normal response time Ti0 of each identification data group is determined; the response time Ti and identification rate Si of the identification process of each anomaly identification data are statistically analyzed; and the validity D of each anomaly identification data is calculated using the formula D = (Ti - Ti0) × Si.
[0023] Preferably, the reliability of each anomaly identification data point is further calculated based on its validity, and the identification complexity Fi of the identification process for each anomaly identification data point is statistically analyzed using the formula... Calculate the reliability XD of the obtained anomaly identification data; where α is the preset proportional coefficient of the reliability of the anomaly identification data, and 0 < α < 1.
[0024] Preferably, the reliability XD of the anomaly identification data is compared and analyzed with a preset reliability threshold XD0; if XD is not greater than XD0, the reliability of the anomaly identification data is determined to be invalid and an invalid label is generated, and the corresponding anomaly identification data is marked as deviation data according to the invalid label; if XD is greater than XD0, the reliability of the anomaly identification data is determined to be valid and a valid label is generated.
[0025] Preferably, the reliability corresponding to the deviation data is obtained, and then expressed using the formula... The deviation occurrence rate PF is calculated, where β is a variable constant parameter;
[0026] The deviation occurrence rate is evaluated, management criteria are formulated, the deviation occurrence rate is matched with all preset deviation sending ranges to obtain the corresponding deviation sending range, and corresponding management instructions are generated. The obtained management instructions are sent to technicians of different levels for processing and analysis.
[0027] Compared with existing technologies, the advantages of this invention, which provides a data processing method based on multidimensional data, are as follows:
[0028] This invention collects historical identification data from an artificial intelligence database, groups and labels it, and filters and cleans the data to avoid duplication, incompleteness, and inaccuracy. An anomaly identification and analysis module receives preprocessed data and acquires anomaly identification data, facilitating better data management and utilization, and enabling efficient data analysis. Validity and reliability are calculated from all anomaly identification data. The reliability of the anomaly identification data is compared with a preset reliability threshold to obtain deviation data. Based on the deviation data, a comprehensive evaluation and management system is implemented, establishing deviation data management and evaluation criteria and generating management instructions at different levels.
[0029] In summary, this invention can calculate the deviation rate based on actual conditions and deviation data, analyze and process the calculation results, determine whether the deviation rate exceeds a preset range by evaluating the numerical value of the deviation rate, generate different levels of management instructions from the system, and send the information to technicians at different levels for processing, thereby ensuring the normal operation of a subsequent data processing method based on multidimensional data. Attached Figure Description
[0030] Figure 1 This is a flowchart of a data processing method based on multidimensional data proposed in this invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] Reference Figure 1 A data processing method based on multidimensional data, comprising:
[0033] The historical identification data acquisition module acquires historical identification data from the target database, cleans the acquired historical identification data, and preprocesses and filters the cleaned data to obtain abnormal identification data.
[0034] The historical identification data acquisition module includes an instruction receiving unit, a data acquisition unit, a data output unit, and a data preprocessing unit.
[0035] The historical identification data acquisition module is connected to the PC, and the instruction receiving unit is used to receive query instructions for historical identification data issued by the system.
[0036] The data acquisition unit is used to obtain the row number and table of the database where the historical identification data is located according to the query instruction, acquire the historical identification data according to the row number and table, and input the acquired historical identification data into the target table;
[0037] The data output unit is used to extract information from the target data in the target table and store the acquired target data;
[0038] The data preprocessing unit is used to clean the stored target data, automatically identify and correct data quality problems through data quality tools, standardize the data format, and preprocess the cleaned data.
[0039] The anomaly identification data are grouped and labeled, and the validity and reliability of each anomaly identification data are calculated to obtain the corresponding validity and reliability.
[0040] Deviation data is obtained by comparing and analyzing the reliability of anomaly identification data with a preset reliability threshold.
[0041] Calculate and analyze the deviation rate of the deviation data, and dynamically formulate deviation data management and evaluation criteria based on the analysis results.
[0042] It should be noted that the application object in this embodiment of the invention can be the data deviation monitoring database of an artificial intelligence recognition platform developed by a technology company. It is used to statistically analyze whether the monitoring and recognition data is abnormal, and to manage and store the abnormal deviation data. The indicators for monitoring and recognition data include image data, voice data, and natural language data of the artificial intelligence recognition platform, as well as other indicators.
[0043] The steps for cleaning the acquired historical identification data, and then preprocessing and filtering the cleaned data to obtain anomaly identification data include:
[0044] Historical recognition data is acquired based on the historical recognition data acquisition module and grouped according to different types, including natural language data group, voice data group and image data group.
[0045] Label each type of data set with Hj, j = 1, 2, 3;
[0046] Artificial intelligence recognition has a wide range of applications. In natural language processing tasks, such as text classification and sentiment analysis, AI systems with high recognition rates can improve the accuracy of text understanding and semantic analysis. In the fields of speech and image recognition, improved AI recognition rates can help systems accurately identify speech and classify objects, scenes, and people in images, which is beneficial for applications such as smart homes, image search, autonomous driving, and security monitoring.
[0047] The recognition rate of the statistical data set Hj; the recognition rate of the data set is labeled as SLj;
[0048] It should be noted that the recognition rate of artificial intelligence refers to the accuracy of correctly identifying or classifying a given task or problem in an artificial intelligence system. It is usually expressed as a percentage, representing the proportion of correctly identified samples out of the total number of samples.
[0049] Set the standard value of the data group recognition rate Bj; and calculate the recognition index ZS corresponding to the data group set using the formula ZS=(SLj / Bj)×100%.
[0050] When plotting the recognition index curve based on the calculated recognition index value, and analyzing the trend of recognition index changes in historical recognition data;
[0051] Obtain the identification index change curve corresponding to any set of identification data; determine the abnormal threshold value through the identification index change curve.
[0052] Several calculation periods are determined based on the anomaly threshold value. The average of the anomaly threshold values for all data groups is taken based on the calculation period. Data below the average value is marked as anomaly identification data.
[0053] For example, taking a certain voice dataset as an example, the recognition rate SL of the voice dataset is obtained as {80%; 75%; 55%; 60%; 45%; 95%}; the standard value of the voice dataset recognition rate is set as B = 90%; then the recognition index ZS of the dataset is {89%; 83%; 61%; 67%; 50%; 101%}; a recognition index curve is plotted with the voice dataset number on the horizontal axis and the recognition index of the voice dataset on the vertical axis; the normal value of the recognition index is set as 80%; the abnormal threshold value is determined by the change curve of the recognition index as {61%; 67%; 50%}; the average of the abnormal threshold values of all voice datasets is taken; then the data below the average value is marked as abnormal recognition data.
[0054] Based on several calculation cycles, the anomaly identification data are grouped according to the corresponding data type and labeled Ai, i = 1, 2, 3, and the identification rate Si corresponding to the anomaly identification data is recorded for each group.
[0055] For each anomaly identification data point, the validity is calculated to obtain the validity D of each anomaly identification data point; based on the validity of each anomaly identification data point, the reliability is further calculated to obtain the reliability XD of the anomaly identification data point.
[0056] Specifically, when calculating the validity of each anomaly identification data, the normal response time Ti0 for each data group is determined; the response time Ti and identification rate Si of the identification process for each anomaly identification data are statistically analyzed; and the validity D of each anomaly identification data is calculated using the formula D = (Ti - Ti0) × Si.
[0057] Furthermore, reliability calculations are performed on each anomaly detection data point based on its validity, and the detection complexity Fi of the detection process for each anomaly detection data point is calculated using the formula. Calculate the reliability XD of the obtained anomaly identification data; where α is the preset proportional coefficient of the reliability of the anomaly identification data, and 0 < α < 1.
[0058] Among them, recognition complexity refers to the complexity of the object being recognized by artificial intelligence. It can be obtained according to the established level standards. The first level of artificial intelligence recognition complexity is to perceive the environment, collect data and conduct preliminary understanding and analysis; the second level of complexity is to reason and make decisions based on existing knowledge and predict future situations; the third level of complexity is to autonomously learn and innovate and apply it to the real environment; in addition, the higher the complexity, the longer the response time.
[0059] The reliability XD of the anomaly identification data is compared and analyzed with the preset reliability threshold XD0;
[0060] If XD is not greater than XD0, the reliability of the anomaly identification data is determined to be invalid and an invalidation label is generated. Based on the invalidation label, the corresponding anomaly identification data is marked as deviation data.
[0061] If XD is greater than XD0, it is determined that the reliability of the anomaly recognition data is valid and a valid label is generated;
[0062] Obtain the reliability corresponding to the deviation data, and calculate the deviation occurrence rate PF through the formula where β is a variable constant parameter;
[0063] Evaluate according to the value of the deviation occurrence rate, formulate management guidelines, traverse and match the deviation occurrence rate with all preset deviation sending ranges to obtain the corresponding deviation sending range and generate the corresponding management instruction, and send the obtained management instruction to technicians at different levels for processing and analysis;
[0064] For example, if PF belongs to (Y3, Y4), the system generates a first-level management instruction;
[0065] If PF belongs to (Y2, Y3), the system generates a second-level management instruction;
[0066] If PF belongs to (Y1, Y2), the system generates a third-level management instruction; where Y1 < Y2 < Y3 < Y4, and Y1, Y2, Y3, Y4 all belong to (0, 1);
[0067] The management levels corresponding to the first-level management instruction, the second-level management instruction, and the third-level management instruction decrease in sequence.
[0068] In the embodiments of the present invention, by obtaining the historical recognition data of the target database and performing grouping and marking, the preprocessing data is received by the anomaly recognition processing and analysis module, and the anomaly recognition data is obtained. The anomaly recognition data calculation module performs data calculation on all anomaly recognition data to obtain the corresponding reliability and validity. The reliability of the anomaly recognition data is compared and analyzed with the preset reliability critical threshold to obtain the deviation data. The comprehensive evaluation management module performs comprehensive evaluation management on the deviation data, formulates the deviation data management evaluation criteria, and generates management instructions at different levels. In summary, the embodiments of the present invention involve data collection and analysis, result generation, and decision-making on optimization measures, and solve the problem of intelligent management of the deviation of recognition data in a data processing method based on multi-dimensional data. In actual situations, more data and context information may be required to make specific decisions and optimization plans.
[0069] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiments, since they are basically based on the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments.
[0070] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0071] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0072] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0073] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0075] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.
[0076] In conclusion, the above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A data processing method based on multidimensional data, characterized in that, include: Obtain historical identification data from the target database, clean the obtained historical identification data, and preprocess and filter the cleaned data to obtain abnormal identification data; Among these measures, data quality tools are used to automatically identify and correct data quality issues and to standardize data formats. Historical recognition data is acquired based on the historical recognition data acquisition module and grouped according to different types, including natural language data group, voice data group and image data group. Label each type of data set with Hj, j=1,2,3; The recognition rate of the statistical data set Hj; the recognition rate of the data set is labeled as SLj; Set the standard value Bj for the recognition rate of the data group; and calculate the recognition index ZS corresponding to the data group set using the formula ZS=(SLj / Bj)×100%; Plot the recognition index curve based on the calculated recognition index values, and analyze the trend of recognition index changes in historical recognition data. Obtain the identification index change curve corresponding to any set of identification data; determine the abnormal threshold value through the identification index change curve. Take the mean of all outlier thresholds in the data set; mark data below the mean as outlier identification data. The anomaly identification data are grouped and labeled, and the validity and reliability of each anomaly identification data are calculated to obtain the corresponding validity and reliability. Among them, the anomaly identification data are grouped according to the corresponding data type and labeled Ai, i=1,2,3, and the identification rate Si corresponding to the anomaly identification data is recorded in groups; For each anomaly identification data point, the validity is calculated to obtain the validity D of each anomaly identification data point; based on the validity of each anomaly identification data point, the reliability is further calculated to obtain the reliability XD of the anomaly identification data point. When calculating the validity of each anomaly identification data set, the normal response time Ti0 for each data set is determined; the response time Ti and identification rate Si of the identification process for each anomaly identification data set are statistically analyzed; and the formula is used to... Calculate the validity D for each anomaly identification data; Based on the validity of each anomaly detection data point, reliability calculations are performed, and the detection complexity Fi of the detection process for each anomaly detection data point is calculated; using the formula... Calculate the reliability XD of the obtained anomaly identification data; where α is the preset proportional coefficient of the reliability of the anomaly identification data, and 0 < α < 1; Deviation data is obtained by comparing and analyzing the reliability of anomaly identification data with a preset reliability threshold. The deviation data is used to calculate and analyze the deviation incidence rate, and deviation data management and evaluation criteria are dynamically formulated based on the analysis results; among these, the reliability of the deviation data is obtained and expressed using a formula. The deviation occurrence rate PF is calculated, where, These are variable constant parameters; XD0 is the preset reliability threshold. The deviation occurrence rate is evaluated, management criteria are formulated, the deviation occurrence rate is matched with all preset deviation sending ranges to obtain the corresponding deviation sending range, and corresponding management instructions are generated. The obtained management instructions are sent to technicians of different levels for processing and analysis.
Citation Information
Patent Citations
Image recognition method and system based on computer network and storage medium
CN117523299A
Impedance matching network adjusting method
CN117559937A