Data standard method and system based on artificial intelligence

Through an artificial intelligence-based data standard method, data density calculation and feature point analysis are used to generate prompt words for data screening and labeling, which solves the problems of data integration and search difficulties and achieves data search accuracy and system availability.

CN120804540APending Publication Date: 2025-10-17ZHOUPU DATA TECH NANJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510960208.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-11
Filing Date
2025-07-11
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

During the data processing process, the diversity and complexity of unstructured data make data integration and search difficult, and the existing technology lacks effective data standardization methods, resulting in a high search error rate and difficulty in meeting the needs of different users.

Method used

Through artificial intelligence-based data standard methods, data density calculation, feature point extraction, information range and location area analysis are used to generate prompt words, perform data screening and labeling, and ensure the precision and accuracy of the data.

Benefits of technology

It improves the accuracy of data search, reduces errors, enhances the system's usability and generalization capabilities, and adapts to changes in user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804540A_ABST
    Figure CN120804540A_ABST
Patent Text Reader

Abstract

The invention discloses a data standard method and system based on artificial intelligence, and belongs to the technical field of data annotation. The method comprises the following steps: collecting all data generated by work to form a data source library, extracting feature points of demand data, and constructing a to-be-labeled data set; further searching an information range of each kind of demand data, and setting cue words of the demand data according to the feature points and the information ranges of the demand data; judging whether to-be-labeled data exists in the collected data or not by utilizing the feature points of the demand data, further judging whether out-of-range data exists in the to-be-labeled data or not by utilizing the information range and the position area, and removing the out-of-range data; marking the screened data by using the cue word obtained by calculation to form a marked data set; performing accuracy calculation on the searched user demand data in real time, judging the accuracy of the annotation and performing early warning; and after early warning is carried out on the annotations, the annotations are optimized according to the demand data which are wrongly searched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data annotation technology, and specifically to a data standardization method and system based on artificial intelligence. Background Art

[0002] With the advancement of science and technology, vast amounts of data are constantly being generated across all industries. This data can come from multiple sources, such as sensors, social media, enterprise databases, and public datasets. This data is often unstructured and includes text, images, and videos. As the importance of data increases, privacy and security issues become more prominent. The protection and compliance of user data have become key issues. However, the generated data is often disorganized and not standardized. The existence of different data sources and formats complicates data integration and processing. When analyzing data, users need to accurately locate the data they need within this vast amount of data. In different fields and work environments, users may require different data standards. Therefore, determining different data standards based on their work environments and needs is crucial. Manually locating the required data within this chaotic data is undoubtedly difficult. Leveraging artificial intelligence to locate data based on user needs within this vast data set can significantly reduce work time. However, AI requires a standard for data retrieval. This standard often uses features within the data as criteria. However, due to the complexity and diversity of data, the location of a single feature can often result in different data. Avoiding errors caused by information location differences within the data is crucial. Summary of the Invention

[0003] The purpose of the present invention is to provide a data standard method and system based on artificial intelligence to solve the problems raised in the above background technology.

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: A data standardization method based on artificial intelligence, comprising the following steps: S100, collect all the data generated by the work to form a data source library, collect the required data found in the user's history, extract the feature points of the required data, and use the data types found by the user to construct a data set to be labeled; Furthermore, the specific steps for constructing the dataset to be labeled using the data types searched by the user are as follows: A data standardization method based on artificial intelligence, comprising the following steps: S101. Collect all data generated during the work, perform preliminary analysis on the data generated during the work, and extract the information interval of each data generated during the work. , represents the first, second, third, …, n information interval in each data generated in the work, n is a positive integer; first, the average of n different information intervals in each data is calculated, and then the average of m data generated in the work is calculated to obtain the data density, the formula is:

[0005] In the formula, J represents the data density generated in the work, G y represents the y information interval in each data extracted, G aver represents the average of n information intervals in a single data; n represents the number of information intervals in each data, m represents the number of data generated in the work; y belongs to 1 to n, x belongs to 1 to m; The calculated density is used to judge the information interval in each data generated in the work, when , it is judged that the data is missing information, and the data missing information is removed as error data; when , it is judged that the data is complete, and the data is usable data; All data generated in the work are judged to obtain all usable data, and the usable data is used to construct a data source library; S102, collect the demand data searched by the user in the user history as , represents the first, second, third, …, h demand data searched in the collected user history, h is a positive integer; the information amount of different information contained in each demand data is , represents the information amount of the first, second, third, …, p information contained in each demand data, the proportion of different information in each demand data in the total information is calculated, and the formula is: ; In the formula, α represents the proportion of each information in the total information in the demand data, X i represents the information amount of each information in the demand data, and p represents the number of information categories in the collected demand data; the proportion of each information in the demand data is calculated, the proportions of p information are compared, and the information category with the largest proportion is selected as the feature point T of the demand data; S103, the feature points of the demand data searched in the collected user history are calculated, and the feature point of each demand data is , represents the feature point of the first, second, third, …, h demand data searched in the collected history, the feature points of the h demand data are compared, the demand data with the same feature point is selected as the same type of demand data, and finally the demand data searched by the user is classified as , The first, second, third,..., g kinds of demand data searched by the user are represented, and the data set contained in each kind of demand data is taken as a to-be-labeled data set, to obtain g to-be-labeled data sets. Feature points of the data searched by the user are extracted, the data are classified by using the feature points, and the to-be-labeled data sets are generated. When the data are labeled, only the data in the to-be-labeled data sets are labeled, so that the labeling workload of the system is reduced; S200, according to the feature points of the demand data, all the demand data in the history are collected, each kind of demand data collected is analyzed, the information range of each kind of demand data is further searched and the position region is divided, the prompt word of the demand data is set according to the feature points of the demand data and the information range, and the information range represents the information types contained in the demand data; Further, the specific steps of setting the prompt word of the demand data according to the feature points of the demand data and the information range are as follows: S201, the demand data searched by the user in the history is collected, the information of each kind of demand data is extracted, the types of the information in each kind of demand data extracted are set as p kinds, and the information types of the different information contained in each kind of demand data extracted in S102 are set as the information range of each kind of demand data; S202, after the information range of each kind of demand data is determined, the positions of the information existing in each data in the data source library are set as , The first, second, third,..., q positions of the information existing in each data are represented, q is a positive integer, the positions of the different types of information in each kind of demand data in the history are extracted and arranged from front to back as , The first, second, third,..., p positions of the information in the demand data are represented, and the information position chain of the demand data in the complete data in the data source library is constructed as , The different positions of the first, second, third,..., p kinds of information in the demand data are represented, The complete data represents one data in the data source library. The positions in the demand data are regionally divided, the divided regions are calculated, and the formula is as follows: ; In the formula, Lg1 represents the first division node of the information position in the demand data, Lg2 represents the second division node of the information position in the demand data, W p represents the last information position in the demand data, represents the first information position in the demand data; according to the two kinds of division nodes calculated, the information positions in the demand data are divided into , and ; the information position regions Set the prompt position as TO, and set the information position area as Set the prompt position as ZH, and set the information position area as Set the prompt position as WE; divide the information position of each demand data into areas, and obtain different information categories contained in the three areas respectively; S203, after dividing the information position of the demand data into areas, set the prompt word as according to the characteristic point of each demand data and the information position area; the mark word contains the characteristic point of each demand data, the number of information categories and the area division; set the prompt word as for g kinds of demand data. , which represents the prompt word set for the 1st, 2nd, 3rd,..., gth demand data.

[0006] By analyzing the information range and position area in the data, the prompt word generated by the characteristic point, information range and position area can make the labeled data more accurate, greatly reduce the error of searching, and avoid the error data caused by different information positions; S300, collect the data in the data source library, and use the characteristic point of the demand data to determine whether there is to-be-labeled data in the collected data, and then further use the information range and position area to determine whether there is data beyond the range in the to-be-labeled data, and remove the data beyond the range; Further, the specific steps of removing the data beyond the range are as follows: S301, use the characteristic point of each demand data in the to-be-labeled data to perform a first step of screening on the data in the data source library, and extract the data with the same characteristic point from the data source library and divide them into the corresponding to-be-labeled data set; S302, after obtaining the data in the to-be-labeled data set, use the information range of each demand data to perform a second step of screening on the information categories contained in the data in the corresponding to-be-labeled data set, and extract the real-time information contained in the data in the to-be-labeled data set as . , which represents the 1st, 2nd, 3rd,..., cth information contained in the data in the to-be-labeled data set, c is a positive integer; after judging the data in the to-be-labeled data set by using the information range of each demand data, the data in the to-be-labeled data set whose information categories contain all kinds of information in the information range are . , which represents the 1st, 2nd, 3rd,..., fth data in the to-be-labeled data set containing all kinds of information in the information range. S303, after screening out the data containing all kinds of information in the information range, the third step of screening is carried out by dividing the region according to the information position of each demand data, the information position of the data screened out in the second step is judged, and all information positions in the data which are the same as the information range are extracted as , The positions of the first, second, third,..., and pth information in the data screened in the second step are indicated, and the position region (TO, ZH, WE) of each demand data is used to judge whether the position of each information in the data is within the corresponding position region. When the position of the information in the data exceeds the corresponding position region, it is judged as out-of-range data and is removed.

[0007] S400, after judging the collected data, the screened data is labeled using the calculated prompt words to form a labeled data set; The three-step screening of data using feature points, information range and position region makes the screened data more accurate, and the position region allows the data to be further refined. When the user searches for data, the wrong information can be removed according to different position regions; Further, the specific steps of labeling the screened data using the calculated prompt words to form a labeled data set are as follows: S401, after the three-step screening in S300, each kind of data to be labeled is input in the data source library, and each kind of data to be labeled is obtained; S402, the screened each kind of data to be labeled is labeled using the prompt words set in S203, specifically , The first, second, third,..., and vth data after screening are labeled; then the user searches for all data using the label of each data, and obtains the data required by the user.

[0008] S500, the labeled data set is used to search for the demand data of the user, the accuracy of the searched demand data of the user is calculated in real time, the accuracy of the label is judged and a warning is given; Further, the specific steps of judging the accuracy of the label and giving a warning are as follows: S501, after the user searches for data using data labeling, the user manually checks the searched data during use, and the data searched incorrectly is , The first, second, third,..., and uth data searched incorrectly by the user are indicated, and u is a positive integer. The records of the use of data by the user in the history are collected, the number of error data in the records is extracted, the accuracy threshold of data searching is calculated, and the formula is: ; In the formula, accy represents the calculated accuracy threshold, Y_sc t represents the number of error data extracted in the tth record, E represents the number of records in the collected history that are damaged by the use of data, and st_c represents the standard deviation of the number of error data. S502, using the calculated accuracy threshold to judge the accuracy Sc of each user data search, the accuracy is the number of error data search, when Sc≥accy, it is judged that the data search process has defects and a warning is issued, when Sc<accy, it is judged that the data search process has no defects.

[0009] S600, after warning the label, the label is optimized according to the demand data of finding errors.

[0010] Further, the specific steps of optimizing the label according to the demand data of finding errors are: S601, when judging the result of user data search, it is found that the search process has defects and a warning is issued, the information category in the error data is extracted, the information position in the information range of the error data is counted, and the maximum position W max and the minimum position W min are selected. S602, input the information position in the information range of the error data into the position area in each demand data, optimize the position area in the demand data by using the information position in the information range of the error data, replace the first information position W1 with the minimum position W min , replace the last information position W p with the maximum position W max , and the new position area obtained after replacement is (W min , Lg1), (Lg1, Lg2) and (Lg2, W max ).

[0011] After real-time judgment and warning of the user search result, the search result is optimized, so that the system can update with the user's work scene and data volume, and the usability and generalization of the system are increased; A data standard system based on artificial intelligence, the data standard system comprising a data collection module, a to-be-labeled data set module, a prompt word generation module, a screening module, a labeling module, an accuracy judgment module and an optimization module; The data collection module is used for collecting the demand data of user search in the history, calculating the data tightness, and pre-processing the collected data to construct a data source library. The to-be-labeled data set module is used for analyzing the data generated in the history, extracting the feature points of each demand data, and constructing a to-be-labeled data set. The prompt word generation module is used for collecting information categories of the demand data to obtain an information range, and analyzing positions of information in the demand data to calculate position regions of each kind of demand data; The screening module is used for screening data in the data source library by using the feature points, the information range and the position regions respectively to obtain the to-be-labeled data; The labeling module is used for labeling the to-be-labeled data by using the prompt words generated by the prompt word generation module; The accuracy judgment module is used for judging the data searched by the user, analyzing whether there is a defect in the searching process and giving a warning; The optimization module is used for optimizing the position regions of the demand data by using information positions of the error data in the searching result when it is found that there is a defect in the searching process.

[0012] The prompt word generation module comprises an information range unit and a position region unit; The information range unit is used for collecting information categories in the demand data to obtain an information range of the demand data; The position region unit is used for extracting positions of each kind of information in the demand data, determining a position range of the information range of the demand data after calculating all the information positions, and dividing the position range.

[0013] The accuracy judgment module comprises an accuracy threshold calculation unit and a judgment unit; The accuracy threshold calculation unit is used for collecting records of losses caused by the user's use data in the history, extracting the number of error data, and calculating an accuracy threshold; The judgment module is used for judging the real-time searching result of the user by using the accuracy threshold, and giving a warning when it is found that there is a defect in the searching process.

[0014] Compared with the prior art, the present application has the following beneficial effects: The present application pre-processes the collected data source library when the user searches data, removes missing data in the data source library, avoids the influence of the missing data on the user's searching, and reduces system error.

[0015] The present application labels data according to prompt words when the user searches data, the prompt words are determined by using data feature points, information ranges and position regions, the searching result is more accurate, and the change of information positions does not cause error of the searched data.

[0016] The present application judges and optimizes the searching process in real time, the searching process can be updated with the user's working environment and data volume, and the usability and generalization of the system are increased. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings are included to provide a further understanding of the application, and are incorporated in and constitute a part of this specification, illustrate embodiments of the application, and together with the description serve to explain the application, and do not limit the application. In the drawings: Fig. 1 is a module distribution diagram of a data standard system based on artificial intelligence according to the application; Fig. 2 is a step schematic diagram of a data standard method based on artificial intelligence according to the application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application.

[0019] Please refer to Figs. 1-2 , the application provides technical solutions: A data standard method based on artificial intelligence, the method comprises the following steps: S101, collecting all data generated in work, pre-analyzing the data generated in work, and extracting information intervals in each data generated in work , representing the first, second, third,..., n information intervals in each data generated in work, n is a positive integer; first, averaging the n different information intervals in each data, and then averaging m data generated in work, to obtain data tightness, the formula is:

[0020] In the formula, J represents the data tightness generated in work, G y represents the yth information interval extracted in each data, G aver represents the average value of n information intervals in a single data; n represents the number of information intervals in each data, and m represents the number of data generated in work; y belongs to 1 to n, and x belongs to 1 to m; Using the calculated tightness to judge the information intervals in each data generated in work, when , it is judged that the data is missing information, and the data missing information is removed as error data; when , it is judged that the information in the data is complete, and the data is usable data; All data generated in work are judged to obtain all usable data, and the usable data is used to construct a data source library; S102, collect the demand data searched by the user in the user history as , indicates the first, second, third,..., hth demand data searched in the collected user history, h is a positive integer; the information amount of different information contained in each demand data is , indicates the information amount of the first, second, third,..., pth information contained in each demand data, the proportion of different information in each demand data in the total information is calculated, and the formula is:

[0021] In the formula, a represents the proportion of each information in the total information in the demand data, X i represents the information amount of each information in the demand data, and p represents the number of information categories in the collected demand data; the proportion of each information in the demand data is obtained by calculation, the proportions of the p information categories are compared, and the information category with the largest proportion is selected as the feature point T of the demand data; S103, the feature points of the demand data searched in the collected user history are calculated, and the feature point of each demand data is obtained , indicates the feature points of the first, second, third,..., hth demand data searched by the user in the collected history, the feature points of the h demand data are compared, the demand data with the same feature points are selected as the same type of demand data, and finally the demand data searched by the user is classified as , indicates the first, second, third,..., gth demand data searched by the user, and the data set contained in each demand data is taken as a to-be-labeled data set, and g to-be-labeled data sets are obtained. The feature points of the data searched by the user are extracted, the data is classified by using the feature points, and the to-be-labeled data set is generated. When the data is labeled, only the data in the to-be-labeled data set is labeled, so that the labeling workload of the system is reduced; S200, according to the feature points of the demand data, all the demand data in the history is collected, each demand data collected is analyzed, the information range of each demand data is further searched and the position area is divided, the prompt word of the demand data is set according to the feature points of the demand data and the information range, and the information range represents the information category contained in the demand data; The specific steps of setting the prompt word of the demand data according to the feature points of the demand data and the information range are as follows: S201, collect the demand data searched by the user in the history, extract the information of each demand data, set the category of the information extracted in each demand data as p, and set the information category of different information contained in each demand data extracted by the S102 as the information range of each demand data; S202、In determining the information range of each demand data, the location of the information in each data in the data source library is , represents the first, second, third,..., q location of information in each data, q is a positive integer; the location of different types of information in the demand data in the history is extracted and arranged from front to back as , represents the location of the first, second, third,..., p type of information in the demand data, and the information location chain of the demand data in the complete data in the data source library is constructed as , represents the different locations of the first, second, third,..., p type of information in the demand data, ; Complete data represents a data in the data source library; the location of the demand data is divided into regions, and the division region is calculated, and the formula is:

[0022] In the formula, Lg1 represents the first division node of the information location in the demand data, Lg2 represents the second division node of the information location in the demand data, W p represents the last information location in the demand data, represents the first information location in the demand data; according to the two division nodes calculated, the information location in the demand data is divided into 、 and ; The information location region is set to the prompt location TO, the information location region is set to the prompt location ZH, and the information location region is set to the prompt location WE; The information location of each demand data is divided into regions, and different information types contained in the three regions are obtained; S203、After the information location in the demand data is divided into regions, the prompt word is set according to the characteristic point and information location region of each demand data as , The flag word contains the characteristic point, information type number and region division of each demand data; the prompt word is set for g kinds of demand data as , represents the prompt word set for the first, second, third,..., g kinds of demand data.

[0023] By analyzing the information range and location region in the data, the prompt word generated by the characteristic point, information range and location region can make the labeled data more accurate, greatly reduce the error of searching, and avoid the error data caused by different information locations; S300, collect data in the data source library, judge whether there is to-be-labeled data in the collected data by using the feature points of the demand data, and then further judge whether there is data beyond the range in the to-be-labeled data by using the information range and the position area, and remove the data beyond the range; The specific steps of removing the data beyond the range are as follows: S301, first-step filtering of data in the data source library by using the feature points of each kind of demand data in the to-be-labeled data, and extracting data with the same feature points from the data source library and dividing them into corresponding to-be-labeled data sets; S302, after obtaining the data in the to-be-labeled data set, second-step filtering of the information types contained in the data in the corresponding to-be-labeled data set by using the information range of each kind of demand data, and extracting real-time information contained in the data in the to-be-labeled data set as , , wherein the first, second, third, …, and cth information contained in the data in the to-be-labeled data set, and c is a positive integer; after judging the data in the to-be-labeled data set by using the information range of each kind of demand data, the data in the to-be-labeled data set whose information types contain all kinds of information in the information range are , , wherein the first, second, third, …, and fth to-be-labeled data set containing all kinds of information in the information range; S303, after filtering out the data containing all kinds of information in the information range, third-step filtering of the data by using the information position division area of each kind of demand data, judging the information position contained in the data filtered out in the second step, and extracting all information positions in the data which are the same as the information range as , , wherein the first, second, third, …, and pth information position in the data filtered out in the second step, and whether the position of each information in the data is in the corresponding position area by using the position area (TO, ZH, WE) of each kind of demand data, when the position of the information in the data is beyond the corresponding position area, the data beyond the range is judged and removed.

[0024] S400, after judging the collected data, labeling the filtered data by using the calculated prompt words to form a labeled data set; The three-step filtering of data by using the feature points, the information range, and the position area makes the filtered data more accurate, and the position area allows the data to be further refined, so that when the user searches for data, the wrong information can be removed according to different position areas; The specific steps of labeling the filtered data by using the calculated prompt words to form a labeled data set are as follows: S401, after the three-step screening in S300, input each data set to be labeled into the data source library to obtain each type of data to be labeled; S402, using the prompt words set in S203 to mark each type of data to be marked after screening, specifically , It indicates the annotation of the 1st, 2nd, 3rd, ... vth data after filtering; then the user uses the annotation of each data to search all the data and obtain the data the user needs.

[0025] S500: Searching for user demand data using the annotated data set, performing accuracy calculation on the searched user demand data in real time, determining the accuracy of the annotations, and issuing an early warning; The specific steps to determine the accuracy of the labeling and issue an early warning are: S501. After the user searches for data using data annotation, the user manually checks the searched data during use and obtains the data with search errors. , Indicates the 1st, 2nd, 3rd, ..., uth data found incorrectly after the user search, where u is a positive integer. Collect the records of data loss caused by the user's use of history, extract the number of erroneous data in the records, and calculate the accuracy threshold of data search. The formula is:

[0026] In the formula, accy represents the calculated accuracy threshold, Y_sc t represents the number of erroneous data in the t-th record extracted, E represents the number of records in the collected history where data was used to cause damage, and st_c represents the standard deviation of the number of erroneous data; S502, using the calculated accuracy threshold to judge the accuracy Sc after each user searches for data, the accuracy is the number of wrong data found, when Sc≥accy, it is judged that there is a defect in the search process and an early warning is issued, when Sc <accy时,判断查找数据过程不存在缺陷。

[0027] S600: After the warning is issued for the annotation, the annotation is optimized according to the demand data for finding errors.

[0028] The specific steps to optimize annotations based on the required data for finding errors are: S601: When the result of the user's search data is judged and it is found that there is a defect in the search process and an early warning is issued, the information type in the error data is extracted, the information position within the information range in the error data is counted, and the maximum position W is selected. max and minimum position W min ; S602, input the information position in the information range of the error data into the position area in each demand data, optimize the position area in the demand data by using the information position in the information range of the error data, replace the first information position W1 with the minimum position W min Replace the last information position W p with the maximum position W max , and the new position area obtained after replacement is (W min , Lg1), (Lg1, Lg2) and (Lg2, W max ).

[0029] After real-time judgment and early warning of the user search result, the search result is optimized, so that the system can update with the user's work scene and data volume, and the usability and generality of the system are increased; A data standard system based on artificial intelligence, the data standard system comprising a data collection module, a to-be-labeled data set module, a prompt word generation module, a screening module, a labeling module, an accuracy judgment module and an optimization module; The data collection module is used to collect demand data of user search in history, calculate data tightness, and pre-process collected data to construct a data source library; The to-be-labeled data set module is used to analyze data generated in history, extract feature points of each demand data, and construct a to-be-labeled data set; The prompt word generation module is used to collect information ranges of information categories of demand data, analyze positions of information in demand data, and calculate position areas of each demand data; The screening module is used to screen data in the data source library by using feature points, information ranges and position areas respectively, and obtain to-be-labeled data; The labeling module is used to label to-be-labeled data by using prompt words generated by the prompt word generation module; The accuracy judgment module is used to judge data searched by a user, analyze whether there is a defect in the search process and give an early warning; The optimization module is used to optimize position areas in demand data by using information positions in error data in search results when a defect is found in the search process.

[0030] The prompt word generation module comprises an information range unit and a position area unit; The information range unit is used to collect information categories in demand data, and obtain information ranges of demand data; The position area unit is used to extract positions of each information in demand data, determine position ranges of information ranges in demand data after calculating all information positions, and divide the position ranges.

[0031] The accuracy judgment module comprises an accuracy threshold calculation unit and a judgment unit; The accuracy threshold calculation unit is configured to collect records of user use data causing loss in history, extract the number of error data therein, and calculate an accuracy threshold; The judgment module is configured to judge the real-time search result of the user by using the accuracy threshold, and issue a warning when defects exist in the search process.

[0032] Embodiment 1: A supply chain system of a certain chain supermarket generates a large amount of order data (such as supplier name, commodity code, delivery time, etc.) every day, and the user needs to quickly extract specific type orders; Collect historical order data, calculate data tightness, remove missing data to build a data source library; extract the demand data of "fresh emergency order" in history as the user query, calculate the proportion of all historical "fresh emergency order" information extracted to obtain the largest proportion of "fresh", and take "fresh" as a feature point; calculate the feature points of the remaining demand data, which can divide the demand data into fresh and vegetable categories according to the feature points; build a data set to be labeled; Extract all information types contained in the "fresh emergency order" demand data to set the information range, including fresh, supplier, shelf life and temperature layer; The information position of the "fresh emergency order" demand data in history is 【1-19】, and three information areas are divided, TO=【1-7】, ZH=【7-13】, and WE=【13-19】; The information types contained in each information area are TO including product category and supplier, ZH including shelf life, and WE including temperature layer; build the prompt word as Mark: fresh→(

product category and supplier

shelf life

temperature layer

[0033] Use the prompt word to label the screened fresh data; When the user queries in real time, the "fresh emergency order" system outputs and displays all the labeled fresh data according to the prompt word.

[0034] It is to be noted that, in the present text, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0035] Finally, it should be noted that the above-mentioned only constitutes the preferred embodiments of the present application and is not intended to limit the present application. Although the present application has been described in detail with reference to the foregoing embodiments, it will be apparent to those skilled in the art that modifications, equivalent replacements, improvements and the like of the technical solutions described in the foregoing embodiments can still be made. Any modifications, equivalent replacements, improvements and the like made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A data standardization method based on artificial intelligence, characterized by: The method includes the following steps: S100. Collect all the data generated during work to form a data source library, collect the required data searched in the user's history, extract the characteristic points of the required data, and construct a dataset to be labeled using the types of data searched by the user; S200. Collect all the required data in history according to the characteristic points of the required data, analyze each type of collected required data, further search for the information range of each type of required data and divide the location area, and set the prompt words for the required data according to the characteristic points and information range of the required data, where the information range represents the types of information included in the required data; S300. Collect the data in the data source library, use the characteristic points of the required data to judge whether there is data to be labeled in the collected data, and then further use the information range and location area to judge whether there is data beyond the range in the data to be labeled, and remove the data beyond the range; S400. After judging the collected data, use the calculated prompt words to label the filtered data to form a labeled dataset; S500. Use the labeled dataset to search for the required data of the user, calculate the accuracy of the searched user required data in real time, judge the accuracy of the label and give an early warning; S600. After giving an early warning for the label, optimize the label according to the required data with search errors.

2. The artificial intelligence-based data standardization method according to claim 1, characterized in that: The specific steps of constructing the dataset to be labeled using the types of data searched by the user in S100 are as follows: S101. Collect all data generated during the work, perform preliminary analysis on the data generated during the work, and extract the information interval of each data generated during the work. , It represents the 1st, 2nd, 3rd, ..., nth information intervals generated in each data during the extraction process, where n is a positive integer. First, the n different information intervals in each data are averaged, and then the m data generated during the process are averaged to obtain the data density. The formula is: ; In the formula, J represents the density of data generated during work, G y Indicates the yth information interval in each extracted data, G aver It represents the average value of n information intervals in a single data; n represents the number of information intervals in each data, m represents the number of data generated in the work; y belongs to 1 to n, and x belongs to 1 to m; The information interval of each data generated in the work is judged by the calculated tightness. When , it is determined that there is a lack of information in the data, and the data with missing information is treated as erroneous data and removed; when When , it is judged that the information in the data is complete and the data is usable data; Judge all the data generated during work to obtain all the available data, and use the available data to construct a data source library; S102: Collect the user's search demand data in the user history. , It represents the 1st, 2nd, 3rd, ..., hth demand data found in the collected user history, where h is a positive integer; the amount of information contained in each demand data is , The information volume of the first, second, third, ..., p types of information contained in each demand data is expressed as follows: ; In the formula, α represents the proportion of each type of information in the total information in the demand data, X i represents the amount of information of each type in the demand data, and p represents the number of information types in the collected demand data. After calculating the proportion of each type of information in the demand data, the proportions of the p types of information are compared, and the type of information with the largest proportion is selected as the characteristic point T of the demand data. S103, calculate the characteristic points of the demand data found in the collected user history, and obtain the characteristic point of each demand data as follows: , Indicates the feature points of the 1st, 2nd, 3rd, ..., hth demand data found by the user in the collected history. The feature points of the h demand data are compared, and the demand data with the same feature points are selected as the same type of demand data. Finally, the demand data found by the user are classified into , It represents the 1st, 2nd, 3rd, ..., gth types of required data that the user searches for, and the data sets contained in each type of required data are used as the data sets to be labeled, thereby obtaining g types of data sets to be labeled.

3. The artificial intelligence-based data standardization method according to claim 2, characterized in that: The specific steps of setting the prompt words for the required data according to the characteristic points and information range of the required data in S200 are as follows: S201. Collect the required data searched by the user in history, extract the information of each type of required data, set the number of types of information in each type of required data extracted as p types, and set the types of information of different information included in each type of required data extracted in S102 as the information range of each type of required data; S202, after determining the information scope of each required data, set the location of the information in each data in the data source library as , Indicates the 1st, 2nd, 3rd, ..., qth position of information in each data, where q is a positive integer; extract the position of different types of information in each demand data in the history and arrange them from front to back as , Indicates the position of the first, second, third, ...p types of information in the demand data, and constructs the information position chain of the demand data in the complete data within the data source library as , Indicates the different positions of the 1st, 2nd, 3rd, ..., pth type of information in the demand data, ; Complete data represents a data in the data source library; the location in the required data is divided into regions, and the division region is calculated. The formula is: ; In the formula, Lg1 represents the first partition node of the information position in the demand data, Lg2 represents the second partition node of the information position in the demand data, and W p Indicates the last information position in the demand data. Indicates the first information position in the demand data; the information position in the demand data is divided into 、 and ; Information location area Set the prompt position to TO, and change the information position area Set the prompt location to ZH, and change the information location area Set the prompt position to WE; divide the information position of each required data into regions, and obtain different types of information contained in three regions; S203, after dividing the information location in the demand data into regions, set prompt words according to the feature points and information location regions of each demand data. The flag words include the characteristic points, number of information types and regional divisions of each demand data; for g types of demand data, the prompt words are set as , Indicates the prompt words set for the 1st, 2nd, 3rd, ..., gth types of demand data.

4. The artificial intelligence-based data standardization method according to claim 3, characterized in that: The specific steps of removing the data beyond the range in S300 are as follows: S301. Use the characteristic points of each type of required data in the data to be labeled to perform the first-step screening on the data in the data source library, extract the data with the same characteristic points in the data source library and divide them into the corresponding datasets to be labeled; S302: After obtaining the data in the dataset to be labeled, the information type contained in the corresponding dataset to be labeled is screened using the information range of each required data, and the real-time information contained in the data in the dataset to be labeled is extracted. , Indicates the 1st, 2nd, 3rd, ..., cth types of information contained in the data set to be labeled, where c is a positive integer. After using the information range of each required data to judge the data in the data set to be labeled, the data in the data set to be labeled that contains all types of information within the information range is obtained. , Represents the data in the 1st, 2nd, 3rd, ...fth datasets to be labeled that contain all types of information within the information range; S303: After filtering out the data containing all types of information within the information range, the third step of filtering is performed using the information location of each required data to divide the area, determine the information location contained in the data filtered in the second step, and extract all information locations in the data that are the same as the information range. , Indicates the positions of the 1st, 2nd, 3rd, ..., pth types of information in the data filtered in the second step. Use the location area (TO, ZH, WE) of each required data to determine whether the position of each information in the data is within the corresponding location area. When the position of the information in the data exceeds the corresponding location area, it is determined as out-of-range data and is removed.

5. The data standardization method based on artificial intelligence according to claim 4, characterized in that: The specific steps of using the calculated prompt words to label the filtered data to form a labeled dataset in S400 are as follows: S401. After three-step screening in S300, input each dataset to be labeled in the data source library to obtain each data to be labeled; S402, using the prompt words set in S203 to mark each type of data to be marked after screening, specifically , It indicates the annotation of the 1st, 2nd, 3rd, ... vth data after filtering; then the user uses the annotation of each data to search all the data and obtain the data the user needs.

6. The artificial intelligence-based data standardization method according to claim 5, characterized in that: The specific steps of judging the accuracy of the label and giving an early warning in S500 are as follows: S501. After the user searches for data using data annotation, the user manually checks the searched data during use and obtains the data with errors. , Indicates the 1st, 2nd, 3rd, ..., uth data found incorrectly after the user search, where u is a positive integer. Collect the records of data loss caused by the user's use of history, extract the number of erroneous data in the records, and calculate the accuracy threshold of the data search. The formula is: ; In the formula, accy represents the calculated accuracy threshold, Y_sc t represents the number of erroneous data in the t-th record extracted, E represents the number of records in the collected history where data was used to cause damage, and st_c represents the standard deviation of the number of erroneous data; S502. Use the calculated accuracy threshold to judge the accuracy Sc after each user searches for data. The accuracy is the number of wrong search data. When Sc≥accy, judge that there are defects in the data search process and give an early warning. When Sc<accy, judge that there are no defects in the data search process.

7. The artificial intelligence-based data standardization method according to claim 6, characterized in that: The specific steps of optimizing the label according to the required data with search errors in S600 are as follows: S601: When the result of the user's search data is judged and it is found that there is a defect in the search process and an early warning is issued, the information type in the error data is extracted, the information position within the information range in the error data is counted, and the maximum position W is selected. max and minimum position W min ; S602: Input the information position within the information range of the error data into the position area of ​​each demand data, optimize the position area of ​​the demand data using the information position within the information range of the error data, and replace the first information position W1 with the minimum position W min Replace, change the last information position W p Use the maximum position W max After replacement, the new location area is (W min , Lg1), (Lg1, Lg2) and (Lg2, W max ).

8. A data standard system based on artificial intelligence, characterized by: The data standard system includes a data collection module, a dataset to be annotated module, a prompt word generation module, a screening module, an annotation module, an accuracy judgment module, and an optimization module; The data collection module is used to collect the demand data searched by users in history, calculate the data density, and pre-process the collected data to build a data source library; The dataset to be annotated module is used to analyze the data generated by historical work, extract the feature points of each required data, and construct the dataset to be annotated; The prompt word generation module is used to collect information types of demand data to obtain information ranges, analyze the positions of information in the demand data, and calculate the position areas of each type of demand data; The screening module is used to screen the data in the data source library using feature points, information range and location area to obtain data to be labeled; The annotation module is used to annotate the data to be annotated using the prompt words generated by the prompt word generation module; The accuracy judgment module is used to judge the data searched by the user, analyze whether there are defects in the search process and issue an early warning; The optimization module is used to optimize the location area in the demand data by using the information location in the erroneous data in the search result after discovering that there is a defect in the search process.

9. The artificial intelligence-based data standard system according to claim 8, characterized in that: The prompt word generation module includes an information range unit and a position area unit; The information range unit is used to collect the information types in the demand data to obtain the information range of the demand data; The location area unit is used to extract the location of each information in the demand data, calculate all information locations, determine the location range of the information range in the demand data, and divide the location range.

10. The artificial intelligence-based data standard system according to claim 8, characterized in that: The accuracy judgment module includes an accuracy threshold calculation unit and a judgment unit; The accuracy threshold calculation unit is used to collect historical records of losses caused by user data usage, extract the number of erroneous data therein, and calculate the accuracy threshold; The judgment module is used to judge the user's real-time search results using the accuracy threshold, and issue an early warning when it is judged that there is a defect in the search process.