Intelligent equipment visual data management and AI model development platform and method
Through the intelligent equipment visual data governance and AI model development platform, using data labeling, model training and version management methods, we solved the problems of insufficient utilization of multi-dimensional data and decreased model accuracy after incremental training, and achieved efficient forgetting management and model performance improvement.
Patent Information
- Application Number
- CN202510855119.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing technologies lack effective utilization of multidimensional data in the visual data management of intelligent equipment, resulting in insufficient improvement in algorithm accuracy and a decrease in model recognition accuracy after incremental training, making it difficult to effectively manage historical versions to locate problems.
Provides a smart equipment visual data governance and AI model development platform, including data labeling modules, model training modules, and version management modules. Through manual and automatic labeling, data enhancement, historical version management, etc., it optimizes model parameter changes, realizes forgetting management, and determines the optimization requirement type within the version range.
It achieves efficient forgetting management of training risk data, reduces storage space requirements, improves model recognition accuracy, and optimizes training requirement types based on differences within version ranges and changes in model parameters, thereby improving the overall performance of the model.
Smart Images

Figure CN120356038B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to a platform and method for intelligent equipment visual data management and AI model development. Background Art
[0002] The existing intelligent algorithm training and integrated verification systems for scientific research systems and smart security have too simple a data fusion method. Currently, algorithm processing is generally performed on a single sensor, some algorithms are integrated at the feature level, and finally the results of the algorithm processing are integrated at the decision level. This processing method is relatively simple, lacks effective use of multi-dimensional data, does not fully explore data correlation information, and thus cannot bring about a guaranteed improvement in algorithm accuracy.
[0003] Therefore, in order to solve the above technical problems, the existing technical solutions obtain and label incremental data and then perform incremental training processing. However, the existing technical solutions have the following problems:
[0004] In the process of visual data management, existing technical solutions often generate training data through automatic labeling. This inevitably leads to a deterioration in the recognition accuracy of certain targets after incremental training. Therefore, how to manage historical versions in a targeted manner so as to quickly and conveniently locate problems arising from visual training has become a technical problem that needs to be solved urgently.
[0005] To solve the above technical problems, this application provides a platform and method for intelligent equipment visual data management and AI model development. Summary of the Invention
[0006] To achieve the purpose of the present invention, the present invention adopts the following technical solutions:
[0007] Specifically, this application provides a smart equipment visual data management and AI model development platform, specifically including:
[0008] Data annotation module, model training module, version management module;
[0009] The data annotation module is responsible for performing data annotation processing to obtain annotated data;
[0010] The model training module is responsible for using the labeled data to perform training processing of the visual model as needed;
[0011] The version management module is responsible for managing different historical versions of the visual model using the training data and the updated training status of the training data in the visual model.
[0012] A further technical solution is that the data annotation processing includes manual annotation and automatic annotation.
[0013] A further technical solution is to further include a data enhancement module responsible for amplifying the labeled data.
[0014] A further technical solution is that the amplification processing includes rotation, cropping, flipping, adversarial generation network, real visual background synthesis of moving targets, and amplification processing methods of real image + simulated visual interference synthesis.
[0015] A further technical solution is to support visual display of changes in model parameters of the visual model during training of the visual model.
[0016] A further technical solution is to manage different historical versions of the visual model, specifically including:
[0017] When the visual model of the training data of the historical version has not been incrementally trained within the most recent preset time period, the need for forgetting management is determined by comparing the incrementally trained visual model with the verification result of the optimized target version in the target database. When the deviation between the accuracy of the verification result of the optimized target version and the visual model is greater than the preset value of the deviation, forgetting processing of the optimized target version is performed.
[0018] In a second aspect, this application provides a historical version management method, which is applied to the above-mentioned intelligent equipment visual data management and AI model development platform, specifically including:
[0019] S1 determines the composition of automatically labeled data under different types of training data based on the training data of the visual model, and determines risk training data in the training data based on the composition;
[0020] S2 determines the composition of training risk data during incremental training of different historical versions, and determines the optimization target version in the historical version based on the similarity between the training risk data and other historical versions in the later period;
[0021] S3 divides historical versions into different version intervals based on training risk data. Based on the composition data of historical versions in different version intervals and the changes in model parameters between the optimization target version and later historical versions, it determines the type of optimization requirements in different version intervals.
[0022] S4 When the optimization requirement types of the optimization target version in different version intervals do not belong to the target requirement types, the forgetting management method of the optimization target version is determined based on the updated training status of the training risk data corresponding to the optimization target version in the visual model and the optimization requirement types in different version intervals.
[0023] The beneficial effects of the present invention are:
[0024] Based on the composition data of historical versions in different version intervals and the changes in model parameters of the optimization target version and later historical versions, the optimization requirement types in different version intervals are determined. That is, the differences in problem location requirements when there are problems in the training risk data corresponding to the version intervals due to the differences in the number of historical versions in different version intervals are taken into account. At the same time, the differences in problems in the training risk data corresponding to the version intervals due to the changes in model parameters of later historical versions are also taken into account, thereby realizing the determination of optimization requirement types from the perspective of location requirements.
[0025] The forgetting management method of the optimized version is determined based on the updated training status of the training risk data in the visual model corresponding to the optimized target version and the optimization requirement types in different version intervals. This ensures the efficiency of forgetting processing of the optimized target version with greater forgetting processing requirements, while also reducing the demand for storage space.
[0026] A further technical solution is that the composition of the automatic annotation of the training data includes the data volume and data volume ratio of the automatically annotated training data during different incremental trainings.
[0027] Specifically, the type of the training data is determined according to the recognition target corresponding to the training data.
[0028] A further technical solution is that the method for determining the risk training data in the training data is:
[0029] Determining the amount of automatically labeled training data of the type under different times of incremental training based on the composition of the training data of the type during the incremental training of the visual model in different historical versions;
[0030] Determining the proportion of data volume under different incremental training times according to the proportion of data volume of the automatically labeled training data of the type under different incremental training times;
[0031] Based on the proportion of data volume under different incremental training times, it is determined whether the training data is risky training data.
[0032] A further technical solution is that the method for determining the forgetting management method of the optimized target version is:
[0033] Using the training risk data corresponding to the optimization target version as matching risk data, and determining historical versions corresponding to different types of matching risk data during incremental training within a recent preset period based on the updated training status of the matching risk data in the visual model, and using the historical versions as updated versions;
[0034] Based on the optimization requirement types within different version intervals, determine the version interval belonging to the second type of requirement and use it as the second type of requirement interval;
[0035] According to the updated versions of the matching risk data of different types and the constituent data of the second type of demand intervals, the forgetting management method of the optimized target version is determined.
[0036] Other features and advantages will be described in the following description. The objectives and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description and drawings.
[0037] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The above and other features and advantages of the present invention will become more apparent by describing in detail example embodiments thereof with reference to the accompanying drawings;
[0039] Figure 1 It is a framework diagram of the intelligent equipment visual data management and AI model development platform;
[0040] Figure 2 It is a flowchart of the management method of historical versions;
[0041] Figure 3 is a flow chart of a method for determining risk training data in training data;
[0042] Figure 4 It is a flowchart of a method for determining an optimization target version among historical versions;
[0043] Figure 5 This is a flowchart of a method for determining an optimization requirement type within a version range. DETAILED DESCRIPTION
[0044] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this specification without creative work should fall within the scope of protection of this specification.
[0045] Example 1
[0046] like Figure 1 As shown, this application provides a smart equipment visual data management and AI model development platform, specifically including:
[0047] Data annotation module, model training module, version management module;
[0048] The data annotation module is responsible for performing data annotation processing to obtain annotated data;
[0049] The model training module is responsible for using the labeled data to perform training processing of the visual model as needed;
[0050] The version management module is responsible for managing different historical versions of the visual model using the training data and the updated training status of the training data in the visual model.
[0051] Furthermore, the data annotation processing includes manual annotation and automatic annotation.
[0052] Specifically, it also includes a data enhancement module, which is responsible for amplifying the labeled data.
[0053] It should be noted that the amplification processing includes rotation, cropping, flipping, adversarial generation network, real visual background synthesis of moving targets, and amplification processing methods of real image + simulated visual interference synthesis.
[0054] It is understandable that when the visual model is being trained, changes in the model parameters of the visual model support visual display.
[0055] Specifically, the management of different historical versions of the visual model includes:
[0056] When the visual model of the training data of the historical version has not been incrementally trained within the most recent preset time period, the need for forgetting management is determined by comparing the incrementally trained visual model with the verification result of the optimized target version in the target database. When the deviation between the accuracy of the verification result of the optimized target version and the visual model is greater than the preset value of the deviation, forgetting processing of the optimized target version is performed.
[0057] It should also be noted that when the training data of the historical version is incrementally trained on the visual model within the most recent preset time period, when a visual model of the optimized target version with similar model parameters appears after incremental training, the verification results between the visual model of the optimized target version with similar model parameters and the optimized target version in the target database are used to determine whether forgetting management is needed, wherein whether the model parameters are similar is determined based on whether the deviation of the model parameters is within a preset range, wherein when the accuracy of the verification result of the optimized target version and the deviation of the visual model is greater than a second deviation preset value, the optimized target version forgetting processing is performed, wherein the deviation preset value is greater than the second deviation preset value.
[0058] Example 2
[0059] According to the proportion of the data volume of the automatically labeled training data of the type under different incremental training times, the proportion of the data volume under different incremental training times is determined. When the average of the proportion of the data volume under different incremental training times is greater than 0.1, the training data is determined to be risky training data.
[0060] The optimization target version is a historical version whose data volume ratio of training risk data during incremental training is less than 0.2, and there is a historical version in other later historical versions whose deviation from the data volume ratio of its training risk data meets the requirements. Specifically, the later versions whose data volume deviations of different risk data of interest from the historical versions are all within the preset deviation range are regarded as similar versions. When the number of similar versions is greater than a threshold, the historical version is determined to be the optimization target version.
[0061] If the number of historical versions within the version interval is greater than the preset version number threshold and the number of optimization target versions is only one, then the optimization requirement type of the optimization target version within the version interval is determined to be a class requirement type; if the number of historical versions is not greater than the preset version number threshold, then the optimization requirement type of the optimization target version within the version interval is determined to be a target requirement type.
[0062] If the number of historical versions is greater than the preset version number threshold and the number of optimization target versions is more than one, it is determined that it belongs to the second type of demand.
[0063] The historical versions corresponding to different types of matching risk data during incremental training within the most recent preset time period are used as updated versions, and the version intervals belonging to the second type of demand are used as second type of demand intervals. When the number of second type of demand intervals is greater than the preset number of demand intervals or there is matching risk data with the number of updated versions greater than the preset update version number threshold, it is determined that the forgetting management method of the optimized target version does not require forgetting management.
[0064] Second, as Figure 2 As shown, this application provides a historical version management method, which is applied to the above-mentioned intelligent equipment visual data management and AI model development platform, specifically including:
[0065] S1 determines the composition of automatically labeled data under different types of training data based on the training data of the visual model, and determines risk training data in the training data based on the composition;
[0066] Furthermore, the composition of the automatic annotation of the training data includes the data volume and data volume ratio of the automatically annotated training data during different incremental trainings.
[0067] Specifically, the type of the training data is determined according to the recognition target corresponding to the training data.
[0068] Specifically, such as Figure 3 As shown, the method for determining the risk training data in the training data is:
[0069] Determining the amount of automatically labeled training data of the type under different times of incremental training based on the composition of the training data of the type during the incremental training of the visual model in different historical versions;
[0070] Determining the proportion of data volume under different incremental training times according to the proportion of data volume of the automatically labeled training data of the type under different incremental training times;
[0071] Based on the proportion of data volume under different incremental training times, it is determined whether the training data is risky training data.
[0072] It can be understood that when there is a preset number of data volumes that do not meet the required incremental training times, the training data is determined to be risky training data. Specifically, when the data volume ratio is greater than a preset threshold, it is determined that it meets the requirements.
[0073] In addition, it should be noted that when there is no risk training data, the verification results between historical versions with similar model parameters in the target database can be used according to the preset period to determine whether forgetting management is needed. Whether the model parameters are similar is determined based on whether the deviation of the model parameters is within the preset range. The historical versions with lower verification result accuracy can be forgotten.
[0074] In another possible embodiment, the method for determining risk training data in the training data is:
[0075] Determining the amount of automatically labeled training data of the type under different times of incremental training based on the composition of the training data of the type during the incremental training of the visual model in different historical versions;
[0076] Whether the training data is risky training data is determined according to the data amount of the automatically labeled training data of the type under different incremental training times.
[0077] It is understandable that when there is a preset amount of data that does not meet the required number of incremental training times, the training data is determined to be risky training data.
[0078] S2 determines the composition of training risk data during incremental training of different historical versions, and determines the optimization target version in the historical version based on the similarity between the training risk data and other historical versions in the later period;
[0079] Furthermore, the composition of the training risk data includes the types of training risk data in the training data during the incremental training of the historical version and the data amounts of different types of training risk data.
[0080] It is understandable that if Figure 4 As shown, the method for determining the optimization target version in the historical version is:
[0081] Determining the amount of different types of training risk data in the historical version based on the composition of the training risk data during incremental training of the historical version, and determining the risk data of interest in the training risk data based on the amount of data;
[0082] Taking a historical version later than the historical version as a later version, and determining the similarity between the later version and the historical version in terms of the amount of different risk data according to the overlap between the risk data of concern of the later version and the historical version;
[0083] According to the similarity between the later version and the historical version in the amount of different risk data, it is determined whether the historical version is the optimization target version.
[0084] Specifically, the risk data of concern is training risk data whose data volume does not meet the requirement, and specifically, the training risk data whose data volume is greater than a preset data volume threshold is used as the risk data of concern.
[0085] It should also be noted that the similarities between the later version and the historical version in terms of the data amounts of different risk data include deviations in terms of the data amounts of different risk data.
[0086] Further, based on the similarity between the later version and the historical version in the amount of different risk data, determining whether the historical version is an optimization target version specifically includes:
[0087] The later versions whose data amount deviations in the risk data different from the historical version are all within the preset deviation range are regarded as similar versions. When the number of similar versions is large, that is, greater than the threshold, the historical version is determined to be the optimization target version.
[0088] Optionally, the method for determining the optimization target version in the historical version is:
[0089] Determining the amount of different types of training risk data in the historical version based on the composition of the training risk data during incremental training of the historical version, and determining the risk data of interest in the training risk data based on the amount of data;
[0090] It should also be noted that in one possible embodiment, if the total amount of different types of training risk data or the proportion of training risk data in the training data in the above steps does not meet the requirements, that is, when the amount of training risk data is large, when the problem is located in the later stage, the probability of there being a problem is relatively high, so it can be determined that it does not belong to the optimization target version.
[0091] In addition, it should be further explained that even if the total amount of different types of training risk data in the above steps meets the requirements, it is still necessary to determine whether the number of types of risk data of concern meets the requirements. Specifically, when the number of types of risk data of concern is greater than a certain threshold, it is determined that it does not belong to the optimization target version.
[0092] Even if the number of risk data types meets the requirements, if the amount of risk data does not meet the requirements, it can be directly determined that it does not belong to the optimization target version, and whether the requirements are met can be determined by means of a threshold.
[0093] A historical version later than the historical version is taken as a later version, and based on the overlap of the risk data of concern between the later version and the historical version, a similarity in the amount of data of the later version and the historical version under different risk data of concern is determined, and based on the similarity, a deviation in the amount of data of the later version under different risk data of concern is determined;
[0094] It should be further explained that when the deviation of the data amount in the later version that does not exist in the historical version under different risk data of concern meets the requirements of the later version, then during the later training and positioning processing, due to the dissimilarity of the training data, the risk of anomalies caused by the intersection of multiple risk data of concern is greater, and the probability of it being needed as a positioning target is greater. In this case, it can be determined that the historical version does not belong to the optimization target version.
[0095] It can also be understood that when the deviation in the amount of data existing in the historical version under different risk data of concern meets the requirements of the later version, it will be regarded as a similar version. When the number of similar versions is large, that is, greater than the threshold, the historical version is determined to be the optimization target version.
[0096] If the number of similar versions is small, it is necessary to proceed to the next step to continue the judgment.
[0097] Whether the historical version is the optimization target version is determined according to the deviation of the data volume of different later versions under different focus risk data and the data volume of the focus risk data.
[0098] It can be understood that in a possible embodiment, the total amount of risk data is used as a basis to determine the threshold number of similar versions under the total amount of data. If and only if the number of similar versions is greater than the threshold number of similar versions under the total amount of data, the historical version is used as the optimization target version.
[0099] S3 divides historical versions into different version intervals based on training risk data. Based on the composition data of historical versions in different version intervals and the changes in model parameters between the optimization target version and later historical versions, it determines the type of optimization requirements in different version intervals.
[0100] Specifically, such as Figure 5 As shown, the method for determining the optimization requirement type within the version range is:
[0101] Determine the number of historical versions in the version interval and the number of optimization target versions based on the constituent data of the historical versions in the version interval;
[0102] Determining the similarity between the model parameters of the optimization target version and the later historical version based on the changes in the model parameters of the optimization target version and the later historical version, and determining the parameter similarity version based on the similarity;
[0103] The optimization requirement type of the optimization target version within the version range is determined according to the number of historical versions, the number of optimization target versions, and the number of parameter-similar versions of different optimization versions.
[0104] It can be understood that the parameter-similar version is a later historical version in which the deviations of different model parameters from the optimization target version meet the requirements. Whether the requirements are met is specifically determined by whether they are within a preset range.
[0105] In a possible embodiment, if the number of historical versions is greater than a preset version number threshold and there is only one optimization target version, it is determined that the optimization requirement type of the optimization target version within the version range is a Class I requirement type.
[0106] It should also be noted that if the number of historical versions is not greater than the preset version number threshold, the optimization requirement type of the optimization target version within the version range is determined to be the target requirement type.
[0107] Furthermore, if the number of historical versions is greater than the preset version number threshold and the number of optimization target versions is more than one, the optimization requirement type of the optimization target version is determined based on the number of parameter-similar versions of the optimization target version within the version range. Specifically, if the number of parameter-similar versions is above the preset reference version number, it is determined to belong to the first type of requirement; if it is not above the preset reference version number, it is determined to be the second type of requirement. In a possible embodiment, the value range of the preset reference version number is 3 to 5.
[0108] It should be noted that, when the optimization target version includes a version interval belonging to the target requirement type, it is impossible to forget it because there is a version with a high degree of correlation for abnormality troubleshooting.
[0109] If there is no version interval of the target requirement type, and different version intervals belong to the same requirement type, it can be determined that they will be forgotten.
[0110] In another embodiment, the method for determining the optimization requirement type within the version range is:
[0111] Determine the number of historical versions in the version interval and the number of optimization target versions based on the constituent data of the historical versions in the version interval;
[0112] In the above steps, if the number of historical versions is not greater than the preset version number threshold, the optimization requirement type of the optimization target version within the version interval is determined to be the target requirement type.
[0113] Furthermore, if the number of historical versions is greater than the preset version number threshold, the number of historical versions is large at this time, so it is necessary to determine whether there is a change deviation parameter, and it is necessary to proceed to the next step to determine the change deviation parameter.
[0114] Determining a change trend of the model parameters of the optimization target version and the later historical versions according to changes in the model parameters of the optimization target version and the later historical versions, and determining a change deviation parameter based on the change trend;
[0115] It can be understood that if the number of change deviation parameters is large, that is, greater than a certain threshold, it means that abnormal traceability processing of change deviation parameters in different dimensions is required. Therefore, the optimization requirement type of the optimization target version within the version range can be directly determined as the target requirement type.
[0116] In addition, if in the change deviation parameters, if the deviation amount of the change parameters between the latest version and the optimization target version is not within the preset parameter deviation range, that is, if there are change deviation parameters with larger changes, then there are change deviation parameters with higher requirements during abnormality troubleshooting, so the optimization requirement type of the optimization target version within the version range can be directly determined as the target requirement type.
[0117] If the deviations of the change parameters between the latest version and the optimization target version are both within the preset parameter deviation range, specifically if the number of change deviation parameters does not meet the requirements or the number of changes in the change deviation parameters does not meet the requirements within the preset change range, it can be determined whether the requirements are met by means of a threshold, and the optimization requirement type of the optimization target version within the version range is determined to be the target requirement type.
[0118] Only when there is no change deviation parameter and the number of reference versions is above the preset reference version number, it is determined to belong to the first type of demand. In other cases, it is necessary to comprehensively consider multiple factors to optimize the determination of the demand type.
[0119] Specifically, the change deviation parameter is a model parameter with a consistent change trend between different adjacent historical versions in the later period. Specifically, if the model parameter data of the later historical version and the previous historical version both become larger or smaller, it means that the model parameter is a change deviation parameter.
[0120] The optimization requirement type of the optimization target version within the version range is determined according to the change of the change deviation parameter, the number of historical versions, and the number of optimization target versions.
[0121] It can be understood that the change of the change deviation parameter is determined according to the deviation of the change parameter between the latest version and the optimization target version.
[0122] It should be noted that if there is only one optimization target version, the optimization requirement type of the optimization target version within the version range is determined to be a Class I requirement type. When there are multiple optimization target versions, it is necessary to determine the changes in the change deviation parameters of different optimization target versions and the number of parameter-similar versions. Specifically, the optimization target whose deviation between the latest version and the optimization target version in different change deviation parameters is less than the preset change threshold and the number of parameter-similar versions is greater than the preset reference version number is regarded as a Class I requirement type. In other cases, it is a Class II requirement type.
[0123] S4 When the optimization requirement types of the optimization target version in different version intervals do not belong to the target requirement types, the forgetting management method of the optimization target version is determined based on the updated training status of the training risk data corresponding to the optimization target version in the visual model and the optimization requirement types in different version intervals.
[0124] Specifically, the method for determining the forgetting management method of the optimization target version is:
[0125] Using the training risk data corresponding to the optimization target version as matching risk data, and determining historical versions corresponding to different types of matching risk data during incremental training within a recent preset period based on the updated training status of the matching risk data in the visual model, and using the historical versions as updated versions;
[0126] Based on the optimization requirement types within different version intervals, determine the version interval belonging to the second type of requirement and use it as the second type of requirement interval;
[0127] According to the updated versions of the matching risk data of different types and the constituent data of the second type of demand intervals, the forgetting management method of the optimized target version is determined.
[0128] Specifically, when the number of the second-category demand intervals is greater than the preset number of demand intervals or there is matching risk data with the number of updated versions greater than the preset update version number threshold, it is determined that the forgetting management method of the optimized target version does not require forgetting processing.
[0129] It should also be noted that when the number of second-category demand intervals is not greater than the preset number of demand intervals and there is matching risk data in which the number of updated versions is not greater than the preset update version number threshold, based on the sum of the number of updated versions under different types of matching risk data, when the sum of the number of updated versions under different types of matching risk data is greater than the preset update version number threshold, it is determined that the forgetting management method of the optimized target version does not require forgetting processing.
[0130] In addition, it needs to be further explained that if the sum of the number of updated versions under different types of matching risk data is not greater than the preset update version number threshold, the update demand is determined by the sum of the number of updated versions under different types of matching risk data and the number of the two types of demand intervals. When the update demand is less than the preset demand threshold, as long as the incremental training of the visual model is performed, that is, the verification result of the incrementally trained visual model and the optimized target version in the target database is used to determine whether forgetting management is needed. When the deviation between the accuracy of the verification result of the optimized target version and the visual model is greater than the preset deviation value, forgetting processing of the optimized target version is performed.
[0131] If the update requirement is less than the preset requirement threshold, when a visual model of the optimized target version with similar model parameters appears after incremental training, the verification result between the visual model of the optimized target version with similar model parameters and the optimized target version in the target database is used to determine whether forgetting management is needed, wherein whether the model parameters are similar is determined based on whether the deviation of the model parameters is within a preset range, wherein when the accuracy of the verification result of the optimized target version and the deviation of the visual model is greater than a second deviation preset value, the optimized target version forgetting processing is performed, wherein the deviation preset value is greater than the second deviation preset value.
[0132] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For relevant details, refer to the descriptions of the method embodiments.
[0133] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0134] The foregoing description is merely one or more embodiments of this specification and is not intended to limit this specification. It will be apparent to those skilled in the art that various modifications and variations may be made to one or more embodiments of this specification. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of one or more embodiments of this specification are intended to be within the scope of the claims of this specification.
Claims
1. A system for intelligent equipment visual data management and AI model development, characterized by: Specifically include: Data annotation module, model training module, version management module; The data annotation module is responsible for performing data annotation processing to obtain annotated data; The model training module is responsible for using the labeled data to perform training processing of the visual model as needed; The version management module is responsible for managing different historical versions of the visual model using the training data and the updated training status of the training data in the visual model; Based on the training data of the visual model, determining the composition of automatically labeled data under different types of training data, and determining risk training data in the training data based on the composition; Determine the composition of training risk data during incremental training for different historical versions, and determine the target version for optimization in the historical version based on the similarity between the training risk data and other historical versions in the later period. Based on the training risk data, historical versions are divided into different version intervals. Based on the composition data of the historical versions in different version intervals and the changes in model parameters between the optimization target version and later historical versions, the optimization requirements in different version intervals are determined. When the optimization requirement types of the optimization target version in different version intervals do not belong to the target requirement type, the forgetting management method of the optimization target version is determined based on the updated training status of the training risk data corresponding to the optimization target version in the visual model and the optimization requirement types in different version intervals; The method for determining the risk training data in the training data is: Determining the amount of automatically labeled training data of the type under different times of incremental training based on the composition of the training data of the type during the incremental training of the visual model in different historical versions; Determining the proportion of data volume under different incremental training times according to the proportion of data volume of the automatically labeled training data of the type under different incremental training times; Determining whether the training data is risky training data based on the proportion of data under different incremental training times; When there is a data volume greater than a preset number and the proportion thereof does not meet the required incremental training times, the training data is determined to be risky training data.
2. The intelligent equipment visual data management and AI model development system according to claim 1, characterized in that: The data annotation processing includes manual annotation and automatic annotation.
3. The intelligent equipment visual data management and AI model development system according to claim 1, characterized in that: It also includes a data enhancement module, which is responsible for amplifying the labeled data.
4. The intelligent equipment visual data management and AI model development system according to claim 3, characterized in that: The amplification processing includes rotation, cropping, flipping, adversarial generation network, real visual background synthesis of moving targets, and amplification processing methods of real image + simulated visual interference synthesis.
5. The intelligent equipment visual data management and AI model development system according to claim 1, characterized in that: The method for determining the forgetting management method of the optimization target version is: Using the training risk data corresponding to the optimization target version as matching risk data, and determining historical versions corresponding to different types of matching risk data during incremental training within a recent preset period based on the updated training status of the matching risk data in the visual model, and using the historical versions as updated versions; Based on the optimization requirement types within different version intervals, determine the version interval belonging to the second type of requirement and use it as the second type of requirement interval; According to the updated versions of the matching risk data of different types and the constituent data of the second type of demand intervals, the forgetting management method of the optimized target version is determined.
6. The intelligent equipment visual data management and AI model development system according to claim 5, characterized in that: When the number of the second-category demand intervals is greater than the preset number of demand intervals or there is matching risk data in which the number of updated versions is greater than the preset update version number threshold, it is determined that the forgetting management method of the optimized target version does not require exception management.
Citation Information
Patent Citations
AI calculation training system
CN119091200A