A method and system for big data collection and annotation

By constructing a quality annotation channel for data acquisition and combining iterative optimization mechanisms for differentiated attention learning and annotation loss, efficient and accurate automatic annotation of monitoring image data was achieved, improving the efficiency and accuracy of image annotation.

CN120707987BActive Publication Date: 2025-12-26BEIJING HONG KONG TECHNOLOGY RESEARCH INSTITUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510826175.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-12-26
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and insufficient accuracy in annotating surveillance image data, especially in multi-node collaborative environments where image acquisition quality varies and annotation consistency is difficult to guarantee.

Method used

By constructing an acquisition quality annotation channel based on image quality resolution factors, introducing a differentiated attention learning mechanism and an annotation loss iterative optimization mechanism, and combining image annotation multivariate factors for federated learning of annotation features and cloud distillation, automatic annotation is achieved.

Benefits of technology

It improves the efficiency and accuracy of image annotation, and solves the problems of low efficiency and insufficient accuracy of existing technologies in the annotation of surveillance image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707987B_ABST
    Figure CN120707987B_ABST
Patent Text Reader

Abstract

The application discloses a kind of method and system for big data acquisition and annotation, it is related to image annotation technical field.The method comprises: according to the Q edge monitoring nodes of monitoring system, obtains Q real-time monitoring set;According to image quality analysis factor, image quality annotation learning is carried out to Q edge monitoring nodes, and acquisition quality annotation channel is constructed;Q real-time monitoring set is adaptively reinforced, and Q reinforcement monitoring set is obtained;Introduce differentiating attention learning mechanism, and carry out annotation feature federation learning to Q edge monitoring nodes, and obtain edge annotation multi-channel;Introduce annotation loss iterative optimization mechanism to carry out cloud distillation to edge annotation multi-channel, and obtain cloud annotation multi-channel;According to cloud annotation multi-channel, Q reinforcement monitoring set is automatically annotated.Solve the technical problems of low efficiency and insufficient accuracy of monitoring image data annotation in the prior art, and achieve the technical effect of improving image annotation efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image annotation, and in particular to a method and system for big data collection and annotation. BACKGROUND

[0002] With the development of intelligent monitoring systems and big data analysis technologies, real-time collection and high-quality annotation of massive video image data have become a key link for intelligent analysis and training model construction. In typical scenarios such as traffic monitoring, safety inspection, and urban management, edge-side devices are increasingly widely deployed, generating a large amount of heterogeneous and dynamically changing image data. In order to support the training and optimization of deep learning models, the collected images need to be accurately and comprehensively annotated. However, in the prior art, image annotation mostly relies on manual or semi-automatic processes, which have problems such as low efficiency, high cost, and insufficient accuracy. In particular, in a multi-node collaborative environment, the quality of image collection is not uniform, and annotation consistency is difficult to guarantee. SUMMARY

[0003] The present application provides a method and system for big data collection and annotation, which solves the technical problems of low efficiency and insufficient accuracy of monitoring image data annotation in the prior art.

[0004] In a first aspect, the present application provides a method for big data collection and annotation, comprising:

[0005] According to Q edge monitoring nodes of a monitoring system, Q real-time monitoring sets are obtained, Q being a positive integer greater than 1; image quality annotation learning is performed on the Q edge monitoring nodes according to an image quality analysis factor, and a collection quality annotation channel is constructed; the Q real-time monitoring sets are adaptively reinforced according to the collection quality annotation channel, and Q reinforced monitoring sets are obtained; a differentiated attention learning mechanism is introduced, and annotation feature federated learning is performed on the Q edge monitoring nodes in combination with an image annotation multi-element factor, and edge annotation multi-channels are obtained; a label loss iterative optimization mechanism is introduced to perform cloud distillation on the edge annotation multi-channels, and cloud annotation multi-channels are obtained; and automatic annotation is performed on the Q reinforced monitoring sets according to the cloud annotation multi-channels.

[0006] In a second aspect, the present application provides a system for big data collection and annotation, comprising:

[0007] The data acquisition module obtains Q real-time monitoring sets according to Q edge monitoring nodes of a monitoring system, Q being a positive integer greater than 1; the label learning module performs image quality label learning on the Q edge monitoring nodes according to an image quality analysis factor, and constructs a collection quality label channel; the data strengthening module performs self-adaptive strengthening on the Q real-time monitoring sets according to the collection quality label channel, and obtains Q strengthened monitoring sets; the federated learning module introduces a differential attention learning mechanism, and performs label feature federated learning on the Q edge monitoring nodes in combination with an image label multi-element factor, and obtains an edge label multi-channel; the distillation optimization module introduces a label loss iterative optimization mechanism to perform cloud distillation on the edge label multi-channel, and obtains a cloud label multi-channel; and the automatic labeling module performs automatic labeling on the Q strengthened monitoring sets according to the cloud label multi-channel.

[0008] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0009] First, Q real-time monitoring sets are obtained according to Q edge monitoring nodes of a monitoring system, Q being a positive integer greater than 1. Then, image quality label learning is performed on the Q edge monitoring nodes according to an image quality analysis factor, and a collection quality label channel is constructed. Then, self-adaptive strengthening is performed on the Q real-time monitoring sets according to the collection quality label channel, and Q strengthened monitoring sets are obtained. Further, a differential attention learning mechanism is introduced, and label feature federated learning is performed on the Q edge monitoring nodes in combination with an image label multi-element factor, and an edge label multi-channel is obtained; a label loss iterative optimization mechanism is introduced to perform cloud distillation on the edge label multi-channel, and a cloud label multi-channel is obtained. Finally, automatic labeling is performed on the Q strengthened monitoring sets according to the cloud label multi-channel. The technical problems of low monitoring image data labeling efficiency and insufficient accuracy in the prior art are solved, and the technical effects of improving image labeling efficiency and accuracy are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0011] Figure 1 A method flowchart for big data acquisition and labeling is provided for the embodiments of the present application;

[0012] Figure 2 A system structure diagram for big data acquisition and labeling is provided for the embodiments of the present application.

[0013] Explanation of reference signs: data collection module 11, annotation learning module 12, data enhancement module 13, federated learning module 14, distillation optimization module 15, automatic annotation module 16. DETAILED DESCRIPTION

[0014] The present application provides a method and system for big data collection and annotation, which solves the technical problems of low efficiency and insufficient accuracy of monitoring image data annotation in the prior art.

[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0016] It should be noted that the terms "comprise" and "have" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0017] Embodiment one, as shown in the present application provides a method for big data collection and annotation, wherein the method comprises: Figure 1

[0018] According to the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, and Q is a positive integer greater than 1.

[0019] In the embodiments of the present application, data collection is performed based on the Q edge monitoring nodes deployed in the monitoring system, wherein the edge monitoring nodes are a plurality of front-end camera devices or image perception terminals with independent video collection capabilities, distributed at different positions in the monitored area. Each edge monitoring node can independently complete video image collection, preliminary data processing and communication tasks.

[0020] Specifically, the system sends a data collection instruction to all Q edge monitoring nodes through the access management module according to a unified time synchronization protocol; after receiving the instruction, each edge monitoring node starts the video stream collection function, captures monitoring image frames in real time according to the preset frame rate (such as 10fps, 25fps, etc.) and image resolution (such as 720p, 1080p); each node continuously collects image frame data within a certain time period to form Q independent real-time monitoring sets.

[0021] ​According to the image quality analysis factor, the Q edge monitoring nodes are subjected to image quality annotation learning, and an acquisition quality annotation channel is constructed. Further, the image quality analysis factor includes global fidelity, local sharpness, and connection smoothness.

[0022] Further, according to the image quality analysis factor, the Q edge monitoring nodes are subjected to image quality annotation learning, and an acquisition quality annotation channel is constructed, including:

[0023] Q image quality record sets of the Q edge monitoring nodes are obtained; according to the image quality analysis factor, the Q image quality record sets are subjected to image quality annotation learning, and Q image quality detection models are obtained; according to the image quality analysis factor, the Q image quality detection models are subjected to same-class evaluation layer aggregation, and a global fidelity evaluation node, a local sharpness evaluation node, and a connection smoothness evaluation node are obtained; the global fidelity evaluation node, the local sharpness evaluation node, and the connection smoothness evaluation node are connected in parallel, and the acquisition quality annotation channel is generated.

[0024] Global fidelity is used to measure the restoration degree between the whole image and the real scene, and is usually modeled by indicators such as color restoration degree, global contrast, and average brightness balance; local sharpness reflects the clarity degree of local details of the image, and can be extracted and evaluated by parameters such as edge strength, local gradient change, and texture feature density; connection smoothness is used to evaluate the spatial continuity and contour integrity of structural elements in the image, and is usually calculated in combination with image connected regions, edge continuity, and spatial distribution consistency.

[0025] Q edge monitoring nodes respectively obtain corresponding image quality record sets, forming Q independent image quality record sets, each of which contains image data and its quality evaluation indicators collected by the corresponding node; taking global fidelity, local sharpness and connection smoothness as image quality analysis factors, feature extraction and label learning are performed on each image quality record set respectively to construct a deep model for image quality recognition, and a multi-branch neural network structure is adopted in the training process, each branch corresponding to a quality factor, and independent supervised loss function is used for targeted feature optimization, thereby obtaining Q image quality detection models, wherein each image quality detection model internally includes a global fidelity evaluation layer, a local sharpness evaluation layer and a connection smoothness evaluation layer for hierarchical and multi-angle evaluation of different types of image quality factors. The global fidelity evaluation layers in all image quality detection models are summarized, and layer weight alignment, distillation fusion or average pooling method is used to generate a unified global fidelity evaluation node; the local sharpness evaluation layer and the connection smoothness evaluation layer are processed in the same way to generate a unified local sharpness evaluation node and a unified connection smoothness evaluation node. Finally, the global fidelity evaluation node, the local sharpness evaluation node and the connection smoothness evaluation node are connected in parallel to construct a multi-channel acquisition quality label channel for unified output of three-factor quality scores or quality label results of image frames.

[0026] Further, according to the image quality analysis factors, the Q image quality record sets are subjected to image quality label learning, comprising:

[0027] According to the Q image quality record sets, a first image quality record set is extracted, and the first image quality record set is subjected to clustering processing according to the image quality analysis factors to generate a first global fidelity record, a first local sharpness record and a first connection smoothness record; the image quality evaluator is subjected to supervised training according to the first global fidelity record, and a fidelity evaluation loss coefficient is obtained every predetermined number of training times; if the fidelity evaluation loss coefficient is less than a fidelity evaluation loss threshold, a first global fidelity evaluation layer is generated; the first local sharpness record and the first connection smoothness record are subjected to loss supervised training according to the image quality evaluator to generate a first local sharpness evaluation layer and a first connection smoothness evaluation layer; the first global fidelity evaluation layer, the first local sharpness evaluation layer and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.

[0028] From the Q image quality record sets, data corresponding to an edge monitoring node is selected, image quality records of the node are extracted, and a first image quality record set is formed; the first image quality record set is clustered according to an image quality analysis factor, and the first image quality record set is divided into three subsets by a multi-dimensional feature clustering algorithm (for example, a combination method based on principal component analysis and K-means clustering), to correspond to a first global fidelity record, a first local sharpness record, and a first connection smoothness record, respectively, to represent representative samples in the corresponding quality dimensions. The first global fidelity record is used as a training sample to construct an image quality evaluator and perform supervised training. The image quality evaluator can be a convolutional neural network structure. A predetermined number of iterations or rounds is set during the training process. At the end of each round of training, a fidelity evaluation loss coefficient is calculated, and the coefficient is compared with a set fidelity evaluation loss threshold. When the loss coefficient decreases and stabilizes to be less than the threshold, it is considered that the learning of the evaluator in the global fidelity aspect has converged. At this time, a first global fidelity evaluation layer can be constructed to represent the stable modeling capability of the image quality evaluator for the global fidelity feature. The above-mentioned image quality evaluator is continuously used to introduce the first local sharpness record and the first connection smoothness record on the basis of the learned fidelity feature, and supervised training is sequentially performed to obtain a first local sharpness evaluation layer and a first connection smoothness evaluation layer by optimizing the corresponding sharpness loss function and the smoothness loss function. After the training of the above-mentioned three evaluation layers is completed, the three evaluation layers are connected in a series or module nesting structure to integrate and generate a first image quality detection model.

[0029] According to the acquisition quality marking channel, the Q real-time monitoring sets are adaptively reinforced to obtain Q reinforced monitoring sets.

[0030] The acquisition quality marking channel is used to perform quality discrimination on each image frame in the Q real-time monitoring sets, and marking results of the image frames in three dimensions of global fidelity, local sharpness, and connection smoothness are respectively output to form corresponding image quality marking sequences. According to a preset image quality expectation sequence (which reflects the expected image quality standard), the quality scores of each image frame are analyzed for differences, and difference values between the images and the expected quality are calculated. Based on the difference values, an adaptive reinforcement algorithm (such as a dynamic weight adjustment or a reinforcement learning strategy) is used to perform targeted processing on the real-time monitoring sets, including but not limited to adjusting acquisition parameters (such as exposure time and focusing settings), optimizing image processing algorithms (such as denoising and sharpening), or selecting low-quality frames. After the adaptive reinforcement processing, the image quality in the original Q real-time monitoring sets is significantly improved, and finally Q reinforced monitoring sets are generated.

[0031] Further, according to the acquisition quality marking channel, the Q real-time monitoring sets are adaptively reinforced to obtain Q reinforced monitoring sets, including:

[0032] According to the acquisition quality marking channel, image quality of each image frame in the Q real-time monitoring sets is marked to obtain a sequence of image qualities of each frame; according to the image quality analysis factor, an expected sequence of image qualities is constructed; according to the expected sequence of image qualities, difference analysis is performed on the sequence of image qualities of each frame to obtain a plurality of sequences of image quality differences; and according to the plurality of sequences of image quality differences, the Q real-time monitoring sets are enhanced to generate the Q enhanced monitoring sets.

[0033] First, the constructed acquisition quality marking channel is used to perform image quality evaluation and marking on the image frames collected by each edge monitoring node in the Q real-time monitoring sets one by one, and the quality marking results of each frame of image in three dimensions of global fidelity, local sharpness and connection smoothness are obtained respectively, and are organized in time sequence to form a corresponding sequence of image qualities. Subsequently, an expected sequence of image qualities is constructed based on the image quality analysis factor, and the expected sequence of image qualities reflects the image quality benchmark expected to be achieved in the actual application scenario, including but not limited to average sharpness standard, minimum fidelity threshold, edge continuity target and the like. Next, the sequence of image qualities and the expected sequence of image qualities are compared frame by frame for difference, and difference factor extraction and analysis operations are performed to obtain deviation indexes of each image frame in three quality dimensions, and finally a plurality of sequences of image quality differences are generated to quantify the deviation degree between the image and the ideal quality. On this basis, the system performs adaptive enhancement processing according to the sequence of image quality differences, including but not limited to weighted reconstruction of image frames, resampling or replacement of low-quality frames, enhancement compensation processing of high-quality frames, dynamic adjustment of image acquisition parameters and the like, so that the overall monitoring image sequence is more consistent in the quality dimension and meets the requirements of subsequent automatic marking. Finally, the Q enhanced monitoring sets are output.

[0034] A differential attention learning mechanism is introduced, and a labeled feature federation learning is performed on the Q edge monitoring nodes in combination with image labeling multi-element factors to obtain an edge labeled multi-channel.

[0035] The differential attention learning mechanism can automatically adjust the attention weights of different regions or features according to the complexity and dynamics of image content and the like.

[0036] For each edge monitoring node, collect its historical image annotation data, covering attribute features, pose features, and additional features, and other multi-factor features, which reflect the multi-dimensional information of the target object and environmental changes. Based on the differentiated attention learning mechanism, the feature traversal and correlation evaluation of the image annotation history set of each node are performed, and different weights are dynamically allocated to highlight key features while suppressing noise and redundant information, thereby optimizing the expression effect of the annotation features. On this basis, the attention-optimized image set of each edge node is supervised and trained, and a local image annotator including an attribute feature annotation model, a pose feature annotation model, and an additional feature annotation model is constructed. Then, through the federated learning framework, the same type of annotation models of each edge node are aggregated, and a comprehensive edge annotation multi-channel is generated using a decentralized data collaboration update method, which includes an attribute feature annotation channel, a pose feature annotation channel, and an additional feature annotation channel.

[0037] Further, a differentiated attention learning mechanism is introduced, and the Q edge monitoring nodes are annotated with multi-factor image annotation federated learning, obtaining an edge annotation multi-channel, including:

[0038] The image annotation multi-factor includes attribute features, pose features, and additional features; based on the image annotation multi-factor, the Q image annotation history sets of the Q edge monitoring nodes are attention-optimized according to the differentiated attention learning mechanism, obtaining Q attention-optimized image sets; the Q attention-optimized image sets are supervised and trained respectively, generating Q image annotators, wherein each image annotator includes an attribute feature annotation model, a pose feature annotation model, and an additional feature annotation model; the same type of annotation models are federated and aggregated according to the Q image annotators, generating the edge annotation multi-channel, which includes an attribute feature annotation channel, a pose feature annotation channel, and an additional feature annotation channel.

[0039] The image annotation multi-factor includes attribute features (such as color, shape), pose features (such as object position, direction), and additional features (such as background information, lighting conditions), which together reflect the diversified information of the target in the image.

[0040] Based on the image annotation multi-element factor, a differential attention learning mechanism is used to perform attention optimization processing on the Q image annotation history sets of the Q edge monitoring nodes. The differential attention learning mechanism dynamically adjusts the attention weights of different feature regions, so that the model can focus more on key features, thereby generating Q attention optimization image sets. Each attention optimization image set highlights the importance of the corresponding edge monitoring node in a specific annotation feature. The Q attention optimization image sets are supervised trained respectively, and Q corresponding image annotators are constructed, each of which includes an attribute feature annotation model, a pose feature annotation model and an additional feature annotation model. Finally, for the same annotation model, a federal aggregation method (such as federal average algorithm) is used for fusion to generate an edge annotation multi-channel, which includes an attribute feature annotation channel, a pose feature annotation channel and an additional feature annotation channel, realizing the collaborative learning and fusion of multi-dimensional annotation features. Specifically, the attribute feature annotation model in the Q image annotators is federally aggregated to generate an attribute feature annotation channel; similarly, a pose feature annotation channel and an additional feature annotation channel are generated.

[0041] Further, the differential attention learning mechanism includes:

[0042] According to the Q image annotation history sets, the qth image annotation history set is extracted, q is a positive integer, 1≤q≤Q; according to the image annotation multi-element factor, the qth image annotation history set is evaluated for feature traversal correlation, to generate a ternary feature correlation evaluation set; according to the image annotation multi-element factor, a differential weight distribution is performed to generate an annotation weight vector; based on the annotation weight vector, the qth image annotation history set is optimized for attention based on the ternary feature correlation evaluation set.

[0043] In the implementation of the differential attention learning mechanism, first, an image annotation history set is extracted from the Q image annotation history sets as the qth image annotation history set (where q is a positive integer, and 1≤q≤Q). Subsequently, based on the preset image annotation multi-factor (including attribute features, posture features and additional features), a comprehensive feature traversal correlation evaluation is performed on the qth image annotation history set; by analyzing the correlation and mutual influence between each feature and other features, a ternary feature correlation evaluation set is generated, which details the correlation strength and importance ranking between features. Next, differential weight allocation is performed according to the image annotation multi-factor; by evaluating the criticality and contribution of each feature in the image annotation task, a labeling weight vector is generated, each element in the labeling weight vector corresponds to a feature, and its value reflects the relative importance of the feature in attention allocation. Finally, based on the generated labeling weight vector, combined with the ternary feature correlation evaluation set, the qth image annotation history set is subjected to attention optimization processing to obtain the qth attention optimized image set. Specifically, the attention weights of different feature regions are adjusted according to the values in the weight vector, so that the model can focus more on key features while suppressing the interference of irrelevant or secondary features.

[0044] The edge annotation multi-channel is cloud distilled by introducing the annotation loss iterative optimization mechanism to obtain a cloud annotation multi-channel.

[0045] The annotation loss iterative optimization mechanism identifies existing annotation loss or bias by continuously monitoring the annotation performance of the edge annotation multi-channel (including attribute feature annotation channel, posture feature annotation channel and additional feature annotation channel). Specifically, each annotation channel is iteratively evaluated, the loss value between its annotation result and the true label is calculated, and the optimization strategy is adjusted according to the size of the loss value.

[0046] Firstly, the multiple annotation channels obtained by the edge side through federated learning are uploaded to the cloud environment, wherein the multiple annotation channels include attribute feature annotation channel, pose feature annotation channel and additional feature annotation channel. For each annotation channel, a lightweight target model is introduced as a student network of the distillation model in the cloud, and the edge annotation channel is used as the teacher network to perform feature migration training by using knowledge distillation. During the training process, the output deviation of the student model and the teacher model on the same input sample is calculated respectively to obtain an initial distillation loss coefficient. If the distillation loss coefficient is greater than a set distillation loss threshold, the distillation loss iterative optimization training based on gradient back propagation is continued until it converges within the threshold range; if the threshold is met, the distillation channel training is completed, and the cloud annotation sub-channel is output. Finally, the cloud annotation sub-channels corresponding to the attribute features, pose features and additional features are integrated and fused to construct the completed cloud annotation multi-channel, so as to realize the optimization migration of the high-performance image annotation model with low computing resource consumption, and facilitate the subsequent automatic annotation task in large-scale data collection.

[0047] Further, the annotation loss iterative optimization mechanism includes:

[0048] The attribute feature annotation channel is subjected to cloud distillation loss iterative optimization to obtain an attribute annotation cloud channel; the pose feature annotation channel is subjected to cloud distillation loss iterative optimization to obtain a pose annotation cloud channel; the additional feature annotation channel is subjected to cloud distillation loss iterative optimization to obtain an additional annotation cloud channel; and the attribute annotation cloud channel, the pose annotation cloud channel and the additional annotation cloud channel are connected to generate the cloud annotation multi-channel.

[0049] Firstly, the attribute feature annotation channel of the edge side is uploaded to the cloud, and the teacher-student network architecture is used to perform knowledge distillation training on the cloud lightweight model. Through multiple rounds of supervised learning and distillation loss analysis, the training parameters are iteratively optimized until the generated attribute annotation student model meets the preset distillation loss threshold to form the final attribute annotation cloud channel. Similarly, the pose feature annotation channel is uploaded to the cloud, and the same distillation method is used for training and optimization to obtain the pose annotation cloud channel. For the additional feature annotation channel, the above distillation process is also used to complete the construction of the additional annotation cloud channel. After the optimization of the three types of cloud channels is completed, the attribute annotation cloud channel, the pose annotation cloud channel and the additional annotation cloud channel are structurally fused through a unified model interface to form the final cloud annotation multi-channel, which is used to support subsequent high-precision image annotation tasks and realize image semantic understanding and annotation automation in full feature dimension.

[0050] Further, the attribute feature annotation channel is subjected to cloud distillation loss iterative optimization to obtain an attribute annotation cloud channel, including:

[0051] According to the attribute feature annotation channel, knowledge distillation is performed on the cloud lightweight model to obtain an attribute annotation first migration channel; loss analysis is performed on the attribute annotation first migration channel according to the attribute feature annotation channel to obtain a first distillation loss coefficient; if the first distillation loss coefficient is greater than or equal to a distillation loss threshold, iterative optimization training of the distillation loss is performed on the attribute annotation first migration channel according to the attribute feature annotation channel, until the attribute annotation cloud channel less than the distillation loss threshold is generated.

[0052] First, the attribute feature annotation channel trained on the edge side is selected as a knowledge source to construct a cloud lightweight model (student model), and the attribute feature annotation channel is used as a teacher model to perform knowledge distillation on the student model to obtain an initial attribute annotation first migration channel; then, the attribute feature annotation result output by the first migration channel is used for distillation loss calculation to obtain a first distillation loss coefficient, which is used to measure the deviation degree of the student model and the teacher model in attribute feature expression; if the first distillation loss coefficient is greater than or equal to a preset distillation loss threshold, it is determined that the student model has not converged to the expected feature expression effect, and iterative optimization training of the first migration channel based on the original attribute feature annotation channel is needed; in each iteration, the student model is guided to approach the teacher model, and the distillation loss coefficient is evaluated in real time until the loss coefficient is less than the distillation loss threshold, at which time the construction of the attribute annotation cloud channel is completed, and the cloud attribute feature expression capability meeting the precision requirement is realized.

[0053] According to the cloud annotation multi-channel, the Q reinforced monitoring sets are automatically annotated.

[0054] The Q reinforced monitoring sets are automatically annotated by calling the cloud annotation multi-channel. Specifically, the attribute annotation cloud channel is used to identify the basic semantic category and appearance feature of the target object in the image, the pose annotation cloud channel is used to analyze the spatial pose information such as target pose, orientation and motion state, and the additional annotation cloud channel is used to supplement the context features or auxiliary description information related to the scene in the label; through the synergistic effect of the three types of annotation channels, full-range and multi-angle automatic annotation of each image frame is realized.

[0055] In summary, the embodiments of the present application have at least the following technical effects:

[0056] Firstly, according to Q edge monitoring nodes of a monitoring system, Q real-time monitoring sets are obtained, Q being a positive integer greater than 1. Next, according to an image quality analysis factor, image quality labeling learning is performed on the Q edge monitoring nodes to construct a collection quality labeling channel. Then, according to the collection quality labeling channel, the Q real-time monitoring sets are adaptively reinforced to obtain Q reinforced monitoring sets. Further, a differential attention learning mechanism is introduced, and image labeling multi-factor is combined to perform labeled feature federated learning on the Q edge monitoring nodes to obtain edge labeled multi-channel. A labeled loss iterative optimization mechanism is introduced to perform cloud distillation on the edge labeled multi-channel to obtain cloud labeled multi-channel. Finally, according to the cloud labeled multi-channel, the Q reinforced monitoring sets are automatically labeled. The technical problems of low monitoring image data labeling efficiency and insufficient accuracy in the prior art are solved, and the technical effects of improving image labeling efficiency and accuracy are achieved.

[0057] Embodiment two, based on the same inventive concept as the method for big data collection and labeling in the foregoing embodiments, as shown in the specification, the present application provides a system for big data collection and labeling, wherein the system comprises: Figure 2

[0058] The data collection module 11 obtains Q real-time monitoring sets according to Q edge monitoring nodes of a monitoring system, Q being a positive integer greater than 1. The labeling learning module 12 performs image quality labeling learning on the Q edge monitoring nodes according to an image quality analysis factor to construct a collection quality labeling channel. The data reinforcement module 13 adaptively reinforces the Q real-time monitoring sets according to the collection quality labeling channel to obtain Q reinforced monitoring sets. The federated learning module 14 introduces a differential attention learning mechanism and combines image labeling multi-factor to perform labeled feature federated learning on the Q edge monitoring nodes to obtain edge labeled multi-channel. The distillation optimization module 15 introduces a labeled loss iterative optimization mechanism to perform cloud distillation on the edge labeled multi-channel to obtain cloud labeled multi-channel. The automatic labeling module 16 automatically labels the Q reinforced monitoring sets according to the cloud labeled multi-channel.

[0059] Further, the labeling learning module 12 is configured to perform the following method:

[0060] Q image quality record sets of the Q edge monitoring nodes are obtained. Image quality labeling learning is performed on the Q image quality record sets according to the image quality analysis factor to obtain Q image quality detection models. The Q image quality detection models are evaluated in the same way according to the image quality analysis factor to obtain a global fidelity evaluation node, a local sharpness evaluation node and a connection smoothness evaluation node. The global fidelity evaluation node, the local sharpness evaluation node and the connection smoothness evaluation node are connected in parallel to generate the collection quality labeling channel.​

[0061] Further, the annotation learning module 12 is configured to perform the following method:

[0062] According to the Q image quality record sets, a first image quality record set is extracted, and the first image quality record set is clustered according to the image quality analysis factor to generate a first global fidelity record, a first local sharpness record and a first connection smoothness record; the image quality evaluator is supervised and trained according to the first global fidelity record, and a fidelity evaluation loss coefficient is obtained every predetermined number of times of training; if the fidelity evaluation loss coefficient is less than a fidelity evaluation loss threshold, a first global fidelity evaluation layer is generated; the first local sharpness record and the first connection smoothness record are continuously supervised and trained according to the image quality evaluator to generate a first local sharpness evaluation layer and a first connection smoothness evaluation layer; and the first global fidelity evaluation layer, the first local sharpness evaluation layer and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.

[0063] Further, the data reinforcement module 13 is configured to perform the following method:

[0064] According to the acquisition quality annotation channel, each image frame in the Q real-time monitoring sets is annotated for image quality to obtain a sequence of image quality of each frame; an image quality expectation sequence is constructed according to the image quality analysis factor; the sequence of image quality of each frame is analyzed for difference according to the image quality expectation sequence to obtain a plurality of image quality difference sequences; and the Q real-time monitoring sets are reinforced according to the plurality of image quality difference sequences to generate the Q reinforced monitoring sets.

[0065] Further, the federal learning module 14 is configured to perform the following method:

[0066] The image annotation multi-factor includes attribute features, posture features and additional features; based on the image annotation multi-factor, the Q image annotation history sets of the Q edge monitoring nodes are optimized for attention according to the differential attention learning mechanism to obtain Q attention-optimized image sets; the Q attention-optimized image sets are respectively supervised and trained to generate Q image annotators, wherein each image annotator includes an attribute feature annotation model, a posture feature annotation model and an additional feature annotation model; the same type of annotation model is aggregated in a federal manner according to the Q image annotators to generate the edge annotation multi-channel, and the edge annotation multi-channel includes an attribute feature annotation channel, a posture feature annotation channel and an additional feature annotation channel.

[0067] Further, the federal learning module 14 is configured to perform the following method:

[0068] According to the Q image label history set, the qth image label history set is extracted, q is a positive integer, 1≤q≤Q;According to the image label multi-element factor, the feature correlation evaluation set of the three elements is generated by performing feature correlation evaluation on the qth image label history set;According to the image label multi-element factor, the label weight vector is generated by performing differential weight allocation;Based on the label weight vector, the attention optimization is performed on the qth image label history set according to the three-element feature correlation evaluation set.

[0069] Further, the distillation optimization module 15 is used to perform the following method:

[0070] The attribute feature labeling channel is iteratively optimized by cloud distillation loss, and an attribute labeling cloud channel is obtained;The pose feature labeling channel is iteratively optimized by cloud distillation loss, and a pose labeling cloud channel is obtained;The additional feature labeling channel is iteratively optimized by cloud distillation loss, and an additional labeling cloud channel is obtained;The attribute labeling cloud channel, the pose labeling cloud channel and the additional labeling cloud channel are connected to generate the cloud labeling multi-channel.

[0071] Further, the distillation optimization module 15 is used to perform the following method:

[0072] According to the attribute feature labeling channel, the attribute labeling first migration channel is obtained by knowledge distillation of the cloud lightweight model;According to the attribute feature labeling channel, the first distillation loss coefficient is obtained by loss analysis of the attribute labeling first migration channel;If the first distillation loss coefficient is greater than or equal to the distillation loss threshold, the attribute labeling cloud channel is generated by iteratively optimizing and training the attribute labeling first migration channel according to the attribute feature labeling channel until the attribute labeling cloud channel is less than the distillation loss threshold.

[0073] Further, the labeling learning module 12 is used to perform the following method:

[0074] The image quality analysis factor includes global fidelity, local clarity and connection smoothness.

[0075] It should be noted that the above sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above describes a specific embodiment of the present application. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.

[0076] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0077] The specification and drawings are to be regarded in all respects as only illustrative and are to be construed in accordance with the scope of the application. It is evident that those skilled in the art can, without departing from the scope of the application, make various changes and modifications of the application. Thus, the present application is intended to embrace all such changes and modifications in the scope of the application and its equivalents.

Claims

1. A method for big data collection and annotation, characterized in that, The method comprises: According to the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, Q is a positive integer greater than 1; According to the image quality analysis factor, image quality labeling learning is performed on the Q edge monitoring nodes to construct a collection quality labeling channel; According to the collection quality labeling channel, the Q real-time monitoring sets are adaptively reinforced to obtain Q reinforced monitoring sets; A differential attention learning mechanism is introduced, and image labeling multi-factor is combined to perform labeled feature federated learning on the Q edge monitoring nodes to obtain an edge labeled multi-channel; An iterative optimization mechanism of labeling loss is introduced to perform cloud distillation on the edge labeled multi-channel to obtain a cloud labeled multi-channel; According to the cloud labeled multi-channel, the Q reinforced monitoring sets are automatically labeled; A differential attention learning mechanism is introduced, and image labeling multi-factor is combined to perform labeled feature federated learning on the Q edge monitoring nodes to obtain an edge labeled multi-channel, comprising: The image labeling multi-factor includes attribute features, posture features and additional features; Based on the image labeling multi-factor, the Q image labeling history sets of the Q edge monitoring nodes are optimized according to the differential attention learning mechanism to obtain Q attention optimized image sets; The Q attention optimized image sets are respectively supervised trained to generate Q image labelers, wherein each image labeler includes an attribute feature labeling model, a posture feature labeling model and an additional feature labeling model; According to the Q image labelers, the same type labeling model is federated aggregated to generate the edge labeled multi-channel, and the edge labeled multi-channel includes an attribute feature labeling channel, a posture feature labeling channel and an additional feature labeling channel.

2. The method for big data collection and annotation of claim 1, wherein, According to the image quality analysis factor, image quality labeling learning is performed on the Q edge monitoring nodes to construct a collection quality labeling channel, comprising: Q image quality record sets of the Q edge monitoring nodes are obtained; According to the image quality analysis factor, image quality labeling learning is performed on the Q image quality record sets to obtain Q image quality detection models; According to the image quality analysis factor, the same type evaluation layers of the Q image quality detection models are aggregated to obtain a global fidelity evaluation node, a local sharpness evaluation node and a connection smoothness evaluation node; The global fidelity evaluation node, the local sharpness evaluation node and the connection smoothness evaluation node are connected in parallel to generate the collection quality labeling channel.

3. The method for big data collection and annotation of claim 2, wherein, According to the image quality analysis factor, image quality labeling learning is performed on the Q image quality record sets, comprising: According to the Q image quality record sets, a first image quality record set is extracted, and the first image quality record set is clustered according to the image quality analysis factor to generate a first global fidelity record, a first local sharpness record and a first connection smoothness record; According to the first global fidelity record, an image quality evaluator is supervised trained, and a fidelity evaluation loss coefficient is obtained every predetermined number of times of training; If the fidelity evaluation loss coefficient is less than a fidelity evaluation loss threshold, a first global fidelity evaluation layer is generated; According to the image quality evaluator, loss supervised training is continued on the first local sharpness record and the first connection smoothness record, to generate a first local sharpness evaluation layer and a first connection smoothness evaluation layer; The first global fidelity evaluation layer, the first local sharpness evaluation layer, and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.

4. The method for big data collection and annotation of claim 1, wherein, According to the acquisition quality labeling channel, the Q real-time monitoring sets are adaptively reinforced to obtain Q reinforced monitoring sets, including: According to the acquisition quality labeling channel, image quality labels are added to each image frame in the Q real-time monitoring sets to obtain a sequence of image quality of each frame; According to the image quality analysis factor, an image quality expectation sequence is constructed; According to the image quality expectation sequence, difference analysis is performed on the sequence of image quality of each frame to obtain a plurality of image quality difference sequences; According to the plurality of image quality difference sequences, the Q real-time monitoring sets are reinforced to generate the Q reinforced monitoring sets.

5. The method for big data collection and annotation of claim 1, wherein, The differential attention learning mechanism includes: According to Q image labeling history sets, a qth image labeling history set is extracted, q is a positive integer, and 1≤q≤Q; According to the image labeling multi-element factor, a feature traversal correlation evaluation set is generated by performing feature traversal correlation evaluation on the qth image labeling history set; According to the image labeling multi-element factor, a labeling weight vector is generated by performing differential weight distribution; Based on the labeling weight vector, the qth image labeling history set is optimized by attention based on the feature traversal correlation evaluation set.

6. The method for big data collection and annotation of claim 1, wherein, The labeling loss iterative optimization mechanism includes: The attribute feature labeling channel is iteratively optimized by cloud distillation loss to obtain an attribute labeling cloud channel; The pose feature labeling channel is iteratively optimized by cloud distillation loss to obtain a pose labeling cloud channel; The additional feature labeling channel is iteratively optimized by cloud distillation loss to obtain an additional labeling cloud channel; The attribute labeling cloud channel, the pose labeling cloud channel, and the additional labeling cloud channel are connected to generate the cloud labeling multi-channel.

7. The method for big data collection and annotation of claim 6, wherein, The attribute feature labeling channel is iteratively optimized by cloud distillation loss to obtain an attribute labeling cloud channel, including: According to the attribute feature labeling channel, an attribute labeling first migration channel is obtained by knowledge distillation of a cloud lightweight model; According to the attribute feature labeling channel, a first distillation loss coefficient is obtained by loss analysis of the attribute labeling first migration channel; If the first distillation loss coefficient is greater than or equal to a distillation loss threshold, the attribute labeling first migration channel is iteratively optimized by distillation loss based on the attribute feature labeling channel until the attribute labeling cloud channel less than the distillation loss threshold is generated.

8. The method for big data collection and annotation of claim 1, wherein, The image quality analysis factor includes global fidelity, local sharpness, and connection smoothness.

9. A system for big data collection and annotation, characterized in that, A method for big data acquisition and labeling according to any one of claims 1-8, the system comprising: A data acquisition module: according to Q edge monitoring nodes of a monitoring system, Q real-time monitoring sets are obtained, Q is a positive integer greater than 1; The labeling learning module learns image quality labeling of the Q edge monitoring nodes according to the image quality analysis factor, and constructs a collection quality labeling channel; The data reinforcement module adaptively reinforces the Q real-time monitoring sets according to the collection quality labeling channel, and obtains Q reinforced monitoring sets; The federated learning module introduces a differential attention learning mechanism, and combines image labeling multi-element factors to perform labeled feature federated learning on the Q edge monitoring nodes, and obtains an edge labeled multi-channel; The distillation optimization module introduces a labeled loss iterative optimization mechanism to perform cloud distillation on the edge labeled multi-channel, and obtains a cloud labeled multi-channel; The automatic labeling module automatically labels the Q reinforced monitoring sets according to the cloud labeled multi-channel; The federated learning module introduces a differential attention learning mechanism, and combines image labeling multi-element factors to perform labeled feature federated learning on the Q edge monitoring nodes, and obtains an edge labeled multi-channel, including: The image labeling multi-element factors include attribute features, posture features, and additional features; Based on the image labeling multi-element factors, the Q image labeling history sets of the Q edge monitoring nodes are optimized by attention according to the differential attention learning mechanism, and Q attention-optimized image sets are obtained; The Q attention-optimized image sets are respectively supervised trained to generate Q image labelers, wherein each image labeler includes an attribute feature labeling model, a posture feature labeling model, and an additional feature labeling model; The same kind of labeling model is aggregated by the Q image labelers to generate the edge labeled multi-channel, and the edge labeled multi-channel includes an attribute feature labeling channel, a posture feature labeling channel, and an additional feature labeling channel.

Citation Information

Patent Citations

  • Model training method and related device

    CN115034836A

  • Multi-channel graph neural network pseudo tag selection method

    CN115526289A