Method and system for big data acquisition and annotation
By constructing an acquisition quality annotation channel and combining iterative optimization mechanism of differentiated attention learning and annotation loss, efficient and accurate automatic annotation of surveillance image data is achieved, solving the problems of low efficiency and insufficient accuracy in existing technologies.
Patent Information
- Application Number
- CN202510826175.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing technology has low efficiency and insufficient accuracy in labeling of surveillance image data, especially in a multi-node collaborative environment where the image acquisition quality is inconsistent and labeling consistency is difficult to ensure.
By constructing an acquisition quality annotation channel based on image quality analysis factors, introducing a differentiated attention learning mechanism and an iterative optimization mechanism for annotation loss, and combining image annotation multivariate factors to perform annotation feature federated learning and cloud distillation, automatic annotation is achieved.
The efficiency and accuracy of image annotation are improved, solving the problems of low efficiency and insufficient accuracy in monitoring image data annotation in the existing technology.
Smart Images

Figure CN120707987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image annotation, and in particular to a method and system for big data collection and annotation. Background Art
[0002] With the development of intelligent monitoring systems and big data analysis technologies, the real-time collection and high-quality annotation of massive video image data have become key links in the construction of intelligent analysis and training models. In typical scenarios such as traffic monitoring, security inspections, and urban management, the deployment of edge devices is becoming increasingly widespread, generating a large amount of heterogeneous and dynamically changing image data. In order to support the training and optimization of deep learning models, the collected images need to be accurately and comprehensively annotated. However, existing image annotation technologies mostly rely on manual or semi-automatic processes, which have problems such as low efficiency, high cost, and insufficient accuracy. Especially in a multi-node collaborative environment, the quality of image acquisition varies and the consistency of annotation is difficult to ensure. Summary of the Invention
[0003] The present application provides a method and system for big data collection and annotation, which solves the technical problems of low efficiency and insufficient accuracy in the annotation of surveillance image data in the prior art.
[0004] In a first aspect of the present application, a method for collecting and annotating big data is provided, the method comprising:
[0005] According to Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1; image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality resolution factor to construct an acquisition quality annotation channel; the Q real-time monitoring sets are adaptively enhanced according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets; a differentiated attention learning mechanism is introduced, and the annotation feature federation learning is performed on the Q edge monitoring nodes in combination with the image annotation multivariate factor to obtain edge annotation multi-channels; a annotation loss iterative optimization mechanism is introduced to perform cloud distillation on the edge annotation multi-channels to obtain cloud annotation multi-channels; and the Q enhanced monitoring sets are automatically annotated according to the cloud annotation multi-channels.
[0006] A second aspect of the present application provides a system for big data collection and annotation, the system comprising:
[0007] Data acquisition module: Based on the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1; annotation learning module: Based on the image quality analysis factor, image quality annotation learning is performed on the Q edge monitoring nodes to construct an acquisition quality annotation channel; data enhancement module: Based on the acquisition quality annotation channel, the Q real-time monitoring sets are adaptively enhanced to obtain Q enhanced monitoring sets; federated learning module: A differentiated attention learning mechanism is introduced to perform annotation feature federation learning on the Q edge monitoring nodes in combination with image annotation multivariate factors to obtain edge annotation multi-channels; distillation optimization module: A annotation loss iterative optimization mechanism is introduced to perform cloud-based distillation on the edge annotation multi-channels to obtain cloud-based annotation multi-channels; automatic annotation module: Based on the cloud-based annotation multi-channels, the Q enhanced monitoring sets are automatically annotated.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0009] First, based on the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1. Next, image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality analysis factor, and an acquisition quality annotation channel is constructed. Then, the Q real-time monitoring sets are adaptively enhanced according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets. Furthermore, a differentiated attention learning mechanism is introduced, and the annotation feature federation learning is performed on the Q edge monitoring nodes in combination with the multivariate factors of image annotation to obtain edge annotation multi-channels; an annotation loss iterative optimization mechanism is introduced to perform cloud distillation on the edge annotation multi-channels to obtain cloud annotation multi-channels. Finally, the Q enhanced monitoring sets are automatically annotated according to the cloud annotation multi-channels. The technical problems of low efficiency and insufficient accuracy of monitoring image data annotation in the existing technology are solved, and the technical effect of improving the efficiency and accuracy of image annotation is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1 A flowchart of a method for collecting and annotating big data provided in an embodiment of the present application;
[0012] Figure 2 A schematic diagram of the system structure for big data collection and annotation provided in an embodiment of the present application.
[0013] Explanation of the accompanying symbols: data acquisition module 11, annotation learning module 12, data enhancement module 13, federated learning module 14, distillation optimization module 15, automatic annotation module 16. DETAILED DESCRIPTION
[0014] This application solves the technical problems of low efficiency and insufficient accuracy in monitoring image data annotation in the prior art by providing a method and system for big data collection and annotation.
[0015] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0016] It should be noted that the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products or devices.
[0017] Example 1, as Figure 1 As shown, the present application provides a method for big data collection and annotation, wherein the method includes:
[0018] According to Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1.
[0019] In this embodiment, data collection is performed using Q edge monitoring nodes deployed in a monitoring system. These edge monitoring nodes are multiple front-end cameras or image sensing terminals with independent video acquisition capabilities, distributed across different locations within the monitored area. Each edge monitoring node can independently complete video image acquisition, preliminary data processing, and communication tasks.
[0020] Specifically, the system issues data collection instructions to all Q edge monitoring nodes through the access management module in accordance with a unified time synchronization protocol. After receiving the instructions, each edge monitoring node starts the video stream acquisition function and captures monitoring image frames in real time according to the preset frame rate (such as 10fps, 25fps, etc.) and image resolution (such as 720p, 1080p). Each node continuously collects image frame data within a certain period of time to form Q independent real-time monitoring sets.
[0021] Image quality annotation learning is performed on the Q edge monitoring nodes according to image quality resolution factors to construct an acquisition quality annotation channel. Further, the image quality resolution factors include global fidelity, local clarity, and connection smoothness.
[0022] Furthermore, image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality analysis factor to construct an acquisition quality annotation channel, including:
[0023] Obtain Q image quality record sets of the Q edge monitoring nodes; perform image quality annotation learning on the Q image quality record sets according to the image quality resolution factors to obtain Q image quality detection models; perform similar evaluation layer aggregation on the Q image quality detection models according to the image quality resolution factors to obtain global fidelity evaluation nodes, local clarity evaluation nodes, and connection smoothness evaluation nodes; and connect the global fidelity evaluation nodes, the local clarity evaluation nodes, and the connection smoothness evaluation nodes in parallel to generate the acquisition quality annotation channel.
[0024] Global fidelity is used to measure the degree of restoration between the overall image and the real scene, and is usually modeled through indicators such as color restoration, global contrast, and average brightness balance; local clarity reflects the clarity of local details in the image, and can be extracted and evaluated through parameters such as edge strength, local gradient changes, and texture feature density; connection smoothness is used to evaluate the spatial continuity and contour integrity of structural elements in the image, and is often calculated in combination with image connected areas, edge continuity, and spatial distribution consistency.
[0025] Corresponding image quality record sets are obtained from Q edge monitoring nodes, forming Q independent image quality record sets. Each image quality record set contains the historical image data collected by the corresponding node and its quality assessment metrics. Using global fidelity, local clarity, and connection smoothness as image quality analysis factors, feature extraction and annotation learning are performed on each image quality record set. A deep model for image quality recognition is constructed. During training, a multi-branch neural network structure is used, with each branch corresponding to a quality factor. Targeted feature optimization is achieved through independent supervised loss functions, resulting in Q image quality detection models. Each image quality detection model includes a global fidelity evaluation layer, a local clarity evaluation layer, and a connection smoothness evaluation layer, which are used to perform hierarchical and multi-faceted evaluations of different image quality factors. The global fidelity evaluation layers from all image quality detection models are aggregated and a unified global fidelity evaluation node is generated using layer weight alignment, distillation fusion, or average pooling. The local clarity evaluation layer and connection smoothness evaluation layer are processed in the same way, generating unified local clarity evaluation nodes and connection smoothness evaluation nodes, respectively. Finally, the global fidelity evaluation node, local clarity evaluation node, and connection smoothness evaluation node are connected in parallel to construct a multi-channel acquisition quality labeling channel, which is used to unify the three-factor quality score or quality label results of the output image frame.
[0026] Furthermore, performing image quality labeling learning on the Q image quality record sets according to the image quality analysis factor includes:
[0027] A first image quality record set is extracted based on the Q image quality record sets, and the first image quality record set is clustered according to the image quality analysis factor to generate a first global fidelity record, a first local clarity record and a first connection smoothness record; an image quality evaluator is supervised and trained based on the first global fidelity record, and a fidelity assessment loss coefficient is obtained each time a predetermined number of trainings are performed; if the fidelity assessment loss coefficient is less than a fidelity assessment loss threshold, a first global fidelity evaluation layer is generated; according to the image quality evaluator, loss supervised training is continued on the first local clarity record and the first connection smoothness record to generate a first local clarity evaluation layer and a first connection smoothness evaluation layer; the first global fidelity evaluation layer, the first local clarity evaluation layer and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.
[0028] Data corresponding to an edge monitoring node is selected from Q image quality record sets, and the image quality record of the node is extracted to form a first image quality record set. The first image quality record set is clustered based on the image quality resolution factor. Using a multidimensional feature clustering algorithm (e.g., a combination of principal component analysis and K-means clustering), the first image quality record set is divided into three subsets, corresponding to a first global fidelity record, a first local sharpness record, and a first connection smoothness record, to represent representative samples under the corresponding quality dimensions. Using the first global fidelity record as a training sample, an image quality evaluator is constructed and supervised training is performed. The image quality evaluator can be a convolutional neural network structure. The training process is performed for a predetermined number of iterations or rounds. At the end of each round of training, a fidelity assessment loss coefficient is calculated and compared with a set fidelity assessment loss threshold. When the loss coefficient decreases and stabilizes to less than the threshold, the evaluator is considered to have converged in terms of global fidelity learning. At this point, a first global fidelity evaluation layer can be constructed to indicate the image quality evaluator's ability to stably model global fidelity features. Continuing with the aforementioned image quality evaluator, based on the learned fidelity features, the first local sharpness record and the first connection smoothness record are introduced, respectively, for supervised training. By optimizing the corresponding sharpness loss function and smoothness loss function, the first local sharpness evaluation layer and the first connection smoothness evaluation layer are obtained, respectively. After the three evaluation layers are trained, they are connected in series or in a modular nested structure to form the first image quality detection model.
[0029] The Q real-time monitoring sets are adaptively enhanced according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets.
[0030] The acquisition quality annotation channel is used to assess the quality of each frame in the Q real-time surveillance sets. The annotation results for each frame are output in three dimensions: global fidelity, local clarity, and connection smoothness, forming a corresponding image quality annotation sequence. Based on a preset image quality expectation sequence (which reflects the desired image quality standard), the quality score of each frame is analyzed for differences, and the difference between each frame and the expected quality is calculated. Based on these difference values, an adaptive enhancement algorithm (such as dynamic weight adjustment or reinforcement learning strategy) is used to perform targeted processing on the real-time surveillance sets. This includes, but is not limited to, adjusting acquisition parameters (such as exposure time and focus settings), optimizing image processing algorithms (such as denoising and sharpening), or filtering low-quality frames. After this adaptive enhancement process, the image quality of the original Q real-time surveillance sets is significantly improved, ultimately generating Q enhanced surveillance sets.
[0031] Furthermore, the Q real-time monitoring sets are adaptively enhanced according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets, including:
[0032] Image quality annotation is performed on each image frame in the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain an image quality sequence for each frame; an image quality expectation sequence is constructed according to the image quality analysis factor; difference analysis is performed on the image quality sequences of each frame according to the image quality expectation sequence to obtain multiple image quality difference sequences; and enhancement processing is performed on the Q real-time monitoring sets according to the multiple image quality difference sequences to generate the Q enhanced monitoring sets.
[0033] First, using the constructed acquisition quality annotation pipeline, image quality assessment and annotation are performed on each image frame captured by each edge monitoring node in the Q real-time monitoring set. Quality annotation results are obtained for each frame along three dimensions: global fidelity, local sharpness, and connection smoothness. These are then organized chronologically into corresponding image quality sequences. Subsequently, an expected image quality sequence is constructed based on image quality analysis factors. This sequence reflects the desired image quality benchmarks for practical application scenarios, including but not limited to average sharpness standards, minimum fidelity thresholds, and edge coherence targets. Next, a frame-by-frame difference comparison is performed between the image quality sequence and the expected image quality sequence. Difference factors are extracted and analyzed to obtain deviation indicators for each image frame along the three quality dimensions. Finally, multiple image quality difference sequences are generated to quantify the degree of deviation from the ideal quality. Based on these difference sequences, the system performs adaptive enhancement processing, including but not limited to weighted reconstruction of image frames, re-sampling or replacement of low-quality frames, enhancement and compensation of high-quality frames, and dynamic adjustment of image acquisition parameters. This ensures that the overall monitoring image sequence is more consistent in terms of quality and meets the requirements of subsequent automatic annotation. Finally, Q enhanced monitoring sets are output.
[0034] A differentiated attention learning mechanism is introduced, and the annotation features of the Q edge monitoring nodes are federated and learned in combination with the multi-factor image annotation to obtain edge annotation multi-channels.
[0035] The differentiated attention learning mechanism can automatically adjust the attention weights of different regions or features according to the complexity, dynamics and other characteristics of the image content.
[0036] For each edge monitoring node, historical image annotation data is collected, covering multiple factors such as attribute features, posture features, and additional features. These features reflect the multidimensional information of the target object and environmental changes. Based on the differentiated attention learning mechanism, feature traversal and association evaluation are performed on the historical image annotation set of each node, and different weights are dynamically assigned to highlight key features while suppressing noise and redundant information, thereby optimizing the expression effect of the annotation features. On this basis, supervised training is performed on the attention-optimized image set of each edge node separately, and a local image annotator is constructed, which includes attribute feature annotation models, posture feature annotation models, and additional feature annotation models. Then, through the federated learning framework, similar annotation models of each edge node are aggregated, and a decentralized data collaborative update method is used to generate a comprehensive edge annotation multi-channel, which includes attribute feature annotation channels, posture feature annotation channels, and additional feature annotation channels.
[0037] Furthermore, a differentiated attention learning mechanism is introduced, and the annotation features of the Q edge monitoring nodes are federated and learned in combination with the multi-factor image annotation to obtain edge annotation multi-channels, including:
[0038] The image annotation multifactors include attribute features, posture features and additional features; based on the image annotation multifactors, the Q image annotation history sets of the Q edge monitoring nodes are optimized according to the differentiated attention learning mechanism to obtain Q attention-optimized image sets; the Q attention-optimized image sets are supervised and trained separately to generate Q image annotators, wherein each image annotator includes an attribute feature annotation model, a posture feature annotation model and an additional feature annotation model; similar annotation models are federated and aggregated according to the Q image annotators to generate the edge annotation multi-channel, which includes an attribute feature annotation channel, a posture feature annotation channel and an additional feature annotation channel.
[0039] The multivariate factors of image annotation include attribute features (such as color and shape), posture features (such as object position and direction) and additional features (such as background information and lighting conditions). These multidimensional features together reflect the diverse information of the target in the image.
[0040] Based on the multivariate factors of image annotation, a differentiated attention learning mechanism is used to optimize the attention of Q image annotation histories of Q edge monitoring nodes. This mechanism dynamically adjusts the attention weights of different feature regions, enabling the model to focus more on key features, thereby generating Q attention-optimized image sets. Each attention-optimized image set highlights the importance of the corresponding edge monitoring node on a specific annotation feature. Supervised training is performed on each of these Q attention-optimized image sets to construct corresponding Q image annotators. Each image annotator includes an attribute feature annotation model, a pose feature annotation model, and an additional feature annotation model. Finally, similar annotation models are fused using a federated aggregation method (such as the federated averaging algorithm) to generate a multi-channel edge annotation system. This multi-channel includes an attribute feature annotation channel, a pose feature annotation channel, and an additional feature annotation channel, achieving collaborative learning and fusion of multi-dimensional annotation features. Specifically, the attribute feature annotation models from the Q image annotators are federated to generate an attribute feature annotation channel. Similarly, pose feature annotation channels and additional feature annotation channels are generated.
[0041] Furthermore, the differentiated attention learning mechanism includes:
[0042] According to Q image annotation history sets, extract the qth image annotation history set, where q is a positive integer, 1≤q≤Q; perform feature traversal association evaluation on the qth image annotation history set according to the image annotation multivariate factors to generate a ternary feature association evaluation set; perform differentiated weight allocation according to the image annotation multivariate factors to generate an annotation weight vector; based on the annotation weight vector, perform attention optimization on the qth image annotation history set according to the ternary feature association evaluation set.
[0043] When implementing the differentiated attention learning mechanism, an image annotation history set is first extracted from Q image annotation history sets, serving as the qth image annotation history set (where q is a positive integer and 1≤q≤Q). Subsequently, a comprehensive feature traversal association evaluation is performed on the qth image annotation history set based on pre-set image annotation multivariate factors (including attribute features, pose features, and additional features). By analyzing the correlation and mutual influence between each feature and other features, a ternary feature association evaluation set is generated, which details the strength of the associations and importance rankings between the features. Next, differentiated weight allocation is performed based on the image annotation multivariate factors. By evaluating the criticality and contribution of each feature in the image annotation task, a labeling weight vector is generated. Each element in the labeling weight vector corresponds to a feature, and its value reflects the relative importance of that feature in attention allocation. Finally, based on the generated labeling weight vector and combined with the ternary feature association evaluation set, attention optimization is performed on the qth image annotation history set, resulting in the qth attention-optimized image set. Specifically, the attention weights of different feature regions are adjusted according to the values in the weight vector, enabling the model to focus more on key features while suppressing interference from irrelevant or minor features.
[0044] An iterative optimization mechanism of labeling loss is introduced to perform cloud-based distillation on the edge labeling multi-channel to obtain the cloud-based labeling multi-channel.
[0045] The iterative optimization mechanism for annotation loss continuously monitors the annotation performance of multiple edge annotation channels (including attribute feature annotation channels, pose feature annotation channels, and additional feature annotation channels) to identify any annotation loss or deviation. Specifically, it iteratively evaluates each annotation channel, calculates the loss between its annotation results and the true label, and adjusts the optimization strategy based on the magnitude of the loss.
[0046] First, multiple annotation channels obtained through federated learning on the edge are uploaded to the cloud environment. These channels include attribute feature annotation channels, pose feature annotation channels, and additional feature annotation channels. For each annotation channel, a lightweight target model is introduced on the cloud as the student network of the distillation model, and the edge annotation channel is used as the teacher network. Feature transfer training is performed using knowledge distillation. During training, the output deviation between the student and teacher models on the same input sample is calculated to obtain the initial distillation loss coefficient. If the distillation loss coefficient exceeds the set distillation loss threshold, iterative optimization training using gradient backpropagation is performed until convergence within the threshold range. If the threshold is met, the distillation channel training is completed, and the output is the cloud-based annotation sub-channel. Finally, the cloud-based annotation sub-channels corresponding to the attribute features, pose features, and additional features are integrated and fused to construct a complete cloud-based annotation multi-channel. This achieves optimized transfer of high-performance image annotation models with low computational resource consumption, facilitating subsequent automatic annotation tasks in large-scale data collection.
[0047] Furthermore, the iterative optimization mechanism for labeling loss includes:
[0048] Perform iterative optimization of cloud distillation loss on the attribute feature labeling channel to obtain the attribute labeling cloud channel; perform iterative optimization of cloud distillation loss on the posture feature labeling channel to obtain the posture labeling cloud channel; perform iterative optimization of cloud distillation loss on the additional feature labeling channel to obtain the additional labeling cloud channel; connect the attribute labeling cloud channel, the posture labeling cloud channel and the additional labeling cloud channel to generate the cloud labeling multi-channel.
[0049] First, the attribute feature annotation channel on the edge is uploaded to the cloud. The cloud-based lightweight model is trained for knowledge distillation using a teacher-student network architecture. Through multiple rounds of supervised learning and distillation loss analysis, the training parameters are continuously optimized iteratively until the generated attribute annotation student model meets the preset distillation loss threshold, forming the final attribute annotation cloud channel. Similarly, the posture feature annotation channel is uploaded to the cloud and trained and optimized using the same distillation method to obtain the posture annotation cloud channel. For the additional feature annotation channel, the above distillation process is also used to complete the construction of the additional annotation cloud channel. After the optimization of all three types of cloud channels is completed, the attribute annotation cloud channel, the posture annotation cloud channel, and the additional annotation cloud channel are structurally fused through a unified model interface to form the final cloud annotation multi-channel, which is used to support the subsequent high-precision image annotation task execution and realize image semantic understanding and annotation automation in all feature dimensions.
[0050] Furthermore, the attribute feature annotation channel is iteratively optimized using cloud distillation loss to obtain the attribute annotation cloud channel, including:
[0051] Perform knowledge distillation on the cloud-based lightweight model according to the attribute feature labeling channel to obtain a first attribute labeling migration channel; perform loss analysis on the first attribute labeling migration channel according to the attribute feature labeling channel to obtain a first distillation loss coefficient; if the first distillation loss coefficient is greater than or equal to a distillation loss threshold, perform distillation loss iterative optimization training on the first attribute labeling migration channel according to the attribute feature labeling channel until the attribute labeling cloud-based channel with a loss less than the distillation loss threshold is generated.
[0052] First, the attribute feature labeling channel completed by edge-side training is selected as the knowledge source to construct a cloud-based lightweight model (student model). The attribute feature labeling channel is used as the teacher model to perform knowledge distillation on the student model to obtain the initial attribute labeling first migration channel. Subsequently, the attribute feature labeling result output by the first migration channel is subjected to distillation loss calculation to obtain the first distillation loss coefficient, which is used to measure the degree of deviation between the student model and the teacher model in attribute feature expression. If the first distillation loss coefficient is greater than or equal to the preset distillation loss threshold, it is determined that the student model has not yet converged to the desired feature expression effect, and it is necessary to continue to perform iterative optimization training on the first migration channel based on the original attribute feature labeling channel. In each round of iteration, the student model is continuously guided to approach the teacher model, and its distillation loss coefficient is evaluated in real time until the loss coefficient is less than the distillation loss threshold. At this time, the construction of the attribute labeling cloud channel is completed, and the cloud-based attribute feature representation capability with satisfactory accuracy is achieved.
[0053] The Q enhanced monitoring sets are automatically labeled according to the cloud-based labeling multi-channel.
[0054] The Q enhanced monitoring sets are automatically annotated by calling multiple cloud-based annotation channels. Specifically, the attribute annotation cloud-based channel is used to identify the basic semantic categories and appearance features of the target objects in the image; the posture annotation cloud-based channel is used to analyze spatial posture information such as the target posture, orientation, and motion state; and the additional annotation cloud-based channel is used to supplement the scene-related contextual features or auxiliary descriptive information in the label. Through the synergistic effect of the three types of annotation channels, all-round and multi-angle automatic annotation of each image frame is achieved.
[0055] In summary, the embodiments of the present application have at least the following technical effects:
[0056] First, based on the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1. Next, image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality analysis factor, and an acquisition quality annotation channel is constructed. Then, the Q real-time monitoring sets are adaptively enhanced according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets. Furthermore, a differentiated attention learning mechanism is introduced, and the annotation feature federation learning is performed on the Q edge monitoring nodes in combination with the multivariate factors of image annotation to obtain edge annotation multi-channels; an annotation loss iterative optimization mechanism is introduced to perform cloud distillation on the edge annotation multi-channels to obtain cloud annotation multi-channels. Finally, the Q enhanced monitoring sets are automatically annotated according to the cloud annotation multi-channels. The technical problems of low efficiency and insufficient accuracy of monitoring image data annotation in the existing technology are solved, and the technical effect of improving the efficiency and accuracy of image annotation is achieved.
[0057] Example 2, based on the same inventive concept as the method for big data collection and annotation in the above embodiment, Figure 2 As shown, the present application provides a system for big data collection and annotation, wherein the system includes:
[0058] Data acquisition module 11: Based on Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1; annotation learning module 12: Image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality analysis factor to construct an acquisition quality annotation channel; data enhancement module 13: Adaptively enhances the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets; federated learning module 14: Introduces a differentiated attention learning mechanism, combines image annotation multivariate factors to perform annotation feature federation learning on the Q edge monitoring nodes, and obtains edge annotation multi-channels; distillation optimization module 15: Introduces an annotation loss iterative optimization mechanism to perform cloud-based distillation on the edge annotation multi-channels to obtain cloud-based annotation multi-channels; automatic annotation module 16: Automatically annotates the Q enhanced monitoring sets according to the cloud-based annotation multi-channels.
[0059] Furthermore, the annotation learning module 12 is used to perform the following method:
[0060] Obtain Q image quality record sets of the Q edge monitoring nodes; perform image quality annotation learning on the Q image quality record sets according to the image quality resolution factors to obtain Q image quality detection models; perform similar evaluation layer aggregation on the Q image quality detection models according to the image quality resolution factors to obtain global fidelity evaluation nodes, local clarity evaluation nodes, and connection smoothness evaluation nodes; and connect the global fidelity evaluation nodes, the local clarity evaluation nodes, and the connection smoothness evaluation nodes in parallel to generate the acquisition quality annotation channel.
[0061] Furthermore, the annotation learning module 12 is used to perform the following method:
[0062] A first image quality record set is extracted based on the Q image quality record sets, and the first image quality record set is clustered according to the image quality analysis factor to generate a first global fidelity record, a first local clarity record and a first connection smoothness record; an image quality evaluator is supervised and trained based on the first global fidelity record, and a fidelity assessment loss coefficient is obtained each time a predetermined number of trainings are performed; if the fidelity assessment loss coefficient is less than a fidelity assessment loss threshold, a first global fidelity evaluation layer is generated; according to the image quality evaluator, loss supervised training is continued on the first local clarity record and the first connection smoothness record to generate a first local clarity evaluation layer and a first connection smoothness evaluation layer; the first global fidelity evaluation layer, the first local clarity evaluation layer and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.
[0063] Furthermore, the data enhancement module 13 is configured to perform the following method:
[0064] Image quality annotation is performed on each image frame in the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain an image quality sequence for each frame; an image quality expectation sequence is constructed according to the image quality analysis factor; difference analysis is performed on the image quality sequences of each frame according to the image quality expectation sequence to obtain multiple image quality difference sequences; and enhancement processing is performed on the Q real-time monitoring sets according to the multiple image quality difference sequences to generate the Q enhanced monitoring sets.
[0065] Furthermore, the federated learning module 14 is configured to perform the following method:
[0066] The image annotation multifactors include attribute features, posture features and additional features; based on the image annotation multifactors, the Q image annotation history sets of the Q edge monitoring nodes are optimized according to the differentiated attention learning mechanism to obtain Q attention-optimized image sets; the Q attention-optimized image sets are supervised and trained separately to generate Q image annotators, wherein each image annotator includes an attribute feature annotation model, a posture feature annotation model and an additional feature annotation model; similar annotation models are federated and aggregated according to the Q image annotators to generate the edge annotation multi-channel, which includes an attribute feature annotation channel, a posture feature annotation channel and an additional feature annotation channel.
[0067] Furthermore, the federated learning module 14 is configured to perform the following method:
[0068] According to Q image annotation history sets, extract the qth image annotation history set, where q is a positive integer, 1≤q≤Q; perform feature traversal association evaluation on the qth image annotation history set according to the image annotation multivariate factors to generate a ternary feature association evaluation set; perform differentiated weight allocation according to the image annotation multivariate factors to generate an annotation weight vector; based on the annotation weight vector, perform attention optimization on the qth image annotation history set according to the ternary feature association evaluation set.
[0069] Furthermore, the distillation optimization module 15 is configured to perform the following method:
[0070] Perform iterative optimization of cloud distillation loss on the attribute feature labeling channel to obtain the attribute labeling cloud channel; perform iterative optimization of cloud distillation loss on the posture feature labeling channel to obtain the posture labeling cloud channel; perform iterative optimization of cloud distillation loss on the additional feature labeling channel to obtain the additional labeling cloud channel; connect the attribute labeling cloud channel, the posture labeling cloud channel and the additional labeling cloud channel to generate the cloud labeling multi-channel.
[0071] Furthermore, the distillation optimization module 15 is configured to perform the following method:
[0072] Perform knowledge distillation on the cloud-based lightweight model according to the attribute feature labeling channel to obtain a first attribute labeling migration channel; perform loss analysis on the first attribute labeling migration channel according to the attribute feature labeling channel to obtain a first distillation loss coefficient; if the first distillation loss coefficient is greater than or equal to a distillation loss threshold, perform distillation loss iterative optimization training on the first attribute labeling migration channel according to the attribute feature labeling channel until the attribute labeling cloud-based channel with a loss less than the distillation loss threshold is generated.
[0073] Furthermore, the annotation learning module 12 is used to perform the following method:
[0074] The image quality analysis factors include global fidelity, local clarity and connection smoothness.
[0075] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0076] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0077] This specification and drawings are merely illustrative of the present application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Obviously, those skilled in the art may make various modifications and variations to this application without departing from the scope of this application. Thus, this application is intended to include such modifications and variations as fall within the scope of this application and its equivalents.
Claims
1. A method for big data collection and annotation, characterized in that: The method comprises: According to the Q edge monitoring nodes of the monitoring system, Q real-time monitoring sets are obtained, where Q is a positive integer greater than 1; Perform image quality annotation learning on the Q edge monitoring nodes according to the image quality analysis factor to construct an acquisition quality annotation channel; Adaptively enhancing the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets; A differentiated attention learning mechanism is introduced, and the annotation features of the Q edge monitoring nodes are federated and learned in combination with the multi-factor image annotation to obtain edge annotation multi-channels. Introducing an iterative optimization mechanism for labeling loss to perform cloud distillation on the edge labeling multi-channel to obtain a cloud labeling multi-channel; The Q enhanced monitoring sets are automatically labeled according to the cloud-based labeling multi-channel.
2. A method for big data collection and annotation according to claim 1, characterized in that: Image quality annotation learning is performed on the Q edge monitoring nodes according to the image quality analysis factor, and an acquisition quality annotation channel is constructed, including: Obtaining Q image quality record sets of the Q edge monitoring nodes; Performing image quality annotation learning on the Q image quality record sets according to the image quality analysis factors to obtain Q image quality detection models; Performing similar evaluation layer aggregation on the Q image quality detection models according to the image quality analysis factor to obtain a global fidelity evaluation node, a local clarity evaluation node, and a connection smoothness evaluation node; The global fidelity evaluation node, the local clarity evaluation node, and the connection smoothness evaluation node are connected in parallel to generate the acquisition quality annotation channel.
3. A method for big data collection and annotation according to claim 2, characterized in that: Performing image quality labeling learning on the Q image quality record sets according to the image quality analysis factors includes: Extracting a first image quality record set from the Q image quality record sets, and performing clustering processing on the first image quality record set according to the image quality resolution factor to generate a first global fidelity record, a first local clarity record, and a first connection smoothness record; Performing supervised training on an image quality assessor according to the first global fidelity record, and obtaining a fidelity assessment loss coefficient after each predetermined number of trainings; If the fidelity assessment loss coefficient is less than the fidelity assessment loss threshold, generating a first global fidelity evaluation layer; According to the image quality evaluator, continue to perform loss supervision training on the first local clarity record and the first connection smoothness record to generate a first local clarity evaluation layer and a first connection smoothness evaluation layer; The first global fidelity evaluation layer, the first local clarity evaluation layer, and the first connection smoothness evaluation layer are connected to generate a first image quality detection model.
4. The method for big data collection and annotation according to claim 1, wherein: Adaptively enhancing the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets, including: Performing image quality annotation on each image frame in the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain an image quality sequence for each frame; constructing an image quality expectation sequence according to the image quality analysis factor; performing difference analysis on the image quality sequences of each frame according to the expected image quality sequence to obtain a plurality of image quality difference sequences; The Q real-time monitoring sets are enhanced according to the multiple image quality difference sequences to generate the Q enhanced monitoring sets.
5. The method for big data collection and annotation according to claim 1, wherein: A differentiated attention learning mechanism is introduced, and the annotation features of the Q edge monitoring nodes are federated and learned in combination with the multi-factor image annotation to obtain edge annotation multi-channels, including: The image annotation multivariate factors include attribute features, posture features and additional features; Based on the image annotation multivariate factor, performing attention optimization on the Q image annotation history sets of the Q edge monitoring nodes according to the differentiated attention learning mechanism to obtain Q attention-optimized image sets; Performing supervised training on the Q attention-optimized image sets to generate Q image annotators, each of which includes an attribute feature annotation model, a posture feature annotation model, and an additional feature annotation model; The same type of annotation models are federated and aggregated according to the Q image annotators to generate the edge annotation multi-channel, which includes an attribute feature annotation channel, a posture feature annotation channel, and an additional feature annotation channel.
6. The method for big data collection and annotation according to claim 1, wherein: The differentiated attention learning mechanism includes: According to the Q image annotation history sets, extract the qth image annotation history set, where q is a positive integer, 1≤q≤Q; Performing feature traversal association evaluation on the qth image annotation history set according to the image annotation multivariate factor to generate a ternary feature association evaluation set; Performing differentiated weight allocation according to the multivariate factors of image annotation to generate an annotation weight vector; Based on the annotation weight vector, attention optimization is performed on the qth image annotation history set according to the ternary feature association evaluation set.
7. The method for big data collection and annotation according to claim 1, wherein: The labeling loss iterative optimization mechanism includes: Perform iterative optimization of the cloud distillation loss on the attribute feature annotation channel to obtain the attribute annotation cloud channel; Perform iterative optimization of the cloud distillation loss on the pose feature annotation channel to obtain the pose annotation cloud channel; Perform iterative optimization of the cloud distillation loss on the additional feature annotation channel to obtain the additional annotation cloud channel; The attribute annotation cloud channel, the posture annotation cloud channel, and the additional annotation cloud channel are connected to generate the cloud annotation multi-channel.
8. The method for big data collection and annotation according to claim 7, wherein: Perform iterative optimization of the cloud distillation loss on the attribute feature annotation channel to obtain the attribute annotation cloud channel, including: Performing knowledge distillation on the cloud-based lightweight model according to the attribute feature annotation channel to obtain a first attribute annotation migration channel; performing a loss analysis on the attribute-annotated first migration channel according to the attribute feature-annotated channel to obtain a first distillation loss coefficient; If the first distillation loss coefficient is greater than or equal to a distillation loss threshold, distillation loss iterative optimization training is performed on the attribute annotation first migration channel according to the attribute feature annotation channel until the attribute annotation cloud channel with a loss less than the distillation loss threshold is generated.
9. The method for big data collection and annotation according to claim 1, wherein: The image quality analysis factors include global fidelity, local clarity and connection smoothness.
10. A system for big data collection and annotation, characterized in that: A system for implementing a method for big data collection and annotation according to any one of claims 1 to 9, comprising: Data acquisition module: obtains Q real-time monitoring sets based on the Q edge monitoring nodes of the monitoring system, where Q is a positive integer greater than 1; Annotation learning module: performs image quality annotation learning on the Q edge monitoring nodes according to the image quality analysis factor, and constructs an acquisition quality annotation channel; Data enhancement module: adaptively enhances the Q real-time monitoring sets according to the acquisition quality annotation channel to obtain Q enhanced monitoring sets; Federated learning module: Introducing a differentiated attention learning mechanism, combined with image annotation multi-factors, performs federated learning of annotation features on the Q edge monitoring nodes to obtain edge annotation multi-channels; Distillation optimization module: Introducing an iterative optimization mechanism for annotation loss to perform cloud-based distillation on the edge annotation multi-channel to obtain cloud-based annotation multi-channel; Automatic labeling module: automatically labels the Q enhanced monitoring sets according to the cloud-based labeling multi-channel.
Citation Information
Patent Citations
Model training method and related device
CN115034836A
Multi-channel graph neural network pseudo tag selection method
CN115526289A
Federal learning backdoor defense method based on attention distillation
CN115630361A
Personalized federal learning method and system
CN119476414A
Automatic data annotation method of ISP image signal processing visual sensor
CN119723010A