Data labeling method and device, computer equipment and readable storage medium

Through the computer vision model, the target objects in the image dataset are automatically labeled, which solves the problem of low manual labeling efficiency in the existing technology, realizes efficient and automated data labeling, and improves the training effect of the image recognition model.

CN119942260APending Publication Date: 2025-05-06SUGON INFORMATION IND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411806328.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing data annotation methods rely on manual annotation, which is longer and has low efficiency, resulting in a smaller amount of data in the target image dataset.

Method used

Computer vision models are used to automatically label the target objects in the image dataset to generate the target image dataset, and used to train the image recognition model.

Benefits of technology

It realizes automatic labeling of computer room image data in each computer room, avoids manual participation, improves the efficiency of data labeling methods, expands the amount of data, and enhances the training effect of image recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942260A_ABST
    Figure CN119942260A_ABST
Patent Text Reader

Abstract

The invention relates to a data labeling method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: receiving an image data set sent by a monitoring server; marking a target object in the image data set according to a computer vision model to obtain a target image data set; sending the target image data set to a monitoring server; the target image data set is used for training an image recognition model on the monitoring server. By adopting the method, the efficiency of the data labeling method can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and in particular to a data annotation method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In order to monitor the safety of the computer room, the image recognition model needs to be deployed to the monitoring server so that the image recognition model can monitor whether there are risks in the computer room based on the real-time collected image data of each computer room. In order to ensure the accuracy of the image recognition model, after the monitoring server is deployed, it is necessary to generate a target image dataset based on the data annotation method and train the image recognition model based on the target image dataset.

[0003] The current data annotation method obtains the image data of each computer room. Then, for each computer room image data, according to the annotation requirements, the staff annotates the target objects with potential safety hazards in the initial computer room image data, obtains the target computer room image data, and constructs the target image data set based on the image data of each target computer room.

[0004] However, the current data annotation method takes a long time to manually annotate each initial image data of the computer room, which is inefficient and the amount of data in the target image data set obtained is small. Therefore, the current data annotation method is inefficient. Summary of the invention

[0005] Based on this, it is necessary to provide a data labeling method, apparatus, computer equipment, computer-readable storage medium and computer program product to address the above technical problems.

[0006] In a first aspect, the present application provides a data labeling method, which is applied to a labeling server, and the method includes:

[0007] Receiving an image data set sent by a monitoring server;

[0008] Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset;

[0009] The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0010] In the above data labeling method, the target objects in the image data set are labeled through the computer vision model to obtain the target image data set, which realizes the automatic labeling of the computer room image data of each computer room, avoids manual labeling, and improves the efficiency of the data labeling method.

[0011] In one embodiment, the image data set includes image data of each computer room and target prompt words, and the target objects in the image data set are labeled according to the computer vision model to obtain the target image data set, including:

[0012] Based on a computer vision model, a target object corresponding to the target prompt word in each of the computer room image data is labeled to obtain mask data corresponding to the target object;

[0013] A target image data set is constructed according to each of the computer room image data and the mask data corresponding to each of the computer room image data.

[0014] In this embodiment, a computer vision model is used to annotate mask data corresponding to the target prompt word in each computer room image data set, thereby obtaining the object that the image recognition model needs to recognize, avoiding manual participation, realizing automatic labeling, and improving the efficiency of the data labeling method.

[0015] In one embodiment, the computer vision model includes a first computer vision model and a second computer vision model, and the target object corresponding to the target prompt word in each of the computer room image data is labeled based on the computer vision model to obtain mask data corresponding to the target object, including:

[0016] Locating the target object corresponding to the target prompt word in each of the computer room image data according to the first computer vision model, and obtaining the target frame data corresponding to the computer room image data;

[0017] Based on the second computer vision model and the target frame data corresponding to the computer room image data, image segmentation processing is performed on each of the computer room image data to obtain mask data corresponding to each of the computer room image data.

[0018] In this embodiment, the image data of each computer room are processed through the first computer vision model and the second computer vision model to obtain each mask data, clarify the position of the target object in each picture data, and realize automatic labeling of the image data of each computer room, avoiding manual participation, improving data security, and at the same time improving the labeling efficiency, thereby improving the efficiency of the data labeling method.

[0019] In one embodiment, constructing a target image data set according to each of the computer room image data and the mask data corresponding to each of the computer room image data includes:

[0020] Adding the mask data corresponding to each of the computer room image data to the computer room image data to obtain the target computer room image data;

[0021] The target computer room image data are combined to obtain a target image data set.

[0022] In this embodiment, by adding mask data to the computer room image data, target computer room image data for training the image recognition model is obtained, and a target image data set is constructed based on each target computer room image data, thereby achieving automatic labeling of the image data set and improving the labeling accuracy and efficiency.

[0023] In a second aspect, the present application provides a data labeling method, which is applied to a monitoring server, and the method includes:

[0024] Obtaining initial computer room image data and initial sample recognition templates of each computer room;

[0025] Building an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and sending the image data set to a labeling server, so that the labeling server labels the image data set through a computer vision model to obtain a target image data set;

[0026] The target image data set sent by the annotation server is received, and an image recognition model is trained according to the target image data set to obtain a target image recognition model.

[0027] In this embodiment, an image dataset is constructed and sent to a labeling server so that the labeling server automatically labels the image dataset through a computer vision model, thereby improving the efficiency of labeling the image dataset. The image recognition model is then trained based on the labeled target image dataset, thereby improving the training efficiency.

[0028] In one embodiment, the step of constructing an image data set based on the initial sample recognition template and the initial computer room image data of each computer room includes:

[0029] Determining each computer room image data in each initial computer room image data of each of the computer rooms;

[0030] Determining a target prompt word from each default prompt word in the initial sample recognition template;

[0031] An image data set is constructed according to each of the computer room image data and the target prompt words.

[0032] In this embodiment, by determining each computer room image data in each initial computer room image data and determining the target prompt word according to the initial sample recognition template, personalized determination of the target object is achieved, and the application scope of the data annotation method is expanded.

[0033] In one embodiment, the step of training the image recognition model according to the target image data set to obtain the target image recognition model comprises:

[0034] Acquire the computer room video data of each of the computer rooms in real time, and perform frame extraction processing on the computer room video data of each of the computer rooms to obtain each frame of computer room image data to be monitored;

[0035] Combining the image data of the to-be-monitored computer rooms of the same frame to obtain image data sets of the to-be-monitored computer rooms;

[0036] Image recognition is performed on each of the to-be-monitored computer room image data sets based on the target image recognition model to obtain an image recognition result, and the safety status of each of the computer rooms is determined based on the image recognition result.

[0037] In this embodiment, image recognition is performed on each image data set of the computer room to be monitored through the target recognition model, and the safety status of the computer room is determined according to the image recognition result, thereby realizing automatic monitoring of the safety of the computer room and improving the monitoring efficiency.

[0038] In a third aspect, the present application also provides a data labeling device, including:

[0039] A receiving module, used for receiving an image data set sent by a monitoring server;

[0040] A labeling module, used for labeling the target object in the image data set according to a computer vision model to obtain a target image data set;

[0041] The sending module is used to send the target image data set to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0042] In a fourth aspect, the present application also provides a data labeling device, including:

[0043] An acquisition device, used to acquire initial computer room image data and initial sample recognition templates of each computer room;

[0044] A construction device is used to construct an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and send the image data set to a labeling server, so that the labeling server labels the computer room image data set through a computer vision model to obtain a target image data set;

[0045] The training device is used to receive the target image data set sent by the annotation server, and train the image recognition model according to the target image data set to obtain the target image recognition model.

[0046] In a fifth aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0047] Receiving an image data set sent by a monitoring server;

[0048] Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset;

[0049] The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0050] In a sixth aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0051] Receiving an image data set sent by a monitoring server;

[0052] Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset;

[0053] The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0054] In a seventh aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0055] Receiving an image data set sent by a monitoring server;

[0056] Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset;

[0057] The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0058] In an eighth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0059] Obtaining initial computer room image data and initial sample recognition templates of each computer room;

[0060] Building an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and sending the image data set to a labeling server, so that the labeling server labels the image data set through a computer vision model to obtain a target image data set;

[0061] The target image data set sent by the annotation server is received, and an image recognition model is trained according to the target image data set to obtain a target image recognition model.

[0062] In a ninth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0063] Receiving an image data set sent by a monitoring server;

[0064] Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset;

[0065] The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0066] In a tenth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0067] Obtaining initial computer room image data and initial sample recognition templates of each computer room;

[0068] Building an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and sending the image data set to a labeling server, so that the labeling server labels the image data set through a computer vision model to obtain a target image data set;

[0069] The target image data set sent by the annotation server is received, and an image recognition model is trained according to the target image data set to obtain a target image recognition model.

[0070] The above-mentioned data annotation method, device, computer equipment, computer-readable storage medium and computer program product receive an image data set sent by a monitoring server; annotate the target object in the image data set according to a computer vision model to obtain a target image data set; send the target image data set to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server. With this method, the target object in the image data set is annotated by a computer vision model to obtain a target image data set, thereby realizing automatic annotation of the image data of each computer room, avoiding manual annotation, and improving the efficiency of the data annotation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0072] Figure 1 An application environment diagram of a data annotation method in an embodiment;

[0073] Figure 2 A schematic diagram of a flow chart of a data labeling method in an embodiment;

[0074] Figure 3 A schematic diagram of a process for constructing a target image data set in one embodiment;

[0075] Figure 4 A schematic diagram of a process for determining mask data in one embodiment;

[0076] Figure 5 A schematic diagram of a process for determining a target image data set in one embodiment;

[0077] Figure 6 is a flowchart of a data annotation method in an exemplary embodiment;

[0078] Figure 7 A schematic diagram of a flow chart of a data labeling method in another embodiment;

[0079] Figure 8 A schematic diagram of a process for constructing an image data set in one embodiment;

[0080] Fig. 9 A schematic diagram of a process for determining an image recognition result in one embodiment;

[0081] Fig.10 is a structural block diagram of a data labeling device in one embodiment;

[0082] Fig.11 is a structural block diagram of a data labeling device in another embodiment;

[0083] Fig.12 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0085] The data annotation method provided in the embodiment of the present application can be applied to Figure 1 In the security monitoring system 100 shown. The security monitoring system 100 includes a labeling server 110, a monitoring server 120 and image acquisition devices 130. The labeling server 110 is used to label the image data set to obtain the target image data set. The monitoring server 120 is used to determine whether the monitored computer room is safe. The image acquisition device 130 is used to collect the initial computer room image data and computer room video data of each computer room. Each image acquisition device 130 is connected to the monitoring server 120 via a transmission line. The labeling server 110 and the monitoring server 120 are connected via a network. The labeling server 110 and the monitoring server 120 can be independent physical servers, or they can be a server cluster or distributed system composed of multiple physical servers, or they can be a cloud server that provides cloud computing services.

[0086] In an exemplary embodiment, Figure 2 As shown, a data annotation method is provided, which is applied to Figure 1 The marking server 110 (hereinafter, the label is omitted, referred to as the marking server) in the embodiment is used as an example to illustrate, including the following steps 202 to 206. Among them:

[0087] Step 202: Receive an image data set sent by a monitoring server.

[0088] The image data set includes image data of each computer room to be labeled and target prompt words.

[0089] In implementation, the annotation server receives the image data set sent by the monitoring server through a network connection.

[0090] Specifically, an image recognition model is set on the annotation server. The operation and maintenance personnel copy the image recognition model in the annotation server to the monitoring server, and configure the environment of the image recognition model on the monitoring server so that the image recognition model can run successfully on the monitoring server. Then, the monitoring server obtains the computer room image data and target prompt words of each computer room, and constructs an image data set based on the computer room image data and target prompt words of each computer room. The monitoring server sends the image data set to the annotation server via a network connection. The annotation server receives the image data set via a network connection.

[0091] Step 204 , annotate the target object in the image dataset according to the computer vision model to obtain the target image dataset.

[0092] The image data set includes image data of each computer room and target prompt words, and the target prompt words represent the target object.

[0093] In implementation, the labeling server performs image recognition on each computer room image data in the image data set according to a computer vision model, and labels the target object corresponding to the target prompt word in the computer room image data to obtain a target image data set.

[0094] Specifically, the annotation server performs image recognition on each computer room image data in the image data set according to the computer recognition model, and annotates the target object corresponding to the target prompt word in the computer room image data to obtain the mask data of the target object. Then, the annotation server determines the target computer room image data according to the mask data of the target object and the computer room image data. The annotation server constructs the target image data set according to each target computer room image data.

[0095] Step 206: Send the target image data set to the monitoring server.

[0096] Among them, the target image dataset is used to train the image recognition model on the monitoring server.

[0097] In implementation, the annotation server sends the target image data set to the monitoring server through a network connection. The monitoring server receives the target image data set and trains an image recognition model based on the target image data set to obtain a target image recognition model. Then, the monitoring server monitors each computer room for potential safety hazards based on the target image recognition model.

[0098] In the above data labeling method, the target objects in the image data set are labeled through the computer vision model to obtain the target image data set, which realizes the automatic labeling of the computer room image data of each computer room, avoids manual labeling, and improves the efficiency of the data labeling method.

[0099] In an exemplary embodiment, the image data set includes image data of each computer room and target prompt words, such as Figure 3 As shown, the specific processing process of step 204 includes steps 302 to 304. Among them:

[0100] Step 302: annotate the target object corresponding to the target prompt word in each computer room image data based on the computer vision model to obtain mask data corresponding to the target object.

[0101] In implementation, the labeling server inputs each computer room image data into a computer vision model, labels the target object corresponding to the target prompt word in the computer room image data through the computer vision model, and obtains mask data corresponding to the target object.

[0102] Specifically, the computer vision model includes a first computer vision model and a second computer vision model. The annotation server performs positioning processing on the target object corresponding to the target prompt word according to the first computer vision model for each computer room image data, and obtains the target frame data of the target object. The target frame data represents the position of the computer room image data where the target object is located. Then, the annotation server performs image segmentation processing on the target frame data and the computer room image data based on the second computer vision model, and obtains the mask data corresponding to the computer room image data.

[0103] Step 304: construct a target image data set according to each computer room image data and the mask data corresponding to each computer room image data.

[0104] In implementation, the annotation server adds the mask data corresponding to each computer room image data to the computer room image data to obtain the target computer room image data, and constructs a target image data set according to each target computer room image data.

[0105] In an exemplary embodiment, if there are multiple target prompt words, each computer room image data has mask data corresponding to each target prompt word. The annotation server adds each target prompt word corresponding to each computer room image data and the mask data corresponding to each target prompt word to the computer room image data to obtain the target computer room image data. Then, the annotation server combines each target computer room image data to obtain the target computer room image data set.

[0106] In this embodiment, a computer vision model is used to annotate mask data corresponding to the target prompt word in each computer room image data set, thereby obtaining the object that the image recognition model needs to recognize, avoiding manual participation, realizing automatic labeling, and improving the efficiency of the data labeling method.

[0107] In an exemplary embodiment, the computer vision model includes a first computer vision model and a second computer vision model, such as Figure 4 As shown, the specific processing process of step 302 includes steps 402 to 404. Among them:

[0108] Step 402: locate the target object corresponding to the target prompt word in each computer room image data according to the first computer vision model, and obtain the target frame data corresponding to the computer room image data.

[0109] Among them, the target prompt word represents the target object with potential safety hazards. The target frame data is used to represent the position of the target object in the computer room image data. The first computer vision model is the DINOv2 model (Dual-Stage Implicit Object-Oriented Network, a dual-stage implicit object-oriented network, a computer vision self-supervised model).

[0110] In implementation, the annotation server inputs each computer room image data and the target prompt word into the first computer vision model for each computer room image data in the image data set, locates the target object in the computer room image data through the first computer vision model, and obtains the target frame data corresponding to the computer room image data.

[0111] Specifically, the labeling server inputs the image data and target prompt words of each computer room image data into the DINOv2 model, and locates the target object in the computer room image data through the DINOv2 model to obtain the target frame data in the computer room image data. The DINOv2 model has a powerful feature extraction capability and can effectively predict the location of the target object even in a new scene that has never been seen, thereby preliminarily realizing the automatic labeling of the target object.

[0112] In an exemplary embodiment, the target prompt word is scattered wires or an empty chair. The annotation server inputs the computer room image data, the scattered wires, and the chair without a human being into the DINOv2 model for each computer room image data in the image data set, and annotates the scattered wires and the empty chair in the computer room image data through the DINOv2 model to obtain the target frame data in the computer room image data. The target frame data represents the position of the scattered wires in the computer room image data, or represents the position of the empty chair in the computer room image data.

[0113] In an optional embodiment, the annotation server pre-fine-tunes the DINOv2 model according to ConvLora to obtain a fine-tuned DINOv2 model. Specifically, the annotation server freezes the backbone network of the DINOv2 model (i.e., the backbone part of visual feature extraction). Then, the annotation server replaces the qkv (Query, Key, Value, Q will represent the query vector for matching with the Key, K will represent the key vector for comparison with the Query, and V will represent the value vector weighted according to the attention weight) implemented by the MLP (Multilayer Perceptron) in the attention mechanism with ConvLora (ConvolutionLow-Rank Adaptation) in the Block of the Transformer (a basic unit in the Transformer model) contained in the DINOv2 model. The annotation server obtains multiple real sample image data of the first computer room. The sample image data of the first computer room is the image data of the computer room containing the target box marked with the target object. Then, the annotation server trains the fine-tuned DINOv2 model according to the sample image data of each computer room until the trained DINOv2 model meets the first computer video model training completion condition, and the annotation server determines the trained DINOv2 model as the target DINOv2 model. Then, for each computer room image data in each computer room image data, the annotation server inputs the computer room image data and the target prompt word into the target DINOv2 model, and locates the target object in the computer room image data through the target DINOv2 model to obtain the target box data in the computer room image data. Replacing qkv with ConvLora can speed up the update efficiency.

[0114] Optionally, the target prompt word is determined according to monitoring requirements, and there may be one or more target prompt words. The embodiment of the present application does not limit the target prompt word.

[0115] Step 404 , based on the second computer vision model and the target frame data corresponding to the computer room image data, image segmentation processing is performed on each computer room image data to obtain mask data corresponding to each computer room image data.

[0116] In the implementation, the second computer vision model is a SAM model (Segment Anything Model, a large image segmentation model).

[0117] During implementation, the labeling server inputs the computer room image data and the target frame data corresponding to the computer room image data into the second computer vision model for each computer room image data, and performs image segmentation processing on the target object in the target frame data through the second computer vision model to obtain mask data corresponding to each computer room image data.

[0118] Specifically, the annotation server inputs the computer room image data and the target frame data corresponding to the computer room image data into the SAM model for each computer room image data. The SAM model in the annotation server preliminarily determines the position of the target object based on the target frame data corresponding to the computer room image data, and performs image segmentation processing on the target object in the computer room image data according to the position of the target object to obtain mask data corresponding to the target object. The mask data marks the target boundary of the target object in detail, providing high-quality annotation data for the subsequent training of the image recognition model.

[0119] In an exemplary embodiment, the target prompt words are empty chairs and scattered wires. The annotation server inputs the computer room image data and the target frame data corresponding to the computer room image data into the SAM model for each computer room image data in each computer room image data. The SAM model in the annotation server preliminarily determines the position of the target object based on the target frame data corresponding to the computer room image data, and performs image segmentation processing on the scattered wires and empty chairs in the computer room image data according to the position of the target object, and obtains mask data corresponding to the scattered wires and mask data corresponding to the empty chairs. The mask data marks the target boundary of the target object in detail, providing high-quality annotation data for subsequent segmentation model training.

[0120] In an optional embodiment, the annotation server pre-fine-tunes the SAM model according to ConvLora to obtain a fine-tuned SAM model. Specifically, the annotation server freezes the position prediction module of the decoder and the encoder module of the prompt word in the SAM model. Then, the annotation server modifies the qkv matrix and applies ConvLoRA in the self-attention layer of the ViT encoder (Vision Transformer, a Transformer-based visual encoder) of the SAM model. In addition, the annotation server adds a lightweight multi-layer perceptron (MLPs, a lightweight version of the multi-layer perceptron Multilayer Perceptron) to the Mask decoder of the SAM model, and adds a lightweight multi-layer perceptron (MLPs) to the mask decoder to support multi-class semantic segmentation prediction. The annotation server obtains multiple real sample image data of the second computer room. The sample image data of the second computer room is the image data of the computer room containing the mask data of the target object annotated. Then, the annotation server trains the fine-tuned SAM model according to each second computer room sample image data until the trained SAM model meets the second computer video model training completion condition, and the annotation server determines the trained SAM model as the target SAM model. For each computer room image data in each computer room image data, the annotation server inputs the computer room image data and the target frame data corresponding to the computer room image data into the target SAM model. The target SAM model in the annotation server preliminarily determines the position of the target object based on the target frame data corresponding to the computer room image data, and performs image segmentation processing on the target object in the computer room image data according to the position of the target object to obtain the mask data corresponding to the target image. The high-precision mask generated by the SAM model is used as the basis to generate a sample set of images. This process avoids manual annotation in traditional methods, greatly reduces manpower and time costs, and ensures the consistency and accuracy of data annotation.

[0121] For example, the annotation server collects the computer room environment video through the computer room camera of the computer room, and extracts the frame of the computer room environment video data to obtain the computer room sample image data. Then, the annotation server annotates the target boxes and segmentation masks of the power lines, network cables, chairs, display screens, keyboards, people and other targets on the ground in each computer room sample image data according to the target prompt words, and obtains the first computer room sample image data and the second computer room sample image data. Each target prompt word is annotated with 500 pictures, covering the data of about 5 computer rooms. The annotation server provides the first computer room sample image data and the second computer room sample image data to DINOv2 and SAM for fine-tuning, and uses the first computer room sample image data and the second computer room sample image data to train a Yolov8-Seg model as the basic model, so that when a new computer room environment is deployed later, transfer learning is used to use the weights trained by the previous model to increase the model convergence speed.

[0122] In this embodiment, the image data of each computer room are processed through the first computer vision model and the second computer vision model to obtain each mask data, clarify the position of the target object in each picture data, and realize automatic labeling of the image data of each computer room, avoiding manual participation, improving data security, and at the same time improving the labeling efficiency, thereby improving the efficiency of the data labeling method.

[0123] In an exemplary embodiment, the computer vision model includes a first computer vision model and a second computer vision model, such as Figure 5 As shown, the specific processing process of step 304 includes steps 502 to 504. Among them:

[0124] Step 502: Add the mask data corresponding to each computer room image data to the computer room image data to obtain the target computer room image data.

[0125] In implementation, the labeling server adds the mask data of the computer room image data to the computer room image data for each computer room image data to obtain the target computer room image data.

[0126] Specifically, the mask data includes the location data of the target object. The annotation server locates the location of the mask data based on the location data of the target data for each computer room image data, and adds the mask data to the computer room image data according to the location of the mask data to obtain the target computer room image data.

[0127] In an exemplary embodiment, the mask data of the target object is the mask data of the empty chair and the mask data of the scattered wires. Each computer room image data of the labeling server is added, and the mask data of the empty chair is added to the computer room image data based on the position data of the empty chair, and the mask data of the scattered wires is added to the computer room image data to obtain the target computer room image data.

[0128] In an optional embodiment, the mask data of the target object includes pixel coordinates of multiple target objects. The annotation server performs standardization processing on each pixel coordinate in the mask data according to a preset format for each mask data to obtain standardized mask data. The annotation server determines the standardized mask data corresponding to the computer room image data for each computer room image data. Then, the annotation server adds the standardized mask data corresponding to the computer room image data to the computer room image data to obtain the target computer room image data.

[0129] Step 504: combine the image data of each target computer room to obtain a target image data set.

[0130] In implementation, the annotation server combines the image data of each target computer room together to obtain an initial target image data set. Then, the annotation server adds the target prompt word to the initial target image data set to obtain a target image data set.

[0131] In an exemplary embodiment, Figure 6 FIG. 1 is a flow chart of a data annotation method in an exemplary embodiment. Figure 6 As shown, the data annotation method includes:

[0132] Step 601 , collecting each initial computer room image data, and determining each computer room image data in each initial computer room image data.

[0133] Step 602, determining the target prompt word.

[0134] Step 603: Input the image data of each computer room and the target prompt words into the DNIO v2 large model.

[0135] Step 604: positioning the target object corresponding to the target prompt word in each computer room image data by using the DNIO v2 large model to obtain the target frame data corresponding to the target object.

[0136] Step 605, input the target frame data and the computer room image data into the SAM large model.

[0137] Step 606: perform image segmentation processing on the target frame data and the computer room image data based on the SAM model to obtain mask data corresponding to the computer room image data.

[0138] Step 607: construct target computer room image data based on the mask data and the computer room image data, and construct a target image data set based on each target computer room image data.

[0139] Step 608: train an image recognition model according to the target image data set to obtain a target image recognition model.

[0140] In this embodiment, by adding mask data to the computer room image data, target computer room image data for training the image recognition model is obtained, and a target image data set is constructed based on each target computer room image data, thereby achieving automatic labeling of the image data set and improving the labeling accuracy and efficiency.

[0141] In an exemplary embodiment, Figure 7 As shown, a data annotation method is provided, which is applied to Figure 1 The monitoring server 120 (hereinafter, the reference numeral is omitted, referred to as the monitoring server) in FIG. 1 is used as an example to illustrate the method, which includes the following steps 702 to 706. Among them:

[0142] Step 702: Acquire initial computer room image data and initial sample recognition template of each computer room.

[0143] In implementation, the image acquisition device of each computer room acquires the initial computer room image data of each computer room. Then, each image acquisition device transmits the acquired initial computer room image data to the monitoring server. The monitoring server receives the initial computer room image data of each computer room. The initial computer room image data is picture data. Then, the monitoring server obtains the initial sample recognition template.

[0144] Specifically, an enterprise will place multiple servers in a computer room to support the operation of various computer devices in the enterprise. In order to ensure the safety of the enterprise's computer room, it is necessary to install an image acquisition device in each computer room. The image acquisition device can collect the initial computer room image data and video data of the computer room. The computer room image data characterizes the environmental characteristics of the computer room and the location characteristics of the items in the computer room. The image acquisition device in each computer room sends the collected initial computer room image data to the monitoring server. The monitoring server receives the initial computer room image data of each computer room and obtains the initial sample recognition template.

[0145] In an exemplary embodiment, the operation and maintenance personnel deploy at least one image acquisition device in each computer room, and the image acquisition device can collect the initial computer room image data of every corner of the computer room without blind spots. Then, the operation and maintenance personnel copy the image recognition model on the annotation server to the monitoring server, and adjust the environment of the monitoring server so that the image recognition model can run smoothly on the monitoring server. Each image acquisition device transmits the collected initial computer room image data to the monitoring server by wired transmission. The monitoring server receives the initial computer room image data of each computer room and stores the initial computer room image data of each computer room locally. The operation and maintenance personnel initiate a model training request for the image recognition model to the monitoring server. The monitoring server responds to the model training request and obtains an initial sample recognition model. The initial sample recognition is used to determine the target prompt word.

[0146] Step 704: construct an image dataset based on the initial sample recognition template and the initial computer room image data of each computer room, and send the image dataset to the annotation server so that the annotation server annotates the image dataset through a computer vision model to obtain a target image dataset.

[0147] In implementation, the monitoring server determines the computer room image data from each of the initial computer room image data of each computer room, and determines the target prompt word based on the initial sample recognition template. Then, the monitoring server constructs an image data set based on the computer room image data and the target prompt word of each computer room, and sends the image data set to the annotation server. The annotation server receives the image data set, and annotates the image data set based on the computer vision model to obtain the target image data set. Among them, the process of annotating the image data set by the annotation server has been described in detail in the above embodiment, and the embodiment of the present application will not be repeated here.

[0148] Specifically, the initial sample recognition template includes each default prompt word. The monitoring server determines each computer room image data in each initial computer room image data for each computer room. The monitoring server determines the target prompt word based on each default prompt word. Then, the monitoring server combines the computer room image data and the target prompt word of each computer room to obtain an image data set. The monitoring server sends the image data set to the annotation server via a network connection.

[0149] Step 706: Receive the target image data set sent by the annotation server, and train the image recognition model according to the target image data set to obtain the target image recognition model.

[0150] Among them, the image recognition model is the YOLOv8 model (You Only Look Once version 8, a deep learning model for target detection).

[0151] In implementation, a training stop condition is pre-set in the monitoring server. The monitoring server receives the target image data set sent by the annotation server through a network connection. The monitoring server trains the image recognition model based on the target image data set until the image recognition model meets the preset training stop condition, and the monitoring server determines the image recognition model as the target image recognition model.

[0152] Specifically, the monitoring server divides the target image data set into a training target image data set and a verification target image data set according to a preset ratio. The monitoring server trains an image recognition model based on the training target image data set to obtain a trained image recognition model. Then, the monitoring server inputs the verification target image data set into the trained image recognition model, and performs image segmentation processing on the verification target image data set through the image recognition model to obtain a verification result. Then, the monitoring server determines whether the verification result meets the preset training completion condition. If the verification result meets the training completion condition, the monitoring server determines the trained image recognition model as the target image recognition model. If the verification result does not meet the training completion condition, the monitoring server executes the step of inputting the verification target image data set into the trained image recognition model until the verification result meets the training completion condition.

[0153] In an exemplary embodiment, the ratio is training target image data set: verification target image data set = 8:2. The monitoring server divides the target image data set into a training target image data set and a verification target image data set at a ratio of 8:2. The monitoring server trains the YOLOv8 model based on the training target image data set to obtain a trained YOLOv8 model. Then, the monitoring server inputs the verification target image data set into the trained YOLOv8 model, and performs image segmentation processing on the verification target image data set through the YOLOv8 model to obtain a verification result. Then, the monitoring server determines whether the verification result meets the preset training completion condition. If the verification result meets the training completion condition, the monitoring server determines the trained YOLOv8 model as the target YOLOv8 model. If the verification result does not meet the training completion condition, the monitoring server executes the step of inputting the verification target image data set into the trained YOLOv8 model until the verification result meets the training completion condition.

[0154] In an exemplary embodiment, the above-mentioned method for training an image recognition model is an incremental learning method. It is to fine-tune a basic image recognition model on a newly collected, automatically annotated data set. Incremental learning allows the model to learn features specific to new scenes while retaining the original knowledge, thereby enhancing adaptability and improving segmentation performance in new scenes. Through model distillation technology, the image recognition model trained with a large amount of local data and the data in the actual user scenario are used to guide the training of the image recognition model, and finally incremental learning is achieved. The incremental learning method is particularly suitable for devices where the initial model performance is poor or needs to adapt to the new environment quickly.

[0155] Optionally, the ratio can be but is not limited to being set to training target image dataset: verification target image dataset = 8:2, or training target image dataset: verification target image dataset = 7:3. The embodiment of the present application does not limit the ratio.

[0156] Optionally, the training stop condition may be, but is not limited to, the number of training rounds reaching a preset training round threshold or the loss value reaching a preset loss threshold. The embodiment of the present application does not limit the training stop condition.

[0157] In this embodiment, an image dataset is constructed and sent to a labeling server so that the labeling server automatically labels the image dataset through a computer vision model, thereby improving the efficiency of labeling the image dataset. The image recognition model is then trained based on the labeled target image dataset, thereby improving the training efficiency.

[0158] In an exemplary embodiment, Figure 8As shown, the specific processing process of constructing the image data set based on the initial sample recognition template and the initial computer room image data of each computer room in step 704 includes steps 802 to 806. Among them:

[0159] Step 802, determining each computer room image data in each initial computer room image data of each computer room.

[0160] In implementation, the monitoring server is provided with the number of images, and the monitoring server determines each computer room image data in each initial computer room image data according to the number of images for each initial computer room image data of each computer room.

[0161] In an optional embodiment, the monitoring server is connected to the display screen via a connecting line. The monitoring server creates a computer room image data page based on the initial computer room image data of each computer room. Then, the monitoring server displays the computer room image data page via the display screen. The operation and maintenance personnel view each initial computer room image data of each computer room via the computer room image data page. Then, the operation and maintenance personnel trigger the initial computer room image data in each initial computer room image data. In response to the triggering operation of the initial computer room image data, the monitoring server determines the triggered initial computer room image data as the computer room image data.

[0162] In particular, each computer room requires at least one computer room image data, and the computer room image data needs to fully display the environment of the computer room.

[0163] Step 804: determine the target prompt word from the default prompt words in the initial sample recognition template.

[0164] The initial sample recognition model includes various default prompt words.

[0165] In implementation, the monitoring server determines the target prompt word based on each default prompt word.

[0166] Specifically, the monitoring server displays each default prompt word in the initial sample recognition template through the computer room image data page. The operation and maintenance personnel view each default prompt word and trigger the default prompt word. In response to the triggering operation of the default prompt word, the monitoring server determines the triggered default prompt word as the target prompt word.

[0167] In an optional embodiment, if the target prompt word does not exist in each default prompt word, the operation and maintenance personnel will input the target prompt word to the monitoring server. The monitoring server obtains the target prompt word input by the operation and maintenance personnel in response to the editing operation of the target prompt word.

[0168] In an exemplary embodiment, each default prompt word is scattered wires, empty chairs, and servers. The monitoring server displays scattered wires, empty chairs, and servers through a computer room image data page. The operation and maintenance personnel view each default prompt word and trigger scattered wires and empty chairs. In response to the triggering operation of scattered wires and empty chairs, the monitoring server determines scattered wires and empty chairs as target prompt words.

[0169] Optionally, the target prompt words may be, but are not limited to, scattered wires and lines, cartons, chairs, open cabinet doors, people, display screens, keyboards, etc. The embodiment of the present application does not limit the target prompt words.

[0170] Step 806: construct an image data set based on the image data of each computer room and the target prompt words.

[0171] In implementation, the monitoring server combines the image data of each computer room and the target prompt words to obtain a target data set.

[0172] In this embodiment, by determining each computer room image data in each initial computer room image data and determining the target prompt word according to the initial sample recognition template, personalized determination of the target object is achieved, and the application scope of the data annotation method is expanded.

[0173] In an exemplary embodiment, after obtaining the target image recognition model, it is necessary to monitor whether there are safety hazards in the computer room based on the target image recognition model. Fig. 9 As shown, the specific processing process of the data labeling method also includes steps 902 to 906. Among them:

[0174] Step 902, acquiring the computer room video data of each computer room in real time, and performing frame extraction processing on the computer room video data of each computer room to obtain each frame of computer room image data to be monitored.

[0175] In the implementation, the image acquisition device in each computer room acquires the computer room video data of each computer room. Then, each image acquisition device transmits the data stream of the computer room video data to the monitoring server in real time. The monitoring server receives the computer room video data of each computer room. Then, the monitoring server extracts the computer room video data according to the preset frame number for the computer room video data of each computer room, and obtains the computer room image data set of each frame to be detected in the computer room.

[0176] Step 904 , combining the image data of the to-be-monitored computer rooms of the same frame to obtain image data sets of the to-be-monitored computer rooms.

[0177] In implementation, the monitoring server combines the image data to be monitored of each computer room of the same frame together to obtain the image data set of each computer room to be monitored.

[0178] Specifically, the monitoring server combines the image data of the computer rooms to be monitored of each computer room in each frame in a time-ordered order to obtain the image data set of the computer rooms to be monitored of the frame.

[0179] Step 906 , performing image recognition on each to-be-monitored computer room image data set based on the target image recognition model to obtain an image recognition result, and determining the safety status of each computer room based on the image recognition result.

[0180] The image recognition result indicates whether there is a target object corresponding to the target prompt word in the image data of the computer room to be monitored.

[0181] In the implementation, the monitoring server inputs the image data set of the monitored computer room into the target image recognition model for each image data set of the monitored computer room, and performs image recognition processing on each image data of the monitored computer room in the image data set of the monitored computer room through the target image recognition model to obtain the image recognition result. Then, the monitoring server determines the security status of each computer room according to the image recognition result.

[0182] Specifically, the target image recognition model is a target YOLOv8 model (You Only Look Once version 8, a deep learning model for target detection). The monitoring server inputs the image data set of the computer room to be monitored into the target YOLOv8 model for each image data set of the computer room to be monitored, and performs image recognition processing on each image data of the computer room to be monitored in the image data set of the computer room to be monitored through the target YOLOv8 model to obtain an image recognition result. Then, the monitoring server determines whether the image recognition result represents the existence of a target object corresponding to the target prompt word in the image data of the computer room to be monitored. If the image recognition result represents the existence of a target object corresponding to the target prompt word in the image data of the computer room to be monitored, the monitoring server determines the target alarm condition corresponding to the query target prompt word. Then, the monitoring server determines whether the image data of the computer room to be monitored meets the target alarm condition based on each image recognition result corresponding to the computer room. When the image data of the computer room to be monitored meets the target alarm condition, an alarm message is sent to the terminal of the operation and maintenance personnel.

[0183] In an exemplary embodiment, the target prompt word is an empty chair. If the image recognition result indicates that there is an empty chair in the frankincense data of the computer room to be monitored, the monitoring server queries the target alarm condition corresponding to the empty chair. The target alarm condition is that the empty chair exists for more than 10 minutes. The monitoring server determines the image recognition results of the computer room in the previous 10 minutes based on the current time as the starting point. If the image recognition results of the previous 10 minutes all indicate that there are empty chairs, the monitoring server determines that the image data of the computer room to be monitored meets the target alarm condition. When the image data of the computer room to be monitored meets the target alarm condition, an alarm message is sent to the terminal of the operation and maintenance personnel. The alarm message includes an image of the computer room.

[0184] Optionally, the alarm information may be sent to the terminal of the operation and maintenance personnel via, but not limited to, email or text message. The implementation of this application does not limit the method of sending the alarm information.

[0185] In this embodiment, image recognition is performed on each image data set of the computer room to be monitored through the target recognition model, and the safety status of the computer room is determined according to the image recognition result, thereby realizing automatic monitoring of the safety of the computer room and improving the monitoring efficiency.

[0186] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0187] Based on the same inventive concept, the embodiment of the present application also provides a data labeling device for implementing the data labeling method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more data labeling device embodiments provided below can refer to the limitations on the data labeling method above, and will not be repeated here.

[0188] In an exemplary embodiment, Fig.10 As shown, a data annotation device 1000 is provided, comprising: a receiving module 1001, an annotation module 1002 and a sending module 1003, wherein:

[0189] The receiving module 1001 is used to receive an image data set sent by a monitoring server.

[0190] The labeling module 1002 is used to label the target object in the image data set according to the computer vision model to obtain the target image data set.

[0191] The sending module 1003 is used to send the target image data set to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

[0192] In an exemplary embodiment, the image data set includes image data of each computer room and target prompt words, and the annotation module 1002 includes:

[0193] The first labeling submodule is used to label the target object corresponding to the target prompt word in each computer room image data based on the computer vision model to obtain mask data corresponding to the target object.

[0194] The first construction submodule is used to construct a target image data set according to each computer room image data and mask data corresponding to each computer room image data.

[0195] In an exemplary embodiment, the computer vision model includes a first computer vision model and a second computer vision model, and the first annotation submodule includes:

[0196] The first positioning submodule is used to locate the target object corresponding to the target prompt word in each computer room image data according to the first computer vision model, and obtain the target frame data corresponding to the computer room image data.

[0197] The first segmentation submodule is used to perform image segmentation processing on each computer room image data based on the second computer vision model and the target frame data corresponding to the computer room image data to obtain mask data corresponding to each computer room image data.

[0198] In an exemplary embodiment, the first building block includes:

[0199] The first adding submodule is used to add the mask data corresponding to each computer room image data to the computer room image data to obtain the target computer room image data.

[0200] The first combining submodule is used to combine the image data of each target computer room to obtain a target image data set.

[0201] In an exemplary embodiment, Fig.11 As shown, a data annotation device 1100 is provided, including: an acquisition module 1101, a construction module 1102 and a sending module 1103, wherein:

[0202] The acquisition module 1101 is used to acquire the initial computer room image data and initial sample recognition template of each computer room.

[0203] Construction module 1102 is used to construct an image dataset based on the initial sample recognition template and the initial computer room image data of each computer room, and send the image dataset to the annotation server so that the annotation server annotates the computer room image dataset through a computer vision model to obtain a target image dataset.

[0204] The training module 1103 is used to receive the target image data set sent by the annotation server, and train the image recognition model according to the target image data set to obtain the target image recognition model.

[0205] In an exemplary embodiment, the building module includes a second building submodule and a first sending submodule. The second building submodule includes:

[0206] The first determining submodule is used to determine each computer room image data in each initial computer room image data of each computer room.

[0207] The second determination submodule is used to determine the target prompt word from each default prompt word in the initial sample recognition template.

[0208] The third construction submodule is used to construct an image data set according to the image data of each computer room and the target prompt words.

[0209] In an exemplary embodiment, the data annotation device 1100 further includes:

[0210] The processing module is used to obtain the computer room video data of each computer room in real time, and perform frame extraction processing on the computer room video data of each computer room to obtain each frame of the computer room image data to be monitored.

[0211] The combination module is used to combine the image data of the to-be-monitored computer rooms of the same frame to obtain the image data sets of the to-be-monitored computer rooms.

[0212] The recognition module is used to perform image recognition on each image data set of the computer room to be monitored based on the target image recognition model, obtain image recognition results, and determine the safety status of each computer room based on the image recognition results.

[0213] Each module in the above data annotation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.

[0214] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.12 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data used in the data annotation method. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data annotation method is implemented.

[0215] Those skilled in the art will understand that Fig.12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0216] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0217] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0218] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0219] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0220] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0221] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A data labeling method, characterized in that: The method is applied to a labeling server, and the method comprises: Receiving an image data set sent by a monitoring server; Annotate the target object in the image dataset according to the computer vision model to obtain a target image dataset; The target image data set is sent to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

2. The method according to claim 1, characterized in that The image data set includes image data of each computer room and target prompt words, and the target objects in the image data set are labeled according to the computer vision model to obtain the target image data set, including: Based on a computer vision model, a target object corresponding to the target prompt word in each of the computer room image data is labeled to obtain mask data corresponding to the target object; A target image data set is constructed according to each of the computer room image data and the mask data corresponding to each of the computer room image data.

3. The method according to claim 2, characterized in that The computer vision model includes a first computer vision model and a second computer vision model, and the target object corresponding to the target prompt word in each of the computer room image data is labeled based on the computer vision model to obtain mask data corresponding to the target object, including: Locating the target object corresponding to the target prompt word in each of the computer room image data according to the first computer vision model, and obtaining the target frame data corresponding to the computer room image data; Based on the second computer vision model and the target frame data corresponding to the computer room image data, image segmentation processing is performed on each of the computer room image data to obtain mask data corresponding to each of the computer room image data.

4. The method according to claim 2, characterized in that: The constructing a target image data set according to each of the computer room image data and the mask data corresponding to each of the computer room image data comprises: Adding the mask data corresponding to each of the computer room image data to the computer room image data to obtain the target computer room image data; The target computer room image data are combined to obtain a target image data set.

5. A data annotation method, characterized in that: The method is applied to a monitoring server, and the method comprises: Obtaining initial computer room image data and initial sample recognition templates of each computer room; Building an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and sending the image data set to a labeling server, so that the labeling server labels the image data set through a computer vision model to obtain a target image data set; The target image data set sent by the annotation server is received, and an image recognition model is trained according to the target image data set to obtain a target image recognition model.

6. The method according to claim 5, characterized in that The constructing of the image data set based on the initial sample recognition template and the initial computer room image data of each computer room comprises: Determining each computer room image data in each initial computer room image data of each of the computer rooms; Determining a target prompt word from each default prompt word in the initial sample recognition template; An image data set is constructed according to each of the computer room image data and the target prompt words.

7. The method according to claim 5, characterized in that After training the image recognition model according to the target image data set to obtain the target image recognition model, the method includes: Acquire the computer room video data of each of the computer rooms in real time, and perform frame extraction processing on the computer room video data of each of the computer rooms to obtain each frame of computer room image data to be monitored; Combining the image data of the computer rooms to be monitored of the same frame of each of the computer rooms to obtain image data sets of each computer room to be monitored; Image recognition is performed on each of the to-be-monitored computer room image data sets based on the target image recognition model to obtain an image recognition result, and the safety status of each of the computer rooms is determined based on the image recognition result.

8. A data labeling device, characterized in that: The device comprises: A receiving module, used for receiving an image data set sent by a monitoring server; A labeling module, used for labeling the target object in the image data set according to a computer vision model to obtain a target image data set; The sending module is used to send the target image data set to the monitoring server; the target image data set is used to train the image recognition model on the monitoring server.

9. A data labeling device, characterized in that: The device comprises: An acquisition module, used to acquire initial computer room image data and initial sample recognition templates of each computer room; A construction module, used to construct an image data set based on the initial sample recognition template and the initial computer room image data of each computer room, and send the image data set to a labeling server, so that the labeling server labels the computer room image data set through a computer vision model to obtain a target image data set; The training module is used to receive the target image data set sent by the annotation server, and train the image recognition model according to the target image data set to obtain the target image recognition model.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 or 5 to 7 are implemented.