Image annotation method and device under complex background, storage medium and computer program product
By using eye tracking equipment and lightweight twin neural networks, we can quickly locate the target gaze point and determine the position of the annotation box, and solve the problem of slow labeling speed and low accuracy of the target sample in complex backgrounds, and achieve efficient and accurate image annotation.
Patent Information
- Application Number
- CN202411357000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-06-13
AI Technical Summary
In complex backgrounds, the labeling speed of image data target samples is slow, the workload is large, the automatic labeling accuracy is low, and the adaptability of complex backgrounds is poor.
By obtaining the image to be marked, receiving eye movement data of the annotated personnel of the eye movement tracking device, determining the target gaze point position, using a selective search algorithm to determine the target candidate area, and using a lightweight twin neural network to perform similarity calculations, and automatically determining the image classification attribute labels in the target gaze box.
It realizes fast and accurate labeling of image data target samples under complex backgrounds, reduces personnel workload, and improves the accuracy and adaptability of automatic labeling.
Smart Images

Figure CN120148032A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image annotation, and particularly to an image annotation method, device, storage medium and computer program product under complex backgrounds. Background Art
[0002] With the wide application of intelligent target recognition technologies in scenarios such as image remote sensing and airborne optoelectronic detection, the deep neural network models adopted often need to be trained using large-scale data samples, and have relatively high requirements for the annotation speed and quality of image data target samples under complex backgrounds such as the ground, sea surface, and cloudy skies.
[0003] However, for the manual annotation of image data under the current complex background, first, the annotator needs to find the location of the target, then draw an annotation box, and finally assign the target type to the sample through human eye recognition. This method is slow and has a large workload. Moreover, for existing automatic annotation, such as the data annotation method described in Patent CN202110340495.2, data annotation is carried out with the aid of a data annotation device, and the annotation of the data to be annotated is completed by listening for and accepting corresponding annotation instructions. Inevitably, complex human preset operation events such as "click operations, voice instructions or touch instructions" are required during the data annotation process in this patent, and the formed data annotation box is "generated according to the determined data annotation start point and data annotation end point". There are certain errors in the fineness of the annotated data, and overall, the workload of the annotator is not effectively reduced, and the annotation efficiency is not significantly improved.
[0004] Aiming at the above problems, how to achieve efficient annotation of image data target samples under complex backgrounds, quickly complete the determination of the target position center, the selection of the target annotation box range, the determination of the target classification attributes, and the verification of the target sample annotation, and improve the speed and accuracy of the target sample annotation, is an urgent problem to be solved currently. Summary of the Invention
[0005] Embodiments of the present application provide an image annotation method, device, storage medium and computer program product under complex backgrounds to solve problems such as slow manual annotation speed, large workload, low accuracy of automatic annotation, and poor adaptability to complex backgrounds in the annotation of image data target samples under complex backgrounds.
[0006] To achieve the above object, the present application adopts the following technical solutions:
[0007] In a first aspect, the present application provides an image annotation method under a complex background. The method includes: obtaining an image to be annotated; receiving eye movement data of an annotator sent by an eye movement tracking device to determine the position of the target fixation point of the annotator; according to the selective search algorithm, taking the position of the target fixation point as the starting position, determining the area with similar features to the position of the target fixation point as the target candidate area; calculating the coordinate position of the target candidate area on the image to be annotated, and using a target annotation box to annotate the target candidate area; according to a lightweight siamese neural network, calculating the similarity between the image within the target annotation box and each image in the reference sample image set to obtain the maximum similarity threshold; and determining and annotating the classification attribute label of the image within the target annotation box according to the maximum similarity threshold.
[0008] A possible design solution is that the method in the first aspect further includes that the reference sample image set is an image set pre-annotated with classification attribute labels by the annotator, and the reference sample image set includes images of at least one category, and each category of images includes images of at least one state.
[0009] A possible design solution is that the method in the first aspect further includes that the eye movement data includes the coordinates and timestamps of the fixation points. Receiving the eye movement data of the annotator sent by the eye movement tracking device to determine the position of the target fixation point of the annotator includes: receiving the data of the fixation point coordinates of the annotator during the period of observing the image to be annotated sent by the eye movement tracking device; and determining the geometric center of the fixation points during the period by aggregating the data of the fixation point coordinates during the period, and the geometric center is the position of the target fixation point of the annotator.
[0010] A possible design solution is that the method in the first aspect further includes that calculating the coordinate position of the target candidate area on the image to be annotated and using a target annotation box to annotate the target candidate area includes: obtaining the upper, lower, left, and right extreme coordinates of the target candidate area, and using a target annotation box to annotate the target candidate area on the image to be annotated according to the extreme coordinates.
[0011] A possible design solution is that the method in the first aspect further includes that according to the lightweight siamese neural network, calculating the similarity between the image within the target annotation box and each image in the reference sample image set to obtain the maximum similarity threshold includes: according to the lightweight siamese neural network, extracting the feature vectors of the image within the target annotation box and the feature vectors of each image in the reference sample image set, and calculating the similarity between the image within the target annotation box and each image in the reference sample image set through the feature vectors; and comparing the similarity values between the image within the target annotation box and each image in the reference sample image set to obtain the maximum similarity threshold.
[0012] In a possible design, the method in the first aspect further includes determining and labeling the classification attribute label of the image within the target annotation box according to the maximum similarity threshold, including: presetting a similarity threshold, and determining whether the maximum similarity threshold is greater than the preset similarity threshold; if the maximum similarity threshold is greater than the preset similarity threshold, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image within the target annotation box, and automatic labeling is performed; if the maximum similarity threshold is less than the preset similarity threshold, the image within the target annotation box is manually verified to determine whether the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold correctly represents the image within the target annotation box. If so, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image within the target annotation box, and automatic labeling is performed. If not, the correct classification attribute label is manually labeled.
[0013] In a second aspect, an image annotation device under complex background is provided. The image annotation device under complex background includes a module for executing the method in the first aspect above.
[0014] In a possible design, the image annotation device under complex background in the second aspect may further include a transceiver. The transceiver may be a transceiver circuit or an interface circuit. The transceiver may be used for the image annotation device under complex background in the second aspect to communicate with other devices.
[0015] In a possible design, the image annotation device under complex background in the second aspect may further include a memory. The memory may be integrated with the processor or may be separately provided. The memory may be used to store the instructions involved in the method in the first aspect.
[0016] In a third aspect, an image annotation device under complex background is provided. The image annotation device under complex background includes: a processor, the processor is coupled to the memory, and the processor is used to execute the instructions stored in the memory so that the image annotation device under complex background executes the method described in the first aspect.
[0017] In a possible design, the image annotation device under complex background described in the third aspect may further include a transceiver. The transceiver may be a transceiver circuit or an interface circuit. The transceiver may be used for the image annotation device under complex background described in the third aspect to communicate with other devices.
[0018] In a fourth aspect, an image annotation device under complex background is provided, including: a processor and a memory; the memory is used to store instructions, and when the processor executes the instructions, the image annotation device under complex background is enabled to execute the method described in the first aspect.
[0019] In a possible design, the image annotation device in the fourth aspect may further include a transceiver. The transceiver may be a transceiver circuit or an interface circuit. The transceiver may be used for the image annotation device in the fourth aspect to communicate with other devices.
[0020] Fifth aspect, a computer-readable storage medium is provided, which includes a stored computer program or instruction. When the computer program or instruction is run, the image annotation method in the first aspect is executed.
[0021] Sixth aspect, a computer program product is provided, including a computer program or instruction. When the computer program or instruction is run, the image annotation method in the first aspect is executed.
[0022] In the embodiments of the present application, the eye movement data obtained by the eye movement tracking device is used to quickly locate the position of the target fixation point, realizing the extraction of the target with high accuracy in a complex background; the selective search centered on the position of the target fixation point can quickly determine the position of the annotation box in a complex background; using a lightweight siamese neural network for similarity comparison can improve the accuracy of automatic annotation and has a low computational complexity; finally, through the similarity threshold verification and subsequent manual intervention, the reliability of the final annotation result can be ensured. The workload of personnel in the entire annotation process is small, and the rapid and accurate annotation of the target samples in the image data under a complex background is realized.
[0023] Other features and advantages of the present application will be described in detail in the subsequent specific implementation part. Description of the Drawings
[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 Flow chart of the image annotation method in the complex background provided by the embodiments of the present application Figure 1 ;
[0026] Figure 2 Flow chart of the image annotation method in the complex background provided by the embodiments of the present application Figure 2 ;
[0027] Figure 3 Annotation illustration of the target sample in the jungle background provided by the embodiments of the present application Figure 1 ;
[0028] Figure 4 Schematic diagram of annotating target samples in a jungle background provided by an embodiment of the present application Figure 2 ;
[0029] Figure 5 Schematic diagram of the structure of an image message processing device provided by an embodiment of the present application Figure 1 ;
[0030] Figure 6 Schematic diagram of the structure of an image message processing device provided by an embodiment of the present application Figure 2 。 Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise clearly and specifically defined.
[0032] Figure 1 Flow schematic diagram of an image annotation method in a complex background provided by an embodiment of the present application Figure 1 This image annotation method in a complex background is applicable to image annotation in a complex background such as an image processing model, a data annotation model, an image annotation device, or other possible ways, which is not limited herein. For the convenience of understanding the present application, the following mainly takes an image processing model as an example for introduction.
[0033] The process of this image annotation method in a complex background is as follows:
[0034] Step S101, obtain the image to be annotated.
[0035] Step S102, receive the eye movement data of the annotator sent by the eye movement tracking device, and determine the position of the target fixation point of the annotator.
[0036] Among them, the eye movement data of the annotator sent by the eye movement tracking device includes the coordinates and timestamps of the fixation points. That is to say, the eye movement tracking device will collect the eye movement data of the annotator, record the data of the fixation point coordinates during the period of observing the image to be annotated, and send this data to the image processing model.
[0037] The image processing model aggregates the data of the fixation point coordinates within a time period. For example, it calculates the average position of the fixation point coordinates or uses statistical methods such as Gaussian mixture models to determine the geometric center of the fixation points within the time period, and then determines the geometric center as the target fixation point position of the annotator.
[0038] In addition, the above-mentioned eye movement tracking device can be an eye tracker or a binocular camera, or other devices, which are not restricted here.
[0039] Step S103: According to the selective search algorithm, starting from the target fixation point position, determine the area with similar characteristics to the target fixation point position as the target candidate area.
[0040] In step S102, the target fixation point, that is, the target position center, has been determined. In this step, the selective search algorithm is performed on the image starting from the target fixation point position, and the boundary is gradually expanded according to the characteristics of the image, and more areas similar to the image characteristics are covered by increasing the boundary of the area as the target candidate area.
[0041] Step S104: Calculate the coordinate position of the target candidate area on the image to be annotated, and use the target annotation box to annotate the target candidate area.
[0042] It can be understood that for the target candidate area, its upper, lower, left, and right extreme coordinates can be obtained. For example: left boundary: X min , right boundary: X max , upper boundary: Y min , lower boundary: Y max . On the image, according to the above extreme coordinates, use the target annotation box to annotate the target candidate area. The upper left coordinate of the annotation box is (X min , Y min ), and the lower right coordinate is (X max , Y max ), which can ensure that the annotation box correctly covers the target candidate area.
[0043] Step S105: According to the lightweight Siamese neural network, calculate the similarity between the image within the target annotation box and each image in the reference sample image set to obtain the maximum similarity threshold.
[0044] Among them, the lightweight Siamese Network is used to calculate the similarity between images. To make the network more efficient, a small network architecture is usually selected, such as MobileNet, EfficientNet, etc. as the feature extractor. For example, when calculating the similarity between images, it includes: Step A: Feature extraction, using a Convolutional Neural Network (CNN) to extract the features of the image within the annotation box and the reference sample image. Step B: Similarity calculation, using the output features of the network to calculate the similarity between images, and methods such as cosine similarity and Euclidean distance can be used.
[0045] Correspondingly, in this step, the image processing model extracts the feature vectors of the image within the target annotation box and each image in the reference sample image set according to the lightweight Siamese Network, and calculates the similarity between the image within the target annotation box and each image in the reference sample image set through the feature vectors; compares the similarity values between the image within the target annotation box and each image in the reference sample image set to obtain the maximum similarity threshold.
[0046] In addition, it should be noted that the reference sample image set is an image set pre-annotated with classification attribute labels by annotators, and the reference sample image set includes images of at least one category, and each category of images includes images of at least one state. For example: The reference sample image set is set for defect detection, and this reference sample image set may include images of three types: "scratch", "dent", and "color change". The "scratch" type of images correspond to several scratch sample images with different lengths and depths, and the "color change" type of images correspond to several color change sample images with different degrees of color change, etc.
[0047] Step S106, determine the classification attribute label of the image within the target annotation box according to the maximum similarity threshold, and annotate it.
[0048] The image processing model presets a similarity threshold according to actual needs, and determines the matching situation of the image by judging whether the maximum similarity threshold is greater than the preset similarity threshold.
[0049] In a possible implementation, the maximum similarity threshold is greater than the preset similarity threshold, indicating a high degree of image matching. At this time, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image within the target annotation box, and this classification attribute label is automatically stored in the database to complete the annotation.
[0050] In another possible implementation, the maximum similarity threshold is less than the preset similarity threshold, indicating that the matching degree of the image may not be high. It is necessary to manually verify the image within the target annotation box and determine whether the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold correctly represents the image within the target annotation box. If so, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image within the target annotation box, and the classification attribute label is automatically stored in the database to complete the annotation. If not, the correct classification attribute label is manually annotated.
[0051] In summary, in the embodiment of the present application, the eye movement data obtained by the eye movement tracking device is used to quickly locate the position of the target fixation point, realizing high-accuracy extraction of the target under complex backgrounds; the selective search centered on the position of the target fixation point can quickly determine the position of the annotation box under complex backgrounds; the lightweight Siamese neural network is used for similarity comparison, which can improve the accuracy of automatic annotation and has low computational complexity; finally, through similarity threshold verification and subsequent manual intervention, the reliability of the final annotation result can be ensured. The workload of personnel in the entire annotation process is small, and the rapid and accurate annotation of the target samples of image data under complex backgrounds is realized.
[0052] The above Figure 1 has described in detail the image annotation method provided by the embodiment of the present application under complex backgrounds. The following Figures 2 - 4 introduces the specific application scenarios of this image annotation method under complex backgrounds in the optoelectronic detection and recognition scenarios of aerial targets such as unmanned aerial vehicles and flying birds.
[0053] Specifically, as shown in Figure 2, the process of this image annotation method under complex backgrounds is as follows:
[0054] Step S201, before annotation, the annotator manually selects a small number (usually several to dozens) of images to form a reference sample image set, and performs refined manual annotation (positioning, drawing a box, and tagging) on the unmanned aerial vehicles and flying birds in it.
[0055] Among them, the annotated unmanned aerial vehicle and flying bird samples should include typical states under different postures, illuminations, etc. In this way, the reference sample image set includes two categories: unmanned aerial vehicles and flying birds.
[0056] Step S202, obtain the image to be annotated.
[0057] Step S203, receive the eye movement data of the annotator sent by the eye movement tracking device, and determine the target fixation point position of the annotator through data analysis.
[0058] Among them, when the annotator focuses on the unmanned aerial vehicle or flying bird target in the current image, the fixation point often stays near the center of the target. Thus, the pixel coordinates of the center of the target position on the image can be roughly determined. For exampleFigure 3 As shown, the positions marked with dots are the target fixation point positions of the annotator.
[0059] Step S204: According to the selective search algorithm built into the computer program, first find the regions with smaller sizes near the target fixation point positions, and then gradually merge the small-sized regions with similar features into large-sized regions, thereby forming target candidate regions, which generally just cover the target regions of drones or flying birds.
[0060] Such as Figure 4 As shown, the regions marked in a shape similar to a drone or a flying bird are the target candidate regions.
[0061] Step S205: Calculate the coordinate positions of the target candidate regions on the image to be annotated, take the extreme coordinate values of the upper, lower, left, and right, and then assign a bounding box to the target candidate regions, which generally just frame the target candidate regions of drones or flying birds.
[0062] Such as Figure 4 As shown, based on the extreme coordinate values of the upper, lower, left, and right of the target candidate regions, use a bounding box to annotate the target candidate regions.
[0063] Step S206: Based on the lightweight Siamese neural network, calculate the pairwise similarity between the images within the bounding box and each image in the reference sample image set in step S201 to obtain the maximum similarity threshold.
[0064] Step S207: According to the setting of the maximum similarity threshold (such as 0.8), determine the target classification attribute label and annotate it.
[0065] Specifically, for the images within the bounding box with a maximum similarity threshold not lower than 0.8, they can be directly stored in the database automatically to complete the annotation. For the images within the bounding box with a maximum similarity threshold lower than 0.8, there may be errors in the bounding box or the classification attribute label, and it is necessary to manually verify the annotation accuracy to check whether it is a drone or a flying bird. Those that pass the verification (with correct bounding box and classification attribute label) can be stored in the database to complete the annotation, while those that do not pass the verification (with incorrect bounding box or classification attribute label) are annotated manually (positioning, drawing a box, and tagging).
[0066] The above has Figures 1 - 4 described in detail the image annotation method provided in the embodiments of the present application in combination with Figures 5 - 6 The following will describe in detail the image annotation device for implementing the image annotation method provided in the embodiments of the present application in combination with
[0067] Figure 5 is the structural schematic Figure 1 of the image annotation device for complex backgrounds provided in the embodiments of the present application. Exemplarily, such as Figure 5As shown in the figure, the image annotation device 500 under complex background includes: a transceiver module 501 and a processing module 502. For the sake of convenience of description, Figure 5 only the main components of the image annotation device under complex background are shown.
[0068] Among them, the transceiver module 501 is used to execute the transceiver function of the above-mentioned image annotation method under complex background, and the processing module 502 is used to execute other functions of the above-mentioned image annotation method except the transceiver function.
[0069] Optionally, the transceiver module 501 may include a sending module ( Figure 5 not shown in the figure) and a receiving module ( Figure 5 not shown in the figure). Among them, the sending module is used to implement the sending function of the image annotation device 500 under complex background, and the receiving module is used to implement the receiving function of the image annotation device 500 under complex background.
[0070] Optionally, the image annotation device 500 under complex background may further include a storage module ( Figure 5 not shown in the figure), and the storage module stores programs or instructions. When the processing module 502 executes the programs or instructions, the image annotation device 500 under complex background can execute the image annotation method in the embodiment of the present application.
[0071] Next, in combination with Figure 6 each component of the image annotation device 600 under complex background will be specifically introduced:
[0072] Among them, the processor 601 is the control center of the image annotation device 600 under complex background, which can be a single processor or a collective term for multiple processing elements. For example, the processor 601 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0073] Optionally, the processor 601 can execute various functions of the image annotation device 600 under complex background by running or executing software programs stored in the memory 602 and calling data stored in the memory 602, such as executing the image annotation method in the embodiment of the present application.
[0074] In a specific implementation, as an example, the processor 601 may include one or more CPUs, such as Figure 6 the CPU0 and CPU1 shown in
[0075] In a specific implementation, as an example, the image annotation device 600 under a complex background may also include multiple processors, such as Figure 6 the processor 601 and the processor 604 shown in. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Among them, the memory 602 is used to store the software program for executing the solution of this application and is controlled by the processor 601 for execution. The specific implementation method may refer to the above method embodiment and will not be elaborated here.
[0076] Optionally, the memory 602 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 602 may be integrated with the processor 601 or exist independently and be coupled to the processor 601 through the interface circuit of the image annotation device 600 under a complex background( Figure 6 not shown in). This application embodiment does not make specific limitations on this.
[0077] The transceiver 603 is used for communication with other communication devices. For example, when the image annotation device 600 under a complex background is the first device, the transceiver 603 may be used for communication with the second device or communication with the third device.
[0078] Optionally, the transceiver 603 may include a receiver and a transmitter( Figure 6 not shown separately in). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0079] Optionally, the transceiver 603 may be integrated with the processor 601 or exist independently, and is coupled to the processor 601 through an interface circuit ( Figure 6 not shown) of the image annotation device 600 under complex backgrounds. The embodiments of the present application do not make specific limitations thereto.
[0080] It can be understood that Figure 6 the structure of the image annotation device 600 under complex backgrounds shown does not constitute a limitation on the image annotation device under complex backgrounds. The actual image annotation device under complex backgrounds may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0081] In addition, for the technical effects of the image annotation device 600 under complex backgrounds, reference may be made to the technical effects of the method described in the above method embodiments, which will not be elaborated herein.
[0082] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0083] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0084] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more sets of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
Claims
1. A method for labeling an image under a complex background, characterized in that: The method comprises: Get the image to be annotated; Receiving eye movement data of an annotator sent from an eye tracking device, and determining the target gaze point position of the annotator; According to the selective search algorithm, taking the target gaze point position as the starting position, determining the area with similar characteristics to the target gaze point position as the target candidate area; Calculating the coordinate position of the target candidate area on the image to be annotated, and annotating the target candidate area with a target annotation frame; According to the lightweight twin neural network, similarity is calculated for the image in the target annotation box and each image in the reference sample image set to obtain a maximum similarity threshold; According to the maximum similarity threshold, the classification attribute label of the image in the target annotation box is determined and annotated.
2. The image annotation method under complex background according to claim 1, characterized in that: The reference sample image set is an image set that is pre-labeled with classification attribute labels by the labeling personnel, and the reference sample image set includes images of at least one category, and each category of images includes images of at least one state.
3. The image annotation method under complex background according to claim 1, characterized in that: The eye movement data includes the coordinates and timestamp of the gaze point, and the step of receiving the eye movement data of the annotator sent by the eye movement tracking device and determining the target gaze point position of the annotator includes: Receiving data of the gaze point coordinates of the annotator during the time period of observing the image to be annotated, sent from the eye tracking device; The geometric center of the gaze point in the time period is determined by aggregating the data of the gaze point coordinates in the time period, and the geometric center is the target gaze point position of the annotator.
4. The image annotation method under complex background according to claim 1, characterized in that: The calculating the coordinate position of the target candidate area on the image to be annotated and annotating the target candidate area with a target annotation frame includes: The upper, lower, left, and right extreme value coordinates of the target candidate area are obtained, and the target candidate area is annotated with a target annotation frame on the image to be annotated according to the extreme value coordinates.
5. The image annotation method under complex background according to claim 2, characterized in that: The method of calculating the similarity between the image in the target annotation frame and each image in the reference sample image set according to the lightweight twin neural network to obtain a maximum similarity threshold includes: According to the lightweight twin neural network, a feature vector of the image in the target annotation frame and a feature vector of each image in the reference sample image set are extracted, and the similarity between the image in the target annotation frame and each image in the reference sample image set is calculated through the feature vectors; The similarity values of the image in the target annotation box and each image in the reference sample image set are compared to obtain a maximum similarity threshold.
6. The image annotation method under complex background according to claim 5, characterized in that: The step of determining the classification attribute label of the image in the target annotation frame according to the maximum similarity threshold and annotating the label comprises: Preset a similarity threshold, and determine whether the maximum similarity threshold is greater than the preset similarity threshold; If the maximum similarity threshold is greater than the preset similarity threshold, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image in the target annotation frame, and is automatically annotated; If the maximum similarity threshold is less than the preset similarity threshold, the image in the target annotation box is manually checked to determine whether the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold correctly describes the image in the target annotation box. If so, the classification attribute label of the image in the reference sample image set corresponding to the maximum similarity threshold is the classification attribute label of the image in the target annotation box and is automatically labeled. If not, the correct classification attribute label is manually labeled.
7. An image annotation device under complex background, characterized in that: The apparatus comprises: a module for executing the method according to any one of claims 1-6.
8. An image annotation device under complex background, characterized in that: The apparatus for labeling images under complex backgrounds comprises: a processor and a memory; the memory is used to store computer instructions, and when the processor executes the instructions, the apparatus for labeling images under complex backgrounds executes the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a computer program or instruction stored therein, and when the computer program or instruction is executed, the method according to any one of claims 1 to 6 is performed.
10. A computer program product, characterized in that The method comprises a computer program or an instruction, and when the computer program or the instruction is executed, the method according to any one of claims 1 to 6 is performed.
Citation Information
Patent Citations
Data annotation methods, devices, terminal equipment and storage media
CN112967359B
Cited By
Statistical analysis-based label automatic selection method, system and device
CN117292173A