Object detection method and device, model training method and device, equipment and medium
Through the hierarchical autoencoder structure, the rectangular box and category information are converted into fixed-length hidden vectors, and combined with the joint training of image feature extractor and diffusion model, the problems of uncertain size of rectangular box sets, complex element arrangement order, and difficult modeling of coordinates and categories in DiffusionDet technology are solved, achieving more efficient and accurate object detection.
Patent Information
- Application Number
- CN202411938991.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
In object detection, the existing DiffusionDet technology has problems such as uncertain size of rectangular box sets, complex element arrangement order to vector spatial structure, and inability to fully model the joint probability relationship between coordinates and categories.
The hierarchical autoencoder structure is adopted to transform the variable-length/discrete/continuous discrete mixed rectangular boxes and category information into fixed-length/ordered/unnoise-resistant hidden vectors, and the speed and accuracy of object detection are improved through joint training of image feature extractor and diffusion model.
It effectively solves the problems of uncertain size of rectangular box sets and complex element arrangement order, improves the speed and accuracy of object detection, and fully models the joint probability relationship between coordinates and categories.
Smart Images

Figure CN119942066A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image detection, and in particular to a technical solution for object detection based on a latent variable diffusion model. Background Art
[0002] Object detection is a classic problem in computer vision. Given an image, its task is to locate the objects of interest in the image and output the coordinates and category of each object.
[0003] According to the stage of technological development, object detection methods can be divided into traditional methods and deep learning-based methods.
[0004] With the revival of deep learning technology, object detection methods have gradually evolved to the deep learning stage since 2014. RCNN uses deep models to replace artificially designed shallow models to extract image features, which greatly improves the accuracy of object detection. Faster RCNN accelerates the candidate box generation step in RCNN, and for the first time, the speed of deep object detection models has been increased to the quasi-real-time level. YOLO is the first single-stage object detection method that combines the candidate box generation and object classification steps into one, sacrificing a little accuracy in exchange for a significant speed increase. It is currently one of the most widely deployed models in practical applications.
[0005] In 2017, Google proposed the famous Transformer model for machine translation. In the following years, Transformer quickly expanded to related fields, and achieved good results in image classification tasks in 2020, and was called ViT (Vision Transformer). In the same year, Facebook used Transformer to implement the detection model DETR without NMS, and in 2022, ViTDet achieved the best performance in object detection at the time.
[0006] In recent years, a study has proposed a detection method based on the diffusion model, DiffusionDet. This method uses the diffusion model to model the conditional probability distribution of the object rectangle relative to the image, and then completes the detection task by iterative sampling. After obtaining the rectangular box, the category of each box is finally estimated, such as: horse, person, etc. With the advantage of the good generalization of the diffusion model, this method achieves the best performance under the Zero-Shot condition.
[0007] Assume that the input is an image I and the output is a set of rectangular boxes A, where each rectangular box contains the information of the rectangular box coordinates (x, y, w, h) and the category class. During the training phase, DiffusionDet uses a diffusion model in the logarithmic gradient domain to model the conditional probability of A relative to I, denoted as Where F represents the model of the image feature extractor, Φ is the parameter of the model F(), Q() represents the diffusion model, θ is the parameter of the diffusion model Q, P() is the ideal unknown conditional probability density function, is the logarithmic gradient. In the inference phase, the Langevin dynamics formula is used to iteratively sample and predict A from Q.
[0008]
[0009] Among them, t is the iteration number, that is, At-1 is the rectangular frame set of the previous iteration, At is the predicted current rectangular frame set, ε is the iteration step, Et is Gaussian white noise, and * represents multiplication. When the iteration step approaches 0 and the number of iterations approaches infinity, the sampling result is close to the ideal value, but in practical applications, considering the amount of calculation, the number of iterations is generally 10 or less.
[0010] However, the existing DiffusionDet technology has the following shortcomings:
[0011] Since the number of targets in the image is uncertain, the size of the rectangular box set A is variable. For ease of operation, DiffusionDet fills the rectangular boxes of all samples with a fixed length (such as 100). However, whether random or regular filling is used, the distribution of the original data will inevitably be destroyed, resulting in training difficulties.
[0012] The rectangular box set is invariant to the order of elements. For example, the geometric meanings of {(0,0,100,100),(100,100,200,200)} and {(100,100,200,200),(0,0,100,100)} are the same, but they correspond to different points in the vector space. When there are a large number of rectangular boxes and the number is variable, the structure of the vector space will be very complex, making training difficult.
[0013] The rectangular box set only contains coordinate information (x, y, w, h), and the prediction of the category information class is placed after the diffusion model, which makes this method unable to fully model the joint probability relationship between coordinates and categories, such as: "saddle" is usually on the back of "horse", "water cup" is usually on the "table", etc.
[0014] Based on this, the present invention studies a new object detection technology solution. Summary of the invention
[0015] The present application provides an object detection method, apparatus, computer equipment, storage medium and program product, which can solve the technical problems mentioned in the background technology.
[0016] In a first aspect, the present application relates to an object detection method, comprising:
[0017] The image to be detected is input into the image feature extractor;
[0018] The image features extracted by the image feature extractor and Gaussian white noise are input into the diffusion model;
[0019] The diffusion model outputs a latent vector, which is input into the decoder of the autoencoder.
[0020] The decoder outputs the object detection results corresponding to the latent vector.
[0021] In a second aspect, the present application relates to an object detection device, comprising
[0022] An acquisition module, used to acquire the image to be detected;
[0023] A feature extraction module is used to extract image features of the image to be detected;
[0024] The latent vector acquisition module is used to obtain the latent vector. The image features and Gaussian white noise are input into the diffusion model, and the diffusion model outputs the latent vector.
[0025] The object detection module is used to obtain the object detection results. The latent vector is input to the decoder in the autoencoder, and the decoder outputs the object detection results corresponding to the latent vector.
[0026] In a third aspect, the present application relates to an object detection model training method, comprising:
[0027] Autoencoder training and joint training of image feature extractor and diffusion model;
[0028] The training of the autoencoder includes: the original object information set is input into the encoder, the latent vector is output from the encoder, the latent vector is input into the decoder, and the decoder outputs the reconstructed object information set;
[0029] The image feature extractor and the diffusion model are jointly trained as follows: the image is input into the image feature extractor to obtain the image features, and the image features, Gaussian white noise and latent vectors are input into the diffusion model.
[0030] In a fourth aspect, the present application relates to an object detection model training device, comprising:
[0031] An autoencoder training device and a joint training device;
[0032] The autoencoder training device includes encoder training and decoder training;
[0033] The joint training device includes image feature extractor training and diffusion model training.
[0034] In a fifth aspect, the present application relates to an electronic device, comprising:
[0035] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first aspect or the third aspect above.
[0036] In a sixth aspect, the present application relates to a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect or the third aspect.
[0037] In a seventh aspect, the present application relates to a computer program product, including a computer program, which implements the method described in the first aspect or the third aspect when executed by a processor.
[0038] The object detection method of the present invention transforms the variable-length / unordered / continuous-discrete mixed rectangular frame and category information into a fixed-length / ordered / noise-resistant latent vector; the hierarchical autoencoder structure provided by the present invention compresses the vector dimension as much as possible while maintaining the good properties of the latent vector space, thereby improving the speed and accuracy of the object detection method. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 is a schematic diagram of a training and inference principle based on an object detection model according to an embodiment of the present invention;
[0041] Figure 2 is a training flow chart of an object information autoencoder provided by an embodiment of the present invention;
[0042] Figure 3 is a flowchart of the joint training of an image feature extractor and a diffusion model provided by an embodiment of the present invention;
[0043] Figure 4 It is a flow chart of the object detection method provided by an embodiment of the present invention.
[0044] Figure 5 It is a structural diagram of a hierarchical object information autoencoder provided in an embodiment of the present invention.
[0045] Figure 6 It is a hierarchical device structure diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0047] Figure 1 It is a schematic diagram of training and reasoning based on an object detection model provided according to an embodiment of the present invention.
[0048] The disclosed embodiments are applicable to the case where a multi-class target detection model is trained in any image. The method can be performed by a training device for an object detection model or an object detection device, which can be implemented in hardware and / or software and can be configured in an electronic device.
[0049] from Figure 1 As can be seen from the figure, the autoencoder is trained first, and then the image feature extraction model and diffusion model are trained. Use the trained model to perform inference.
[0050] Figure 2 It is a training flow chart of the object information autoencoder provided in an embodiment of the present invention.
[0051] refer to Figure 2 ,The autoencoder training includes : the original object information set is input into the encoder, the latent vector is output from the encoder, the latent vector is input into the decoder, and the decoder outputs the reconstructed object information set.
[0052] Furthermore, the training of the autoencoder is driven by two objective functions: rectangular box coordinate reconstruction and label classification.
[0053] The present invention proposes to use an autoencoder to transform a variable-length / unordered / continuous-discrete mixed rectangular frame and category set, recorded as an object information set B = {(x, y, w, h, class)...}, into a fixed-length / ordered / noise-resistant latent vector Z, and at the same time obtain Z to reconstruct the object information set The inverse transform of .
[0054] In some embodiments, the rectangular box (x, y, w, h) is a four-dimensional real vector. To unify the numerical range, it is normalized to [0, 1] according to the length and width of the image, the upper left corner of the image is [0, 0], and the lower right corner is [1, 1]. The category label class is an integer, and the embodiment further preferably uses a lookup table to map it to a D-dimensional vector. The size of the lookup table is (number of categories + 1) × D, +1 is reserved for the background category (or negative sample), and the elements of the lookup table are trainable parameters. Assuming that there are N objects in the image I, its object information Bi can be represented by an N×(D+4)-dimensional vector.
[0055] Since the number of objects N in each image may be different, the size of the N×(D+4)-dimensional vector may not be consistent. To facilitate probabilistic modeling, it is further proposed to use the cross-attention and self-attention mechanisms to construct a hierarchical encoder, transforming the N×(D+4)-dimensional vector into a fixed-length latent vector Z, whose size is K×L×C, where K is the number of query tokens in the cross-attention layer, indicating the maximum number of objects in the image, and is generally recommended to be set to 100 for typical scenarios; L is the number of encoder layers, and is preferably recommended to be set to about 6; C is the dimension of the cross-attention and self-attention layers, which is generally set to between 512-1024.
[0056] In the embodiment, the latent vector corresponding to the i-th layer is recorded as Zi, and the encoding formula implemented by the encoder part in the object information autoencoder can be expressed as:
[0057] Z={Z1,Z2,...ZL}
[0058] Zi=CrossAtti(SelfAtti(...SelfAtt1(Proj(B))))
[0059] Then, the hidden vector Zi is concatenated layer by layer using the feature concatenation layer Access the decoder, through L self-attention layers and a linear projection layer Finally, the final output is a vector of size K×(D+4), the first D columns of which represent the category information, and the last 4 columns represent the coordinates of the rectangular box.
[0060] The decoding formula implemented by the decoder part of the object information autoencoder can be expressed as:
[0061]
[0062] Among them, Proj() is a linear projection, Concat() is a feature concatenation operation, SelfAtt() is a self-attention operation, and CrossAtt() is a cross-attention operation. The corresponding self-attention operation of the i-th self-attention layer in the encoder and decoder is marked as SelfAtti(), and the corresponding cross-attention operation is marked as CrossAtti().
[0063] As a preferred embodiment, the autoencoder includes an encoder and a decoder, the encoder includes a linear projection layer and L self-attention layers, each self-attention layer is correspondingly provided with a cross-attention layer, the decoder includes a linear projection layer and L self-attention layers, and in addition, a feature splicing layer is provided between adjacent self-attention layers.
[0064] As a preferred embodiment, the reconstruction loss function of the autoencoder consists of three parts:
[0065] (1) The original object information set B and the reconstructed object information set To match the objects in the image, and obtain the optimal matching relationship, the Hungarian matching algorithm can be used for matching.
[0066] (2) Calculate the reconstruction loss of the rectangular box of the matched object.
[0067] (3) Calculate the classification loss for the matched objects.
[0068] For rectangular box reconstruction, the GIoU loss function is preferably used, and for object classification, the Softmax regression loss function is preferably used.
[0069] In addition, during the training process, it is also possible to preferably add noise to the latent vector Z in a manner similar to a variational autoencoder, so as to obtain a latent vector expression that is more robust to noise.
[0070] After the autoencoder is trained, the subsequent steps can use the encoder or decoder part respectively.
[0071] The encoder part of the trained object information autoencoder transforms the object information set B into a latent vector Z with better properties.
[0072] refer to Figure 3 ,The joint training of the image feature extractor and the diffusion model includes: the image is input to the image feature extractor to obtain the image features, and the image features ,Gaussian white noise and latent vector are input to the diffusion model.
[0073] Furthermore, when the image feature extractor is jointly trained with the diffusion model, the latent vector input into the diffusion model is the encoder in the trained autoencoder transforming the original object information set B into the latent vector Z.
[0074] The model F(I; Φ) used by the image feature extractor can be implemented by a convolutional network or a Transformer, and the feature map size of its output is H×W×C. The feature map is then input as a condition into the diffusion model Q(Z, F(I; Φ); θ), which can be implemented by UNet or Transformer. The feature map and the diffusion model are preferably connected by a cross attention layer, or they can be connected by other methods.
[0075] Similar to the principle of the ordinary diffusion model, Q and F jointly model the gradient of the logarithmic conditional probability density function of Z with respect to the image I:
[0076]
[0077] Its loss function is:
[0078]
[0079] Among them, E is Gaussian white noise, α is the noise intensity, Φ is the model parameter of the image feature extractor, and θ is the function parameter of the diffusion model.
[0080] After the autoencoder, image feature extractor, and diffusion model are trained, they can be combined for object detection.
[0081] Figure 4 is a flow chart of an object detection method provided according to an embodiment of the present invention. The method can be performed by an object detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 4 , the method specifically includes the following:
[0082] The image to be detected is input into the image feature extractor;
[0083] The image features extracted by the image feature extractor and Gaussian white noise are input into the diffusion model;
[0084] The diffusion model outputs a latent vector, which is input into the decoder of the autoencoder.
[0085] The decoder outputs the object detection results corresponding to the latent vector.
[0086] In this embodiment, the image feature extractor is not limited, for example, a convolutional neural network or a Transformer network may be used.
[0087] Furthermore, in the decoder, a linear projection layer and L self-attention layers are included.
[0088] Furthermore, a feature splicing layer is provided between adjacent self-attention layers.
[0089] Furthermore, after L layers of self-attention layers and linear projection layers, the final output is the reconstructed object information set That is, the object detection result, which includes the object rectangle and category information.
[0090] refer to Figure 5 , the specific working process of the object information autoencoder is preferably:
[0091] In the encoder, there is a linear projection layer and L self-attention layers. Each self-attention layer is provided with a corresponding cross-attention layer. After the object information set B is input into the self-attention layer 1 via the linear projection layer, it passes through the corresponding cross-attention layer to obtain the latent vector of the first layer, denoted as Z1; after Z1 is input into the self-attention layer 2, it passes through the corresponding cross-attention layer to obtain the latent vector of the second layer, denoted as Z2; … After the latent vector ZL-1 of the L-1th layer is input into the self-attention layer L, it passes through the corresponding cross-attention layer to obtain the latent vector of the Lth layer, denoted as ZL.
[0092] In the decoder, it also includes a linear projection layer and L self-attention layers. In addition, a feature splicing layer is set between adjacent self-attention layers. After the latent vector ZL of the Lth layer is input into the self-attention layer L, the output of the self-attention layer L and the latent vector ZL-1 of the L-1th layer are input into the self-attention layer L-1 using the feature splicing layer. The output of the self-attention layer L-1 and the latent vector ZL-2 of the L-2th layer are input into the self-attention layer L-2 using the feature splicing layer, ... The output of the self-attention layer 2 and the latent vector Z1 of the 1st layer are input into the self-attention layer 1 using the feature splicing layer. After the output of the self-attention layer 1 passes through the linear projection layer, the final output is the reconstructed object information set.
[0093] In the inference process, i.e., the object detection process, the image I and Gaussian white noise E are input into the diffusion model Q to obtain the gradient of the conditional probability density function. Based on the Langevin dynamics formula, Zt-1 can be iteratively denoised along the gradient direction. After several iterations, the decoder is used to restore the corresponding object rectangle and category information from the latent vector. Among them, Zt-1 represents the object information set of the previous iteration, Zt is the predicted current object information set, ε is the iteration step, and Et is Gaussian white noise.
[0094] refer to Figure 4 In the embodiment, given an image to be detected I, the model F() of the image feature extractor is first used to extract the feature map. Then the image feature map and Gaussian white noise are input into the diffusion model, and based on the diffusion model Q(), the target latent vector Z is obtained by sampling using the Langevin dynamics formula.
[0095]
[0096] Finally, the decoder part of the trained object information autoencoder is used to restore Z to the final object detection result.
[0097] In specific implementation, the above process can be automatically run using computer software technology and is applicable to multi-class target detection in any image.
[0098] In another embodiment of the present invention, an object detection device is provided, comprising:
[0099] An acquisition module, used to acquire the image to be detected;
[0100] A feature extraction module is used to extract image features of the image to be detected;
[0101] The latent vector acquisition module is used to obtain the latent vector. The image features and Gaussian white noise are input into the diffusion model, and the diffusion model outputs the latent vector.
[0102] The object detection module is used to obtain the object detection results. The latent vector is input to the decoder in the autoencoder, and the decoder outputs the object detection results corresponding to the latent vector.
[0103] In another embodiment of the present invention, there is also provided an object detection model training device, comprising an autoencoder training device and a joint training device;
[0104] The autoencoder training device includes encoder training and decoder training;
[0105] The joint training device includes image feature extractor training and diffusion model training.
[0106] In addition, the present application also relates to an electronic device, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned object detection method or the above-mentioned object detection model training method.
[0107] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor may call the logic instructions in the memory to execute the above-mentioned object detection method or the above-mentioned object detection model training method.
[0108] The present application also relates to a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the above-mentioned object detection method or the above-mentioned object detection model training method.
[0109] The present application also relates to a computer program product, comprising a computer program, which implements the above-mentioned object detection method or the above-mentioned object detection model training method when executed by a processor.
[0110] In order to understand the technical effect of the present invention, the object detection method proposed in the present invention is tested based on a certain interface analysis task, and the corresponding comparative data with the prior art DiffusionDet method are shown in Table 1. The coordinates and category information of the elements such as buttons, icons, groups, lists, etc. that developers focus on in the interface image are basically accurate, which provides a good foundation for subsequent functions such as automatic interface editing and code generation.
[0111] Based on the AP (Average Precision) indicators of various elements in Table 1, it can be seen that the method based on the latent variable diffusion model shows a higher accuracy than DiffusionDet in the tested interface analysis task.
[0112] Table 1
[0113] AP Button icon Grouping List DiffusionDet 75.7 61.4 70.2 62.8 Hidden variable diffusion model 79.4 73.5 84.1 83.6
[0114] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0115] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An object detection method, characterized in that: include: The image to be detected is input into the image feature extractor; The image features extracted by the image feature extractor and Gaussian white noise are input into the diffusion model; The diffusion model outputs a latent vector, which is input into the decoder of the autoencoder. The decoder outputs the object detection results corresponding to the latent vector.
2. The object detection method according to claim 1, characterized in that: In the decoder, there is a line projection layer and multiple self-attention layers.
3. The object detection method according to claim 2, characterized in that: A feature splicing layer is set between adjacent self-attention layers.
4. An object detection device, characterized in that: include: An acquisition module, used to acquire the image to be detected; A feature extraction module is used to extract image features of the image to be detected; The latent vector acquisition module is used to obtain the latent vector. The image features and Gaussian white noise are input into the diffusion model, and the diffusion model outputs the latent vector. The object detection module is used to obtain the object detection results. The latent vector is input to the decoder in the autoencoder, and the decoder outputs the object detection results corresponding to the latent vector.
5. A method for training an object detection model, characterized in that: include: Autoencoder training and joint training of image feature extractor and diffusion model; The training of the autoencoder includes: the original object information set is input into the encoder, the latent vector is output from the encoder, the latent vector is input into the decoder, and the decoder outputs the reconstructed object information set; The image feature extractor and the diffusion model are jointly trained as follows: the image is input into the image feature extractor to obtain the image features, and the image features, Gaussian white noise and latent vectors are input into the diffusion model.
6. The object detection model training method according to claim 5, characterized in that: The encoder is responsible for transforming the original object information set into a latent vector, and the decoder is responsible for transforming the latent vector into a reconstructed object information set.
7. An object detection model training device, comprising An autoencoder training device and a joint training device; The autoencoder training device includes encoder training and decoder training; The joint training device includes image feature extractor training and diffusion model training.
8. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the method as claimed in any one of claims 1 to 3 or any one of claims 5 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make the computer execute the method according to any one of claims 1 to 3 or any one of claims 5 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 3 or any one of claims 5 to 6 is implemented.