Target detection method and device applied to terminal equipment, equipment and storage medium
Patent Information
- Application Number
- CN202411037743.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-30
AI Technical Summary
[0004]本发明实施例提供一种应用于端设备的目标检测方法、装置、计算机设备及存储介质,以解决端设备的目标检测效率较低的问题
[0041]上述应用于端设备的目标检测方法、装置、计算机设备及存储介质,通过在特征图像中预测目标特征点的偏移范围,进而将偏移范围中的特征点拷贝到端设备的缓存单元中,然后,在偏移范围内对特征图像中进行采样,得到每个偏移量对应的特征值,以根据注意力权重对特征值进行特征融合,得到用于生成目标检测结果的融合特征值。可见,本实施例中只需要拷贝一次数据,即特征图像中偏移范围内的特征值,因此,在偏移范围内对特征图像中进行采样得到每个偏移量对应的特征值后,直接根据缓存单元中对应的特征值进行后续的注意力机制运算即可,无需再进行重复的拷贝操作,相较于传统的在端设备上运行的可形变注意力机制,可以达到提高端设备的目标检测效率的目的。
Smart Images

Figure CN121437836B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and in particular to a target detection method, apparatus, device, and storage medium for use in edge devices. Background Technology
[0002] With the rapid development of science and technology, attention mechanisms, as one of the key technologies for object detection, have been widely used in various technical fields, such as image, speech, and video.
[0003] While deformable attention mechanisms, as a variant of attention mechanisms, have lower computational complexity compared to traditional attention mechanisms, in practical applications, each feature point predicts dozens or even hundreds of offset feature points. Sampling these offset feature points requires dozens or even hundreds of small data copy operations. Each time a feature point is sampled, it needs to be copied to the cache unit of the end device. These fragmented data copy operations significantly increase the latency of memory access on the end device, especially on devices with limited memory bandwidth. This significantly impacts the overall performance of the end device, resulting in lower target detection efficiency. Summary of the Invention
[0004] This invention provides a target detection method, apparatus, computer device, and storage medium for use in terminal devices, in order to solve the problem of low target detection efficiency in terminal devices.
[0005] A target detection method applied to an end device, the method comprising:
[0006] Obtain a feature image and a feature sequence corresponding to the feature image, wherein the feature sequence and the feature image contain predetermined target feature points;
[0007] The feature image is input into the first linear layer to predict the offset range of the target feature point in the feature image, so that all feature points in the feature image that are within the offset range are copied to the cache unit of the terminal device, and the target feature point is the center point of the offset range;
[0008] The feature sequence is input into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset;
[0009] Based on the offset, the feature image is sampled within the offset range to obtain a feature value corresponding to each offset;
[0010] Based on the attention weight corresponding to each offset, feature fusion is performed on the feature value corresponding to each offset to obtain a fused feature value, which is used to generate the target detection result.
[0011] Optionally, in the above method, the step of inputting the feature image into the first linear layer and predicting the offset range of the target feature points in the feature image includes:
[0012] The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image;
[0013] The target radius corresponding to the detected target is determined according to the target category;
[0014] The offset range is obtained based on the compensation radius and the target radius.
[0015] Optionally, in the above method, sampling the feature image within the offset range according to the offset to obtain the feature value corresponding to each offset includes:
[0016] Based on the offset, the offset coordinates of the target feature point after offset are obtained;
[0017] The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
[0018] Optionally, in the above method, the step of inputting the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset includes:
[0019] The feature sequence is input into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head;
[0020] The offset predicted by each attention head is normalized to obtain the attention weight corresponding to each offset.
[0021] Optionally, in the above method, the step of performing feature fusion on the feature values corresponding to each offset based on the attention weight corresponding to each offset to obtain fused feature values includes:
[0022] The feature value corresponding to the offset and the attention weight corresponding to the offset are added together to obtain the first feature value corresponding to each attention head;
[0023] The first feature value corresponding to each attention head is input into the third linear layer to obtain the fused feature value.
[0024] Optionally, in the above method, obtaining the feature image and the feature sequence corresponding to the feature image includes:
[0025] Acquire the image to be detected;
[0026] Feature extraction is performed on the image to be detected to obtain the feature image corresponding to the image to be detected;
[0027] The feature image is divided into grids to obtain a set of feature points corresponding to the feature image;
[0028] All feature points in the feature point set are sorted to obtain the feature sequence corresponding to the feature image.
[0029] A target detection device for use in end devices, the device comprising:
[0030] The feature acquisition module is used to acquire a feature image and a feature sequence corresponding to the feature image, wherein the feature sequence and the feature image contain predetermined target feature points;
[0031] The range segmentation module is used to input the feature image into the first linear layer, predict the offset range of the target feature point in the feature image, and copy all feature points in the feature image that are within the offset range to the cache unit of the terminal device, wherein the target feature point is the center point of the offset range;
[0032] The weight calculation module is used to input the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset;
[0033] The feature value sampling module is used to sample the feature image within the offset range according to the pixel coordinates of the offset, so as to obtain the feature value corresponding to each offset.
[0034] The feature fusion module is used to perform feature fusion on the feature values corresponding to each offset according to the attention weight corresponding to each offset, so as to obtain fused feature values, which are used to generate target detection results.
[0035] Optionally, in the above-described apparatus, the range division module is used for:
[0036] The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image;
[0037] The target radius corresponding to the detected target is determined according to the target category;
[0038] The offset range is obtained based on the compensation radius and the target radius.
[0039] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements a target detection method applied to an end device as described in any of the preceding claims.
[0040] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the target detection method applied to an end device as described in any of the preceding claims.
[0041] The aforementioned target detection method, apparatus, computer device, and storage medium applied to end devices predict the offset range of target feature points in a feature image, then copy the feature points within the offset range to the cache unit of the end device. Next, sampling is performed on the feature image within the offset range to obtain the feature value corresponding to each offset. The feature values are then fused according to attention weights to obtain fused feature values used to generate the target detection result. It is evident that this embodiment only requires copying the data once, i.e., the feature values within the offset range of the feature image. Therefore, after sampling the feature image within the offset range to obtain the feature value corresponding to each offset, subsequent attention mechanism operations can be performed directly based on the corresponding feature values in the cache unit, eliminating the need for repeated copying operations. Compared to traditional deformable attention mechanisms running on end devices, this improves the target detection efficiency of the end device. Attached Figure Description
[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a flowchart illustrating an implementation of a target detection method applied to an end device according to an embodiment of the present invention;
[0044] Figure 2 This is a partial implementation flowchart of a target detection method applied to an end device according to an embodiment of the present invention;
[0045] Figure 3 This is a partial implementation flowchart of a target detection method applied to an end device in one embodiment of the present invention;
[0046] Figure 4 This is a partial implementation flowchart of a target detection method applied to an end device according to an embodiment of the present invention;
[0047] Figure 5 This is a partial implementation flowchart of a target detection method applied to an end device according to an embodiment of the present invention;
[0048] Figure 6 This is a partial implementation flowchart of a target detection method applied to an end device according to an embodiment of the present invention;
[0049] Figure 7 This is a schematic diagram of the target detection device applied to an end device in one embodiment of the present invention;
[0050] Figure 8 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0052] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0053] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0054] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0055] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0057] This invention discloses a target detection method, apparatus, computer device, and storage medium applied to an end device. It predicts the offset range of target feature points in a feature image, then copies the feature points within the offset range to a cache unit on the end device. Next, it samples the feature image within the offset range to obtain feature values corresponding to each offset. These feature values are then fused according to attention weights to obtain fused feature values used to generate the target detection result. As can be seen, this embodiment only requires copying the data once, i.e., the feature values within the offset range of the feature image. Therefore, after sampling the feature image within the offset range to obtain the feature values corresponding to each offset, subsequent attention mechanism operations can be performed directly based on the corresponding feature values in the cache unit, eliminating the need for repeated copying operations. Compared to traditional deformable attention mechanisms running on end devices, this improves the target detection efficiency of the end device. Specific embodiments are described below.
[0058] like Figure 1 The diagram shows a flowchart of a target detection method for edge devices disclosed in this invention. This method is applicable to edge devices with image processing capabilities, such as mobile phones, smart home appliances, various sensors, cameras, drones, and other intelligent devices. The method in this embodiment may specifically include the following steps:
[0059] S101: Obtain the feature image and the corresponding feature sequence. The feature sequence and the feature image contain predetermined target feature points.
[0060] The target feature points may be in the same or different positions in different feature images. The target feature points can be learned in advance through a linear layer. The specific representation of the target feature points can be the position coordinates of the target feature points in the feature image.
[0061] Specifically, in this embodiment, the feature image and the corresponding feature sequence can be obtained through the following steps, such as... Figure 2 As shown:
[0062] S201: Image to be detected has been obtained.
[0063] The image to be detected can be an image captured by an end device, or an image transmitted from a device with a shooting function received by the end device. End devices include, but are not limited to, mobile phones, smart home appliances, various sensors, cameras, drones, and other smart devices. Thus, the image to be detected can be acquired through the end device. Furthermore, the detection targets in the image to be detected in this embodiment can be people, vehicles, animals, plants, etc. By capturing an image containing these detection targets as the image to be detected, the detection targets in the image to be detected can be identified.
[0064] S202: Perform feature extraction on the image to be detected to obtain the feature image corresponding to the image to be detected.
[0065] The image to be detected is input into the feature extraction model to obtain the feature image corresponding to the image to be detected.
[0066] Specifically, the feature extraction model in this embodiment includes, but is not limited to, the ResNet structure model, the AlexNet structure model, and the LeNet5 structure model. The feature extraction model extracts features from the image to be detected to obtain the feature image corresponding to the image to be detected. If necessary, the scale of the extracted feature image can also be adjusted by the feature extraction model to meet the needs of subsequent attention mechanism operations.
[0067] S203: Perform grid-based segmentation on the feature image to obtain the set of feature points corresponding to the feature image.
[0068] The feature image is segmented into grids, which can be 10×10, 20×20, or even 100×100 grids. In this embodiment, the grid segmentation rule for the feature image can be selected according to actual needs. The specific grid segmentation rule is not limited in this embodiment.
[0069] For example, taking the 10×10 grid division of the feature image as an example, if the feature image is divided into 9 equal-distance divisions along the horizontal and vertical directions, 100 feature points can be obtained. These 100 feature points are the feature point set corresponding to the feature image.
[0070] For example, taking the 20×20 grid division of the feature image as an example, if the feature image is divided into 19 equal-distance divisions along the horizontal and vertical directions, 400 feature points can be obtained. These 400 feature points are the feature point set corresponding to the feature image.
[0071] S204: Sort all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0072] Specifically, in this embodiment, all feature points in the feature point set can be sorted according to a preset sorting rule, which can be as follows:
[0073] In the first method, as described in this embodiment, all feature points in the feature point set can be sorted according to the magnitude of their horizontal coordinate values. For example, taking a 10×10 grid division of the feature image as an example, the feature point with coordinates (1,10) is taken as the first feature point of the feature sequence, the feature point with coordinates (2,10) is taken as the second feature point of the feature sequence, and so on, until the feature point with coordinates (10,10) is taken as the tenth feature point of the feature sequence, then the feature point with coordinates (1,9) is taken as the eleventh feature point of the feature sequence, and so on, until the feature point with coordinates (10,9) is taken as the twentieth feature point of the feature sequence, and so on, thus completing the sorting of all feature points in the feature point set and obtaining the feature sequence corresponding to the feature image.
[0074] In the second method, this embodiment sorts all feature points in the feature point set according to the magnitude of their ordinate values. For example, taking a 10×10 grid division of the feature image as an example, the feature point with coordinates (1,10) is taken as the first feature point of the feature sequence, the feature point with coordinates (1,9) is taken as the second feature point of the feature sequence, and so on, until the feature point with coordinates (1,1) is taken as the tenth feature point of the feature sequence, then the feature point with coordinates (2,10) is taken as the eleventh feature point of the feature sequence, and so on, until the feature point with coordinates (2,1) is taken as the twentieth feature point of the feature sequence, and so on, thus sorting all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0075] The third approach, in this implementation, is to randomly arrange all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0076] In summary, this embodiment divides a large image into smaller image blocks by meshing, which enables parallel processing of a large number of image blocks in batches, greatly reducing the blur detection time of large images and reducing local noise interference in the obtained target image to be identified.
[0077] S102: Input the feature image into the first linear layer, predict the offset range of the target feature point in the feature image, and copy all feature points in the feature image that are within the offset range to the cache unit of the end device, with the target feature point being the center point of the offset range.
[0078] It is important to note that in this embodiment, the first linear layer can be pre-trained, and then the feature map can be input into the first linear layer to predict the offset range of the target feature points in the feature image. Specifically, in this embodiment, a radius label can be assigned based on the width and height of the detected target in the training sample image. The training sample image with the radius label is then input into the first linear layer for training. After training is completed, the feature image is input into the first linear layer, and the first linear layer predicts the offset range of the target feature points in the feature image.
[0079] Specifically, the cache unit in this embodiment includes, but is not limited to, RAM (random access memory), etc., which copies all feature points in the feature image that are within the offset range to the RAM memory of the end device, so that the required feature points can be retrieved at any time during subsequent attention mechanism operations without having to perform fragmented memory copy operations.
[0080] S103: Input the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset.
[0081] The feature sequence is input into a second linear layer to predict the offsets and coordinates of the target feature points. Then, the second linear layer is followed by a softmax function for normalization, resulting in the attention weights for each offset. The initial weights are normalized using a softmax function to obtain the attention weights for each offset.
[0082] In a specific implementation, the second linear layer in this embodiment can be an MK-channel linear projection operator, wherein the feature sequence is fed into a 3MK-channel linear projection operator, where the first 2MK channels encode the sampling offset, and the remaining MK channels are fed into the softmax operator to obtain attention weights.
[0083] Specifically, in this embodiment, the offset of the target feature point and the attention weight corresponding to each offset can be obtained through the following steps, such as... Figure 3 As shown:
[0084] S301: Input the feature sequence into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head.
[0085] The number of attention heads and the number of offsets predicted by each attention head can be set according to actual needs, such as 3, 5, or 7 offsets predicted by each attention head. In this embodiment, the number of predicted offsets is not limited. Different attention heads can predict the same number of offsets, but the offsets predicted by different attention heads can be different from each other. For example, the offset direction and offset distance can be different between different offsets.
[0086] S302: Normalize the predicted offset for each attention head to obtain the attention weight corresponding to each offset.
[0087] Specifically, in this embodiment, the offset predicted by each attention head can be normalized using the softmax function to obtain the attention weight corresponding to each offset.
[0088] In summary, this embodiment predicts the offset and then samples the feature image based on the offset to obtain feature points, replacing the traditional attention mechanism that randomly samples the feature image to obtain feature points. This can effectively reduce the number of samplings and improve the efficiency of target detection.
[0089] S104: Based on the offset, sample the feature image within the offset range to obtain the feature value corresponding to each offset.
[0090] The offset coordinates are used to sample the target feature points based on the offset. After sampling based on the offset coordinates, the feature values corresponding to the offset are obtained.
[0091] Specifically, in this embodiment, the nearest neighbor interpolation algorithm can be used to sample the feature image to obtain the feature value corresponding to each offset, as shown below. Figure 4 As shown:
[0092] S401: Obtain the offset coordinates of the target feature point after offset based on the offset amount.
[0093] Based on the coordinates and offset of the target feature point, obtain the offset coordinates of the target feature point after offset.
[0094] Specifically, in this embodiment, the position coordinates of the target feature point can be used as a reference, and the position coordinates of the target feature point can be offset according to different offsets to obtain the offset coordinates of the target feature point.
[0095] S402: The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
[0096] The offset coordinates of the target feature point after offset are input into the nearest neighbor interpolation algorithm calculation formula, and then the nearest neighbor interpolation algorithm is used to sample the feature image. The sampled feature value is used as the feature value corresponding to the offset.
[0097] The nearest neighbor interpolation algorithm in this embodiment can be as follows:
[0098] f(x,y)=f(round(x),round(y))
[0099] Where round(x) and round(y) represent the nearest integer coordinates of the target feature point, x represents the x-coordinate of the target feature point, y represents the y-coordinate of the target feature point, and f(x,y) represents the feature value obtained by sampling through the nearest neighbor interpolation algorithm.
[0100] As is well known, existing deformable attention mechanisms primarily use bilinear interpolation to sample the feature image based on an offset, thereby obtaining the feature value corresponding to the offset. The calculation formula for the bilinear interpolation algorithm is shown below:
[0101] f(x,y)=(1-dy)(1-dy)f(x1,y1)+dx(1-dy)f(x2,y1)+(1-dx)dyf(x1,y2)+dxdyf(x2,y2)
[0102] Where (x1,y1), (x2,y1), (x1,y2), and (x2,y2) represent the position coordinates of the four nearest feature points of the target feature point in the feature image, f(x1,y1), f(x2,y1), f(x1,y2), and f(x2,y2) represent the feature values of the four nearest feature points of the target feature point in the feature image, dx and dy represent the offsets, and f(x,y) represents the sampled feature values.
[0103] As can be seen, the bilinear interpolation algorithm not only needs to consider the coordinates of the four nearest feature points, but also requires multiple multiplications and additions when sampling the feature image based on the offset and the coordinates of the four nearest feature points. Therefore, compared with the nearest neighbor interpolation algorithm, the nearest neighbor interpolation algorithm requires fewer computations to complete the offset sampling of the feature image. Thus, the target detection algorithm in this embodiment has higher detection efficiency.
[0104] S105: Based on the attention weight corresponding to each offset, feature fusion is performed on the feature value corresponding to each offset to obtain fused feature value, which is used to generate target detection results.
[0105] Based on the attention weight corresponding to each offset and the feature value corresponding to each offset, the attention weight corresponding to each feature value is determined. The attention weights corresponding to each feature value are added together, and then passed through a linear layer to obtain the fused feature value.
[0106] Furthermore, when multiple attention heads exist, the feature sequence is input into the second linear layer to obtain the offset of the target feature point predicted by each attention head and the attention weight corresponding to each offset. Each attention head can predict multiple corresponding offsets, and the number of offsets can be set according to actual needs. For example, the number of offsets predicted by each attention head can be 4, 5, etc. In this embodiment, the number of attention heads and the number of offsets predicted by each attention head are not limited.
[0107] In this implementation, the Transformer Decoder interacts with the fused feature values to generate a corresponding feature representation. This feature representation contains the location and category information of the detected target in the image. The feature representation is then passed to two independent feed-forward networks to predict the target bounding box and category label, i.e., the target detection result. The category label represents the classification category of the detected target within the bounding box. The target bounding box can be obtained by inversely normalizing the target's location information and transforming it to the original image coordinate system; that is, the image within the target bounding box is the detected target image. The decoder in this embodiment is pre-trained with object query network parameters, which guide the network model to output the target bounding box and category label. After the fused feature values are input into the decoder, the object query network parameters interact with the fused feature values to generate a corresponding feature representation, which is then passed to two independent feed-forward networks to predict the target detection result.
[0108] Specifically, in this embodiment, feature fusion can be performed on the feature values corresponding to each offset based on the attention weights corresponding to each offset, to obtain fused feature values, such as... Figure 5 As shown:
[0109] S501: Add the feature values corresponding to the offset and the attention weights corresponding to the offset to obtain the weighted feature values corresponding to each attention head.
[0110] S502: Input the weighted feature value corresponding to each attention head into the third linear layer to obtain the fused feature value.
[0111] Based on the attention weight corresponding to each offset and the feature value corresponding to each offset, the attention weight corresponding to each feature value is determined. The attention weights corresponding to each feature value are added together to obtain the weighted feature value. Then, the weighted feature value is linearly mapped through the third linear layer to obtain the fused feature value.
[0112] In summary, this invention discloses a target detection method applied to edge devices. It predicts the offset range of target feature points in a feature image, copies the feature points within the offset range to a cache unit on the edge device, and then samples the feature image within the offset range to obtain feature values corresponding to each offset. These feature values are then fused according to attention weights to obtain fused feature values used to generate the target detection result. As can be seen, this embodiment only requires copying the data once—the feature values within the offset range of the feature image. Therefore, after sampling the feature image within the offset range to obtain the feature values corresponding to each offset, subsequent attention mechanism operations can be performed directly based on the corresponding feature values in the cache unit, eliminating the need for repeated copying operations. Compared to traditional deformable attention mechanisms running on edge devices, this method can improve the target detection efficiency of edge devices.
[0113] based on Figure 1 In its specific implementation, step S102 in this embodiment can be achieved through the following steps, such as... Figure 6 As shown:
[0114] S601: Input the feature image into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image.
[0115] The feature image is input into the first linear layer, which then detects the target category in the feature image and predicts the supplementary radius. The target category includes, but is not limited to, people, vehicles, etc.
[0116] S602: Determine the target radius corresponding to the detected target based on the target category.
[0117] Specifically, in this embodiment, different target radii can be assigned according to different target categories. The target can be a fixed value. For example, taking a person as an example, the target radius is 6 centimeters. In a feature image with a length of 30 centimeters and a width of 20 centimeters, when the target category of the detected target is determined to be a person, the corresponding target radius can be determined to be 6 centimeters.
[0118] S603: Obtain the offset range based on the compensation radius and the target radius.
[0119] The offset range determined by the compensation radius and the target radius can be a circular range, which is a circular range centered on the target feature point.
[0120] It is important to note that the offset range determined solely by the target radius may be too small to support the convergence of the target detection model. Therefore, another learnable radius, namely the compensation radius, can be set in the model to adjust the target radius, and then the offset range can be determined based on the adjusted target radius.
[0121] Specifically, in this embodiment, a compensation radius can be added to the target radius to obtain a new target radius. Based on the new target radius, a circular range centered on the target feature point is determined, and this circular range is the offset range.
[0122] In addition, the compensation radius in this embodiment can be a negative or positive value. When the compensation radius is a positive value, the area of the offset range can be increased by adding the compensation radius to the target radius. When the compensation radius is a negative value, the area of the offset range can be reduced by adding the compensation radius to the target radius.
[0123] In summary, this embodiment dynamically adjusts the target radius by learning the compensation radius, and then determines the offset range based on the adjusted target radius. This not only limits the range of attention offset, but also excludes sampling points that are far from the target feature points, i.e., sampling points that are outside the offset range, which can further improve the target detection efficiency of edge devices.
[0124] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0125] like Figure 7 The diagram shown is a schematic representation of a target detection device for edge devices disclosed in this invention. This device is suitable for edge devices with image processing capabilities, such as mobile phones, smart home appliances, various sensors, cameras, drones, and other intelligent devices. The method in this embodiment may specifically include the following steps:
[0126] The feature acquisition module 701 is used to acquire a feature image and a feature sequence corresponding to the feature image. The feature sequence and the feature image contain predetermined target feature points.
[0127] The range division module 702 is used to input the feature image into the first linear layer, predict the offset range of the target feature point in the feature image, and copy all feature points in the feature image that are within the offset range to the cache unit of the end device, with the target feature point being the center point of the offset range.
[0128] The weight calculation module 703 is used to input the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset;
[0129] The feature value sampling module 704 is used to sample the feature image within the offset range according to the pixel coordinates of the offset, and obtain the feature value corresponding to each offset.
[0130] The feature fusion module 705 is used to perform feature fusion on the feature values corresponding to each offset according to the attention weight corresponding to each offset, so as to obtain fused feature values, which are used to generate target detection results.
[0131] In summary, the target detection device disclosed in this embodiment, applied to an end device, predicts the offset range of target feature points in a feature image, copies the feature points within the offset range to a cache unit on the end device, and then samples the feature image within the offset range to obtain feature values corresponding to each offset. These feature values are then fused according to attention weights to obtain fused feature values used to generate the target detection result. It is evident that this embodiment only requires copying data once, namely the feature values within the offset range of the feature image. Therefore, after sampling the feature image within the offset range to obtain the feature values corresponding to each offset, subsequent attention mechanism operations can be performed directly based on the corresponding feature values in the cache unit, eliminating the need for repeated copying operations. Compared to traditional deformable attention mechanisms running on end devices, this significantly improves the target detection efficiency of the end device.
[0132] In one implementation, the range partitioning module 702 is used for:
[0133] The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image;
[0134] Determine the target radius corresponding to the target being detected based on the target category;
[0135] The offset range is obtained based on the compensation radius and the target radius.
[0136] In one implementation, the eigenvalue sampling module 704 is used for:
[0137] Based on the offset, obtain the offset coordinates of the target feature point after offset;
[0138] The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
[0139] In one implementation, the weight calculation module 703 is used for:
[0140] The feature sequence is input into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head;
[0141] The offset predicted for each attention head is normalized to obtain the attention weight corresponding to each offset.
[0142] In one implementation, the feature fusion module 705 is used for:
[0143] The first feature value corresponding to each attention head is obtained by weighting the feature value corresponding to the offset and the attention weight corresponding to the offset.
[0144] The first feature value corresponding to each attention head is input into the third linear layer to obtain the fused feature value.
[0145] In one implementation, the feature acquisition module 701 is used for:
[0146] Acquire the image to be detected;
[0147] Feature extraction is performed on the image to be detected to obtain the feature image corresponding to the image to be detected;
[0148] The feature image is divided into grids to obtain the set of feature points corresponding to the feature image;
[0149] Sort all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0150] For specific limitations regarding the target detection device applied to end devices, please refer to the relevant limitations on the target detection method applied to end devices mentioned above, which will not be repeated here. Each module in the aforementioned target detection device applied to end devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0151] This application also discloses a computer device, which can be a server, and its internal structure diagram can be as follows. Figure 7As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a target detection method applied to an end device.
[0152] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0153] Obtain the feature image and the corresponding feature sequence. The feature sequence and the feature image contain predetermined target feature points.
[0154] The feature image is input into the first linear layer to predict the offset range of the target feature point in the feature image. All feature points in the feature image that are within the offset range are copied to the cache unit of the end device, with the target feature point being the center point of the offset range.
[0155] The feature sequence is input into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset;
[0156] Based on the offset, the feature image is sampled within the offset range to obtain the feature value corresponding to each offset;
[0157] Based on the attention weight corresponding to each offset, feature fusion is performed on the feature value corresponding to each offset to obtain fused feature value, which is used to generate target detection results.
[0158] In one implementation, the feature image is input into a first linear layer to predict the offset range of target feature points in the feature image, including:
[0159] The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image;
[0160] Determine the target radius corresponding to the target being detected based on the target category;
[0161] The offset range is obtained based on the compensation radius and the target radius.
[0162] In one implementation, the feature image is sampled within the offset range based on the offset to obtain the feature value corresponding to each offset, including:
[0163] Based on the offset, obtain the offset coordinates of the target feature point after offset;
[0164] The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
[0165] In one implementation, the feature sequence is input into the second linear layer to obtain the offsets of the target feature points and the attention weights corresponding to each offset, including:
[0166] The feature sequence is input into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head;
[0167] The offset predicted for each attention head is normalized to obtain the attention weight corresponding to each offset.
[0168] In one implementation, feature fusion is performed on the feature values corresponding to each offset based on the attention weights corresponding to each offset, resulting in fused feature values, including:
[0169] The feature value corresponding to the offset and the attention weight corresponding to the offset are added together to obtain the first feature value corresponding to each attention head;
[0170] The first feature value corresponding to each attention head is input into the third linear layer to obtain the fused feature value.
[0171] In one implementation, obtaining the feature image and the corresponding feature sequence includes:
[0172] Acquire the image to be detected;
[0173] Feature extraction is performed on the image to be detected to obtain the feature image corresponding to the image to be detected;
[0174] The feature image is divided into grids to obtain the set of feature points corresponding to the feature image;
[0175] Sort all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0176] This application also discloses a computer-readable storage medium that, when executed by a processor in a computer device, enables the computer device to perform various steps of any embodiment of a target detection method for an end device disclosed in this invention. The computer-readable storage medium may be non-volatile or volatile.
[0177] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0178] Obtain the feature image and the corresponding feature sequence. The feature sequence and the feature image contain predetermined target feature points.
[0179] The feature image is input into the first linear layer to predict the offset range of the target feature point in the feature image. All feature points in the feature image that are within the offset range are copied to the cache unit of the end device, with the target feature point being the center point of the offset range.
[0180] The feature sequence is input into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset;
[0181] Based on the offset, the feature image is sampled within the offset range to obtain the feature value corresponding to each offset;
[0182] Based on the attention weight corresponding to each offset, feature fusion is performed on the feature value corresponding to each offset to obtain fused feature value, which is used to generate target detection results.
[0183] In one implementation, the feature image is input into a first linear layer to predict the offset range of target feature points in the feature image, including:
[0184] The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image;
[0185] Determine the target radius corresponding to the target being detected based on the target category;
[0186] The offset range is obtained based on the compensation radius and the target radius.
[0187] In one implementation, the feature image is sampled within the offset range based on the offset to obtain the feature value corresponding to each offset, including:
[0188] Based on the offset, obtain the offset coordinates of the target feature point after offset;
[0189] The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
[0190] In one implementation, the feature sequence is input into the second linear layer to obtain the offsets of the target feature points and the attention weights corresponding to each offset, including:
[0191] The feature sequence is input into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head;
[0192] The offset predicted for each attention head is normalized to obtain the attention weight corresponding to each offset.
[0193] In one implementation, feature fusion is performed on the feature values corresponding to each offset based on the attention weights corresponding to each offset, resulting in fused feature values, including:
[0194] The first feature value corresponding to each attention head is obtained by weighting the feature value corresponding to the offset and the attention weight corresponding to the offset.
[0195] The first feature value corresponding to each attention head is input into the third linear layer to obtain the fused feature value.
[0196] In one implementation, obtaining the feature image and the corresponding feature sequence includes:
[0197] Acquire the image to be detected;
[0198] Feature extraction is performed on the image to be detected to obtain the feature image corresponding to the image to be detected;
[0199] The feature image is divided into grids to obtain the set of feature points corresponding to the feature image;
[0200] Sort all feature points in the feature point set to obtain the feature sequence corresponding to the feature image.
[0201] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0202] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0203] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A target detection method applied to an end device, characterized in that, The method includes: Obtain a feature image and a feature sequence corresponding to the feature image, wherein the feature sequence and the feature image contain predetermined target feature points; The feature image is input into the first linear layer to predict the offset range of the target feature point in the feature image, so that all feature points in the feature image that are within the offset range are copied to the cache unit of the terminal device, and the target feature point is the center point of the offset range; The feature sequence is input into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset; Based on the offset, the feature image is sampled within the offset range to obtain a feature value corresponding to each offset; Based on the attention weight corresponding to each offset, feature fusion is performed on the feature value corresponding to each offset to obtain a fused feature value, which is used to generate the target detection result; The step of inputting the feature image into the first linear layer and predicting the offset range of the target feature points in the feature image includes: The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image; The target radius corresponding to the detected target is determined according to the target category; The offset range is obtained based on the compensation radius and the target radius.
2. The target detection method applied to an end device as described in claim 1, characterized in that, The step of sampling the feature image within the offset range according to the offset to obtain the feature value corresponding to each offset includes: Based on the offset, the offset coordinates of the target feature point after offset are obtained; The feature value corresponding to the offset is obtained by sampling based on the offset coordinate using the nearest neighbor interpolation algorithm.
3. The target detection method applied to end devices as described in claim 1, characterized in that, The step of inputting the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset includes: The feature sequence is input into the second linear layer to obtain multiple offsets of the target feature points predicted by each attention head; The offset predicted by each attention head is normalized to obtain the attention weight corresponding to each offset.
4. The target detection method applied to end devices as described in claim 3, characterized in that, The step of fusing features based on the attention weight corresponding to each offset to obtain fused feature values includes: The feature value corresponding to the offset and the attention weight corresponding to the offset are weighted and calculated to obtain the first feature value corresponding to each attention head; The first feature value corresponding to each attention head is input into the third linear layer to obtain the fused feature value.
5. The target detection method applied to an end device as described in claim 1, characterized in that, The acquisition of the feature image and the corresponding feature sequence includes: Acquire the image to be detected; Feature extraction is performed on the image to be detected to obtain the feature image corresponding to the image to be detected; The feature image is divided into grids to obtain a set of feature points corresponding to the feature image; All feature points in the feature point set are sorted to obtain the feature sequence corresponding to the feature image.
6. A target detection device applied to an end device, characterized in that, The device includes: The feature acquisition module is used to acquire a feature image and a feature sequence corresponding to the feature image, wherein the feature sequence and the feature image contain predetermined target feature points; The range segmentation module is used to input the feature image into the first linear layer, predict the offset range of the target feature point in the feature image, and copy all feature points in the feature image that are within the offset range to the cache unit of the terminal device, wherein the target feature point is the center point of the offset range; The weight calculation module is used to input the feature sequence into the second linear layer to obtain the offset of the target feature point and the attention weight corresponding to each offset; The feature value sampling module is used to sample the feature image within the offset range according to the pixel coordinates of the offset, so as to obtain the feature value corresponding to each offset. The feature fusion module is used to perform feature fusion on the feature values corresponding to each offset according to the attention weight corresponding to each offset, so as to obtain fused feature values, which are used to generate target detection results; The range division module is also used for: The feature image is input into the first linear layer to obtain the target category and compensation radius of the detected target in the feature image; The target radius corresponding to the detected target is determined according to the target category; The offset range is obtained based on the compensation radius and the target radius.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the target detection method applied to an end device as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the target detection method applied to the end device as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Semantic segmentation method and related equipment
CN112749706A
Visual small target detection method based on collaborative saliency
CN117496121A