Image detection methods, apparatus, electronic devices and storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2026-08-14
AI Technical Summary
相关技术中的目标跟踪,在目标清晰、遮挡较少的情况下效果较好,但是在人员密集、遮挡严重、目标近似或者目标发生部分变化的情况下跟踪会失败
[0076]This disclosure involves detecting at least one object in an acquired image to be detected, extracting a first feature for each object, then obtaining a target 3D model of the target object, and extracting a second feature from the target 3D model. Finally, based on the second feature and the first feature of each object, the target object among the at least one object is determined; that is, the object whose first feature matches the second feature is identified as the target object. By extracting a relatively comprehensive set of features, namely the second feature, from the target object's 3D model, the target object can be accurately identified from at least one object, even in situations involving dense objects, severe occlusion, approximate targets, or partial changes in the target, thus enabling accurate tracking of the target object.
Smart Images

Figure CN116563773B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and specifically to an image detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of artificial intelligence technology, various automated processing methods can now be performed on images, such as target tracking in video. Target tracking involves tracking a target in video captured by a camera based on captured or input target information. Target tracking technologies generally perform well when the target is clear and has minimal occlusion, but tracking may fail in situations with dense crowds, severe occlusion, similar targets, or where the target has undergone partial changes. Summary of the Invention
[0003] To overcome the problems existing in the related technologies, the present disclosure provides an image detection method, apparatus, electronic device and storage medium to solve the defects in the related technologies.
[0004] According to a first aspect of the present disclosure, an image detection method is provided, comprising:
[0005] Detect at least one object in the acquired image to be detected, and extract a first feature for each object;
[0006] Obtain the target 3D model of the target object, and extract the second feature of the target 3D model;
[0007] A target object is determined from the at least one object based on the second feature and the first feature of each of the objects, wherein the first feature of the target object matches the second feature.
[0008] In some embodiments, the image to be detected is a frame image from a first video;
[0009] After determining the target object among the at least one object, the method further includes:
[0010] Based on the first feature and / or the second feature of the target object, the target object is detected in every frame of the first video except for the image to be detected.
[0011] In some embodiments, after determining the target object among the at least one object, the method further includes:
[0012] Based on the first feature of the target object, the target 3D model is optimized; and / or,
[0013] After detecting the target object in each frame of the first video except for the image to be detected, the method further includes:
[0014] The target 3D model is optimized based on the first feature of the target object detected in each frame of the image.
[0015] In some embodiments, prior to determining the target object among the at least one object based on the second feature and the first feature of each of the objects, the method further includes:
[0016] Using a decision tree, the first category information of each object is determined based on the first feature of each object;
[0017] Based on the pre-input second category information and / or the third category information in the target 3D model, determine the object to be detected among the at least one object, wherein the first category information of the object to be detected matches the second category information and / or the third category information;
[0018] The step of determining the target object among the at least one object based on the second feature and the first feature of each object includes:
[0019] Based on the second feature and the first feature of each of the objects to be detected, a target object is determined among the at least one objects to be detected.
[0020] In some embodiments, obtaining the target 3D model of the target object includes:
[0021] The target 3D model is obtained from the model library based on the pre-input target object identifier and / or pre-extracted third features, wherein the third features are extracted from the pre-input target object data.
[0022] In some embodiments, the image detection method further includes:
[0023] Detect multiple similar objects in at least one frame of the second video, and extract a fourth feature for each of the similar objects;
[0024] After each frame of image is detected, a 3D model of the corresponding similar object is constructed or optimized based on the fourth feature of each similar object, and the fifth feature of the 3D model of each similar object is extracted. Based on the pre-extracted third feature and the fourth and / or fifth features of each similar object, it is determined whether each similar object is retained as a similar object.
[0025] If one similar object remains and is still considered a similar object after a preset number of frames are detected, the similar object is determined to be the target object, and the 3D model of the similar object is determined to be the target 3D model, and the target 3D model is saved to the model library.
[0026] In some embodiments, detecting multiple similar objects in at least one frame of the second video includes:
[0027] In the case of detecting the first frame image in the second video, multiple objects in the first frame image are detected, and each of the objects is regarded as a similar object;
[0028] In the case of detecting non-first frame images in the second video, each similar object in the non-first frame image is detected based on the fourth and fifth features of each similar object in the previous frame image.
[0029] In some embodiments, constructing or optimizing the 3D model of the corresponding similar object based on the fourth feature of each of the similar objects includes:
[0030] In the case of detecting the first frame image in the second video, a three-dimensional model of the corresponding similar object is constructed based on the fourth feature of each of the similar objects;
[0031] In the case of detecting non-first frame images in the second video, the 3D model of the corresponding similar object is optimized based on the fourth feature of each similar object.
[0032] In some embodiments, determining whether each similar object should be retained as a similar object based on a pre-extracted third feature and a fourth and / or fifth feature of each similar object includes:
[0033] In the case of detecting the first frame image in the second video, based on the pre-extracted third feature and the fourth feature of each of the similar objects, it is determined whether each of the similar objects is retained as a similar object;
[0034] In the case of detecting non-first frame images in the second video, it is determined whether each of the similar objects should be retained as a similar object based on the pre-extracted third feature and the fourth and fifth features of each of the similar objects.
[0035] In some embodiments, the image detection method further includes:
[0036] Using a decision tree, the third category information of the similar objects is determined based on the fourth feature of each similar object;
[0037] Saving the target 3D model to the model library includes:
[0038] The third category information of the target object is added to the target 3D model, and the target 3D model is saved to the model library.
[0039] According to a second aspect of the present disclosure, an image detection apparatus is provided, comprising:
[0040] A detection module is used to detect at least one object in the acquired image to be detected, and to extract a first feature of each object;
[0041] The model module is used to acquire the target 3D model of the target object and extract the second feature of the target 3D model;
[0042] A determining module is configured to determine a target object among the at least one object based on the second feature and a first feature of each of the objects, wherein the first feature of the target object matches the second feature.
[0043] In some embodiments, the image to be detected is a frame image from a first video;
[0044] The image detection device further includes a tracking module for:
[0045] After determining the target object among the at least one objects, the target object is detected in each frame of the first video, excluding the image to be detected, based on the first feature and / or the second feature of the target object.
[0046] In some embodiments, the image detection device further includes an optimization module for:
[0047] After determining the target object among the at least one object, the target 3D model is optimized based on a first feature of the target object; and / or,
[0048] After detecting the target object in each frame of the first video except for the image to be detected, the target 3D model is optimized based on the first feature of the target object detected in each frame.
[0049] In some embodiments, the image detection device further includes a filtering module for:
[0050] Before determining the target object among the at least one object based on the second feature and the first feature of each object, a decision tree is used to determine the first category information of each object based on the first feature of each object;
[0051] Based on the pre-input second category information and / or the third category information in the target 3D model, determine the object to be detected among the at least one object, wherein the first category information of the object to be detected matches the second category information and / or the third category information;
[0052] The determining module is specifically used for:
[0053] Based on the second feature and the first feature of each of the objects to be detected, a target object is determined among the at least one objects to be detected.
[0054] In some embodiments, when the model module is used to obtain the target 3D model of the target object, it is specifically used for:
[0055] The target 3D model is obtained from the model library based on the pre-input target object identifier and / or pre-extracted third features, wherein the third features are extracted from the pre-input target object data.
[0056] In some embodiments, the image detection device further includes a modeling module for:
[0057] Detect multiple similar objects in at least one frame of the second video, and extract a fourth feature for each of the similar objects;
[0058] After each frame of image is detected, a 3D model of the corresponding similar object is constructed or optimized based on the fourth feature of each similar object, and the fifth feature of the 3D model of each similar object is extracted. Based on the pre-extracted third feature and the fourth and / or fifth features of each similar object, it is determined whether each similar object is retained as a similar object.
[0059] If one similar object remains and is still considered a similar object after a preset number of frames are detected, the similar object is determined to be the target object, and the 3D model of the similar object is determined to be the target 3D model, and the target 3D model is saved to the model library.
[0060] In some embodiments, when the modeling module is used to detect multiple similar objects in at least one frame of the second video, it is specifically used to:
[0061] In the case of detecting the first frame image in the second video, multiple objects in the first frame image are detected, and each of the objects is regarded as a similar object;
[0062] In the case of detecting non-first frame images in the second video, each similar object in the non-first frame image is detected based on the fourth and fifth features of each similar object in the previous frame image.
[0063] In some embodiments, when the modeling module is used to construct or optimize a 3D model of a corresponding similar object based on the fourth feature of each similar object, it is specifically used for:
[0064] In the case of detecting the first frame image in the second video, a three-dimensional model of the corresponding similar object is constructed based on the fourth feature of each of the similar objects;
[0065] In the case of detecting non-first frame images in the second video, the 3D model of the corresponding similar object is optimized based on the fourth feature of each similar object.
[0066] In some embodiments, when the modeling module determines whether each of the similar objects should be retained as a similar object based on the pre-extracted third feature and the fourth and / or fifth features of each of the similar objects, it is specifically used for:
[0067] In the case of detecting the first frame image in the second video, based on the pre-extracted third feature and the fourth feature of each of the similar objects, it is determined whether each of the similar objects is retained as a similar object;
[0068] In the case of detecting non-first frame images in the second video, it is determined whether each of the similar objects should be retained as a similar object based on the pre-extracted third feature and the fourth and fifth features of each of the similar objects.
[0069] In some embodiments, the image detection device further includes a classification module for:
[0070] Using a decision tree, the third category information of the similar objects is determined based on the fourth feature of each similar object;
[0071] When the modeling module saves the target 3D model to the model library, it is specifically used for:
[0072] The third category information of the target object is added to the target 3D model, and the target 3D model is saved to the model library.
[0073] According to a third aspect of the present disclosure, an electronic device is provided, the electronic device including a memory and a processor, the memory being used to store computer instructions executable on the processor, and the processor being used to execute the computer instructions based on the image detection method described in the first aspect.
[0074] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0075] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0076] This disclosure involves detecting at least one object in an acquired image to be detected, extracting a first feature for each object, then obtaining a target 3D model of the target object, and extracting a second feature from the target 3D model. Finally, based on the second feature and the first feature of each object, the target object among the at least one object is determined; that is, the object whose first feature matches the second feature is identified as the target object. By extracting a relatively comprehensive set of features, namely the second feature, from the target object's 3D model, the target object can be accurately identified from at least one object, even in situations involving dense objects, severe occlusion, approximate targets, or partial changes in the target, thus enabling accurate tracking of the target object. Attached Figure Description
[0077] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0078] Figure 1 This is a flowchart illustrating an exemplary embodiment of the image detection method disclosed herein;
[0079] Figure 2 This is a flowchart illustrating a model building process according to an exemplary embodiment of this disclosure;
[0080] Figure 3 This is a schematic diagram of the structure of an image device shown in an exemplary embodiment of the present disclosure;
[0081] Figure 4 This is a structural block diagram of an electronic device illustrated in an exemplary embodiment of the present disclosure. Detailed Implementation
[0082] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0083] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0084] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0085] In related technologies, the accuracy of target tracking based on traditional matching principles or deep learning algorithms depends entirely on the model's detection accuracy. This accuracy, in turn, is highly dependent on the quality and quantity of data. Therefore, target tracking algorithms in these technologies perform well when the target is clear and there is minimal occlusion, but fail when there are dense crowds, severe occlusion, or similar targets. Furthermore, once a model can identify a person, it will fail to detect them if they change clothes or hairstyle. In police investigations, after retrieving large amounts of video data, the only way to determine the presence of the target (missing child, criminal) is through human observation (police officers and other personnel related to the person being identified), and tracking the target person is extremely inefficient when tracing them through videos extracted from different locations. Therefore, how to accurately and automatically determine whether certain individuals appear in scene videos at different times and locations from a large volume of video footage has become a pressing problem.
[0086] Firstly, at least one embodiment of this disclosure provides an image detection method, please refer to the appendix. Figure 1 It illustrates the process of the method, including steps S101 to S103.
[0087] The image detection method is used to detect target objects from an image to be detected. The target object can be one of at least one object in the image to be detected. For example, if there are multiple people in the image to be detected, the target object can be one of them, i.e., the target person. The target person can be a missing person, a criminal suspect, or a person under guardianship, etc.
[0088] The image to be detected can be an image captured by an image acquisition device, or a frame image from a video to be detected recorded by the image acquisition device. It is understood that, where each frame image in the video to be detected recorded by the image acquisition device can be used as the image to be detected, the method provided in this application embodiment can be used to sequentially detect each frame image, thereby completing the detection of the video to be detected.
[0089] In one possible scenario, the image to be detected can be a sampled image from the video to be detected, an image extracted from the video according to preset sampling requirements (e.g., the first frame, last frame, or a preset number of frames), or an image randomly selected from the video to be detected. The detection of the target object is performed on the sampled image. If the target object is detected in the sampled image, it means that the target object exists in the video to be detected, and the target object can then be detected and tracked. If the target object is not detected in the sampled image, it means that the target object does not exist in the video to be detected. Taking finding a missing person in surveillance video as an example, videos captured by cameras in different locations can be retrieved. With a large number of videos, this method can first determine which videos contain the target person, and then perform target person detection and tracking on these videos. This facilitates accurate and automated determination of whether the target person appears in the scene video at different times and locations within a large number of videos, and further, it can determine the target person's movement trajectory, thus facilitating accurate and efficient finding of missing persons.
[0090] Alternatively, this method can be executed by electronic devices such as terminal devices or servers. Terminal devices can be user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistant (PDA) handheld devices, computing devices, in-vehicle devices, wearable devices, etc. This method can be implemented by a processor calling computer-readable instructions stored in memory. Alternatively, the method can be executed by a server, such as a local server or a cloud server.
[0091] In step S101, at least one object in the acquired image to be detected is detected, and a first feature of each object is extracted.
[0092] This involves using a pre-trained neural network to detect at least one object in an image. The detection result can be a bounding box, which is the smallest rectangle that can enclose the object. For example, if the object is a person in the image, a pre-trained person detection neural network can detect at least one person and present the detection result as the person's bounding box. The image can be input into a neural network such as a convolutional neural network, which can then output the location of the bounding box for at least one object in the image. The location of the bounding box can be represented by the coordinates of its top-left and bottom-right corners. In one possible example, every object in the image can be detected.
[0093] Each detected object can be feature extracted using image feature extraction algorithms commonly used in related technologies. The extracted first feature is used to characterize at least one dimension of the object's properties. For example, when the object is a person, the extracted first feature can be used to characterize the person's height, skin color, hairstyle, gender, age, and other properties.
[0094] In one possible embodiment, a neural network integrating object detection and feature extraction functions can be pre-trained. Then, in this step, this neural network is used to simultaneously detect at least one object in the image to be detected and extract the first feature of each object. That is, the neural network can have an object detection module and a feature extraction module, both of which can be constructed using convolutional neural networks. The image to be detected can then be input into the neural network. The neural network first inputs the image to be detected into the object detection module, which then inputs the detection boxes of at least one object into the feature extraction module. The feature extraction module extracts features from the objects within the detection boxes and outputs the first feature of each object within the detection boxes.
[0095] In step S102, the target three-dimensional model of the target object is obtained, and the second feature of the target three-dimensional model is extracted.
[0096] Among them, the target 3D model is the 3D model of the target object, that is, the 3D model constructed in advance based on the features of the target object. Compared with feature carriers such as 2D images and text descriptions, the 3D model can carry more comprehensive and specific features.
[0097] The target 3D model can be obtained from a model library. The model library can pre-store 3D models of multiple objects, and each 3D model has an identifier. For example, a 3D model of a person can be identified using their name, ID number, etc. Therefore, the target 3D model can be obtained from the model library based on the pre-inputted identifier of the target object and / or pre-extracted third features, wherein the third features are extracted from the pre-input data of the target object. If the user pre-inputs the identifier of the target object, such as the name or ID number of the target person, the corresponding 3D model can be found in the model library using this identifier and obtained as the target 3D model. If the user pre-inputs the data of the target object, such as an image or text description of the target person, the third features can be extracted from the target object data first, and then compared with the features of each 3D model in the model library. The 3D model whose comparison result meets the requirements (e.g., the error is within a preset error range) is obtained as the target 3D model.
[0098] A second feature of the target 3D model can be extracted using a pre-trained neural network. This second feature can characterize multiple dimensions of the target object, such as the target person's height, skin color, hairstyle, gender, and age. Optionally, the target 3D model can be input into the neural network, which can then output the second feature of the target 3D model.
[0099] In step S103, a target object among the at least one object is determined based on the second feature and the first feature of each of the objects, wherein the first feature of the target object matches the second feature.
[0100] This process involves comparing the second feature with the first feature of each object to determine if the first feature matches the second feature. Objects with a matching first feature are then identified as target objects. A match between the first and second features can mean that the error between them is within a pre-defined error range, or that their similarity is within a pre-defined similarity range. For example, a decision tree can be used to determine the similarity between the first and second features, and then objects with similarity within the target similarity range are identified as target objects.
[0101] Furthermore, when the similarity between the first feature of each object and the second feature of the target 3D model is not within the target similarity range, a suspected similarity range can be used. This range is wider than the target similarity range, and at least one suspected object can be identified through this range. That is, objects whose first and second features fall within this range are identified as suspected objects, and these suspected objects can then be presented to the user for identification. For example, if a target person undergoes cosmetic surgery, their first feature changes, but their skin color, height, and bone structure remain unchanged, thus maintaining a certain similarity to the second feature. Therefore, the identified target object is a suspected object, requiring further manual identification to determine if it is indeed the target person who has undergone partial changes. This allows for accurate detection and tracking of criminals undergoing cosmetic surgery or other alterations.
[0102] Furthermore, after identifying the target object, the target 3D model can be optimized based on the target object's first feature. This allows for further optimization of the target 3D model during the target object detection process, resulting in more comprehensive and specific features and improving the accuracy of detection using the target 3D model.
[0103] Furthermore, when the image to be detected is a frame image from the first video, such as a sampled image from the first video, after determining the target object, the target object can be detected in every frame of the first video except for the image to be detected, based on the first feature and / or the second feature of the target object. This means tracking the target object in the first video, thereby facilitating the determination of the target object's trajectory and range of motion. Moreover, after detecting the target object in every frame of the first video except for the image to be detected, the target 3D model can be optimized based on the first feature of the target object detected in each frame. This means optimizing the target 3D model using the first feature of the target object in each frame, thereby continuously increasing the comprehensiveness and specificity of the target 3D model, as well as the accuracy of target tracking, through iterative processing.
[0104] This embodiment detects at least one object in an image to be detected, extracts a first feature for each object, then obtains a target 3D model of the target object, and extracts a second feature from the target 3D model. Finally, based on the second feature and the first feature of each object, the target object among at least one object is determined; that is, the object whose first feature matches the second feature is identified as the target object. Through the target 3D model of the target object, relatively comprehensive features, namely the second feature, can be extracted. This allows for accurate identification of the target object from at least one object, even in situations with dense objects, severe occlusion, similar targets, or targets that have undergone partial changes, thus enabling accurate tracking of the target object. For example, if we want to track a person who has undergone plastic surgery, using detection methods in related technologies would likely result in the correct target being missed. However, using the detection method in this embodiment, the target person can be successfully detected using their unchanged skin color, height, skeletal structure, and other features. Therefore, the detection method in this embodiment is very helpful in finding missing elderly people and children and tracking criminals.
[0105] In some embodiments of this disclosure, before determining the target object among the at least one object based on the first feature of each of the objects and the second feature of the target 3D model, a decision tree may also be used to determine the first category information of each of the objects based on the first feature of each of the objects, and then to determine the object to be detected among the at least one object based on the pre-input second category information and / or the third category information in the target 3D model, wherein the first category information of the object to be detected matches the second category information and / or the third category information.
[0106] The first category information of an object can include the object's category in at least one dimension, such as the category of a person in terms of height, hairstyle, skin color, gender, and age. Users can input the second category information of the target object based on known information about the target object, or the target 3D model can carry the third category information of the target object. The second and third category information can be the category of the target object in at least one dimension, such as the category of a person in terms of height, hairstyle, skin color, gender, and age. Decision trees can efficiently determine the first category information of each object.
[0107] Matching the first category information with the second and / or third category information can be done by having the first category information be the same as the second and / or third category information. For example, if an object's first category information is 1.75m tall, short-haired, male of East Asian descent, and its second category information is also 1.75m tall, short-haired, male of East Asian descent, then the first category information matches the second category information.
[0108] Filtering by category information can eliminate at least some objects in a given set of objects. Furthermore, decision trees are highly efficient at determining the first category, and comparing category information is more efficient than comparing features.
[0109] Based on this, when determining the target object among the at least one objects based on the first feature of each object and the second feature of the target 3D model, the target object among the at least one object to be detected can be determined based on the first feature and the second feature of each object to be detected. That is, the target object is determined only among the objects to be detected, thereby avoiding comparison of the first and second features of other objects and improving detection efficiency.
[0110] Some embodiments of this disclosure also include, for example Figure 2 The model building process shown includes steps S201 to S203.
[0111] In step S201, multiple similar objects in at least one frame of the second video are detected, and a fourth feature of each of the similar objects is extracted.
[0112] The second video is the video to be detected and tracked. It can detect each frame of the second video, or multiple frames within the second video at preset intervals, such as starting from the first frame and detecting one frame every N frames, where N is greater than or equal to 1.
[0113] Multiple similar objects can be all or part of an image. When detecting the first frame of the second video, multiple objects in the first frame are detected, and each object is considered a similar object. When detecting non-first frame images in the second video, each similar object in the non-first frame is detected based on the fourth and fifth features of each similar object in the previous frame. That is, each object in the first frame is considered a similar object, while whether each object in a non-first frame is considered a similar object depends on the determination result of the previous frame. The fourth and fifth features are used to characterize at least one dimension of the object's properties. For example, when the object is a person, the extracted first feature can be used to characterize the person's height, skin color, hairstyle, gender, age, etc.
[0114] In step S202, after each frame of image is detected, a three-dimensional model of the similar object is constructed or optimized based on the fourth feature of each similar object, and the fifth feature of the three-dimensional model of each similar object is extracted. Based on the pre-extracted third feature and the fourth and / or fifth features of each similar object, it is determined whether each similar object is retained as a similar object.
[0115] The third feature is extracted from the pre-input target object data. The pre-input target object data can be key information for detecting the target object, such as an image or text description of the target person. The third feature is used to characterize at least one dimension of the object's properties. For example, if the object is a person, the extracted first feature can be used to characterize the person's height, skin color, hairstyle, gender, age, etc.
[0116] When detecting the first frame image in the second video, a three-dimensional model of the corresponding similar object is constructed based on the first feature of each of the similar objects; when detecting non-first frame images in the detected video, the three-dimensional model of the corresponding similar object is optimized based on the fourth feature of each of the similar objects.
[0117] When detecting the first frame image in the second video, based on the pre-extracted third feature and the fourth feature of each similar object, it is determined whether each similar object should be retained as a similar object. Optionally, if the error between the third feature and the fourth feature of a similar object is within a preset error range, the similar object is retained as a similar object; otherwise, the similar object is no longer retained as a similar object. When detecting images that are not the first frame image in the second video, based on the pre-extracted third feature and the fourth and fifth features of each similar object, it is determined whether each similar object should be retained as a similar object. Optionally, if the error between the third feature and the fourth feature of a similar object is within a preset error range, and the error between the third feature and the fifth feature of the similar object is within a preset error range, the similar object is retained as a similar object; otherwise, the similar object is no longer retained as a similar object.
[0118] In this step, each frame of image is used to optimize the 3D model of each similar object, making the 3D model of each similar object more comprehensive and specific. As a result, the features of each similar object become more comprehensive and specific, and the difference between the non-target object and the third feature becomes greater. Consequently, the number of similar objects decreases, meaning the interference of the target object decreases. Through repeated iterations, the 3D model of each object can be made more accurate, and the detection range can be continuously narrowed, gradually approaching the target object.
[0119] In step S203, if there is one remaining similar object and it is still considered a similar object after a preset number of frames are detected, the similar object is determined to be the target object, the three-dimensional model of the similar object is determined to be the target three-dimensional model, and the target three-dimensional model is saved to the model library.
[0120] Through repeated iterations, the number of similar objects can be gradually reduced until only one similar object remains. Then, the 3D model of the similar object can be optimized using multiple frames of images. The similar object can be verified using a more comprehensive and specific fifth feature. If the similar object is always retained as a similar object, then the similar object is the target object.
[0121] When saving the target 3D model, the pre-input target object identifier and data can be saved along with the 3D model, facilitating retrieval of the target 3D model when detecting target objects in other videos. Alternatively, a decision tree can be used to determine the third category information of each similar object based on its fourth feature, adding this third category information to the target 3D model before saving it to the model library. This allows for preliminary screening of multiple objects in other videos based on the target object's third category information, improving detection efficiency and reducing computational load. The decision tree is a tree-structured category graph, where each node represents a category. Each node can have parent and child nodes. The N highest-level categories of the fourth feature can be determined from the top of the decision tree. Further refinements of the fourth feature are then determined within the child nodes of each highest-level category until no further refinements can be determined. The determined category information is then used as the third category information of the fourth feature.
[0122] In this embodiment, similar objects are continuously detected in at least one frame of the first video. After each frame is detected, a 3D model of each similar object is constructed or optimized. The features of the 3D model are used to filter similar objects, thereby continuously reducing the number of similar objects while optimizing the 3D model of each similar object. This reduces the interference of the target object and gradually approaches the final detection result of the target object. Progressive target detection is performed in the iterative process, and multiple 3D models, including the target 3D model, are constructed simultaneously, providing an accurate feature carrier for subsequent detection of the target object.
[0123] The model building process shown in steps S201 to S203 will be explained below with a specific example. In this example, there are five people in the second video, and these five people appear continuously in the second video. The user provides photos of the target people. First, the third feature of the target person is extracted from the photo. Then, the first frame of the second video is detected, and five people are detected. The fourth feature of each person is extracted, and a 3D model of the corresponding person is constructed based on the fourth feature. The fifth feature of each person's 3D model is then extracted. Based on the similarity between the third feature and the fourth feature of each person, and the similarity between the third feature and the fifth feature of each person, people with similarity scores higher than a preset threshold are retained as similar people. For example, three out of five people are retained as similar people. Next, other frames of the second video are detected, and three similar people are detected from the frames based on the fourth and fifth features of the three similar people. The fourth feature of each detected similar person is extracted, and the 3D model of that similar person is optimized using the fourth feature. The fifth feature of the optimized 3D model of each similar person is then extracted. The remaining similar people are then filtered in the same way as those in the first frame. If only one similar person remains after multiple frames, that similar person can be identified as the target person, and the 3D model of that target person is the target 3D model. The target 3D model can then be saved, completing the model construction.
[0124] The 3D model constructed in this embodiment can carry more comprehensive and specific features, thus solving the problems of insufficient detection and recognition and slow detection speed in situations such as occlusion, small scale, and local differences. This improves the accuracy of occlusion and small object detection, and enables accurate identification and tracking of targets even in the presence of local differences.
[0125] According to a second aspect of the embodiments of this disclosure, an image detection apparatus is provided; please refer to the appendix. Figure 3 ,include:
[0126] The detection module 301 is used to detect at least one object in the acquired image to be detected and extract a first feature of each object;
[0127] Model module 302 is used to acquire the target three-dimensional model of the target object and extract the second feature of the target three-dimensional model;
[0128] The determining module 303 is configured to determine a target object among the at least one object based on the second feature and the first feature of each of the objects, wherein the first feature of the target object matches the second feature.
[0129] In some embodiments of this disclosure, the image to be detected is a frame image in a first video;
[0130] The image detection device further includes a tracking module for:
[0131] After determining the target object among the at least one objects, the target object is detected in each frame of the first video, excluding the image to be detected, based on the first feature and / or the second feature of the target object.
[0132] In some embodiments of this disclosure, the image detection apparatus further includes an optimization module for:
[0133] After determining the target object among the at least one object, the target 3D model is optimized based on a first feature of the target object; and / or,
[0134] After detecting the target object in each frame of the first video except for the image to be detected, the target 3D model is optimized based on the first feature of the target object detected in each frame.
[0135] In some embodiments of this disclosure, the image detection device further includes a screening module for:
[0136] Before determining the target object among the at least one object based on the second feature and the first feature of each object, a decision tree is used to determine the first category information of each object based on the first feature of each object;
[0137] Based on the pre-input second category information and / or the third category information in the target 3D model, determine the object to be detected among the at least one object, wherein the first category information of the object to be detected matches the second category information and / or the third category information;
[0138] The determining module is specifically used for:
[0139] Based on the second feature and the first feature of each of the objects to be detected, a target object is determined among the at least one objects to be detected.
[0140] In some embodiments of this disclosure, when the model module is used to obtain the target 3D model of the target object, it is specifically used for:
[0141] The target 3D model is obtained from the model library based on the pre-input target object identifier and / or pre-extracted third features, wherein the third features are extracted from the pre-input target object data.
[0142] In some embodiments of this disclosure, the image detection apparatus further includes a modeling module for:
[0143] Detect multiple similar objects in at least one frame of the second video, and extract a fourth feature for each of the similar objects;
[0144] After each frame of image is detected, a 3D model of the corresponding similar object is constructed or optimized based on the fourth feature of each similar object, and the fifth feature of the 3D model of each similar object is extracted. Based on the pre-extracted third feature and the fourth and / or fifth features of each similar object, it is determined whether each similar object is retained as a similar object.
[0145] If one similar object remains and is still considered a similar object after a preset number of frames are detected, the similar object is determined to be the target object, and the 3D model of the similar object is determined to be the target 3D model, and the target 3D model is saved to the model library.
[0146] In some embodiments of this disclosure, when the modeling module is used to detect multiple similar objects in at least one frame of the second video, it is specifically used for:
[0147] In the case of detecting the first frame image in the second video, multiple objects in the first frame image are detected, and each of the objects is regarded as a similar object;
[0148] In the case of detecting non-first frame images in the second video, each similar object in the non-first frame image is detected based on the fourth and fifth features of each similar object in the previous frame image.
[0149] In some embodiments of this disclosure, when the modeling module is used to construct or optimize a 3D model of a corresponding similar object based on the fourth feature of each similar object, it is specifically used for:
[0150] In the case of detecting the first frame image in the second video, a three-dimensional model of the corresponding similar object is constructed based on the fourth feature of each of the similar objects;
[0151] In the case of detecting non-first frame images in the second video, the 3D model of the corresponding similar object is optimized based on the fourth feature of each similar object.
[0152] In some embodiments of this disclosure, when the modeling module determines whether each similar object should be retained as a similar object based on a pre-extracted third feature and a fourth and / or fifth feature of each similar object, it is specifically used for:
[0153] In the case of detecting the first frame image in the second video, based on the pre-extracted third feature and the fourth feature of each of the similar objects, it is determined whether each of the similar objects is retained as a similar object;
[0154] In the case of detecting non-first frame images in the second video, it is determined whether each of the similar objects should be retained as a similar object based on the pre-extracted third feature and the fourth and fifth features of each of the similar objects.
[0155] In some embodiments of this disclosure, the image detection device further includes a classification module for:
[0156] Using a decision tree, the third category information of the similar objects is determined based on the fourth feature of each similar object;
[0157] When the modeling module saves the target 3D model to the model library, it is specifically used for:
[0158] The third category information of the target object is added to the target 3D model, and the target 3D model is saved to the model library.
[0159] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments of the method in the first aspect, and will not be elaborated upon here.
[0160] According to a third aspect of the embodiments of this disclosure, please refer to the appendix. Figure 4 The diagram illustrates, for example, a block diagram of an electronic device. For instance, device 400 could be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0161] Reference Figure 4 The device 400 may include one or more of the following components: a processing component 402, a memory 404, a power supply component 406, a multimedia component 408, an audio component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.
[0162] Processing component 402 typically controls the overall operation of device 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.
[0163] Memory 404 is configured to store various types of data to support the operation of device 400. Examples of this data include instructions for any application or method operating on device 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0164] The power supply component 406 provides power to the various components of the device 400. The power supply component 406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 400.
[0165] Multimedia component 408 includes a screen that provides an output interface between the device 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, swipe, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the device 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0166] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when device 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.
[0167] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0168] Sensor assembly 414 includes one or more sensors for providing status assessments of various aspects of device 400. For example, sensor assembly 414 can detect the on / off state of device 400, the relative positioning of components such as the display and keypad of device 400, image detection of changes in the position of device 400 or a component of device 400, the presence or absence of user contact with device 400, orientation or acceleration / deceleration of device 400, and temperature changes of device 400. Sensor assembly 414 may also include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0169] Communication component 416 is configured to facilitate wired or wireless communication between device 400 and other devices. Device 400 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0170] In an exemplary embodiment, device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the image detection method of the aforementioned electronic device.
[0171] Fourthly, in exemplary embodiments, this disclosure also provides a non-transitory computer-readable storage medium including instructions, such as a memory 404 including instructions, which can be executed by a processor 420 of device 400 to complete the image detection method of the electronic device described above. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0172] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0173] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image detection method, characterized in that, include: Detect at least one object in the acquired image to be detected, and extract a first feature for each object, wherein the image to be detected is the current image frame within the video to be detected; Obtain a target 3D model of the target object and extract a second feature of the target 3D model, wherein the target 3D model is constructed based on the target object in at least one historical image frame within the video to be detected; Using a decision tree, the first category information of each object is determined based on the first feature of each object; Based on the pre-input second category information and / or the third category information in the target 3D model, determine the object to be detected among the at least one object, wherein the first category information of the object to be detected matches the second category information and / or the third category information; Based on the second feature and the first feature of each of the objects to be detected, a target object is determined among the at least one objects to be detected, wherein the first feature of the target object matches the second feature.
2. The image detection method according to claim 1, characterized in that, The image to be detected is a frame image from the first video; After determining the target object among the at least one object, the method further includes: Based on the first feature and / or the second feature of the target object, the target object is detected in every frame of the first video except for the image to be detected.
3. The image detection method according to claim 2, characterized in that, After determining the target object among the at least one object, the method further includes: Based on the first feature of the target object, the target 3D model is optimized; and / or, After detecting the target object in each frame of the first video except for the image to be detected, the method further includes: The target 3D model is optimized based on the first feature of the target object detected in each frame of the image.
4. The image detection method according to claim 1, characterized in that, The process of obtaining the target 3D model of the target object includes: The target 3D model is obtained from the model library based on the pre-input target object identifier and / or pre-extracted third features, wherein the third features are extracted from the pre-input target object data.
5. The image detection method according to claim 4, characterized in that, Also includes: Detect multiple similar objects in at least one frame of the second video, and extract a fourth feature for each of the similar objects; After each frame of image is detected, a 3D model of the corresponding similar object is constructed or optimized based on the fourth feature of each similar object, and the fifth feature of the 3D model of each similar object is extracted. Based on the pre-extracted third feature and the fourth and / or fifth features of each similar object, it is determined whether each similar object is retained as a similar object. If one similar object remains and is still considered a similar object after a preset number of frames are detected, the similar object is determined to be the target object, and the 3D model of the similar object is determined to be the target 3D model, and the target 3D model is saved to the model library.
6. The image detection method according to claim 5, characterized in that, The detection of multiple similar objects in at least one frame of the second video includes: In the case of detecting the first frame image in the second video, multiple objects in the first frame image are detected, and each of the objects is regarded as a similar object; In the case of detecting non-first frame images in the second video, each similar object in the non-first frame image is detected based on the fourth and fifth features of each similar object in the previous frame image.
7. The image detection method according to claim 5, characterized in that, The step of constructing or optimizing the 3D model of the corresponding similar object based on the fourth feature of each similar object includes: In the case of detecting the first frame image in the second video, a three-dimensional model of the corresponding similar object is constructed based on the fourth feature of each of the similar objects; In the case of detecting non-first frame images in the second video, the 3D model of the corresponding similar object is optimized based on the fourth feature of each similar object.
8. The image detection method according to claim 5, characterized in that, The step of determining whether each similar object should be retained as a similar object based on the pre-extracted third feature and the fourth and / or fifth features of each similar object includes: In the case of detecting the first frame image in the second video, based on the pre-extracted third feature and the fourth feature of each of the similar objects, it is determined whether each of the similar objects is retained as a similar object; In the case of detecting non-first frame images in the second video, it is determined whether each of the similar objects should be retained as a similar object based on the pre-extracted third feature and the fourth and fifth features of each of the similar objects.
9. The image detection method according to claim 5, characterized in that, Also includes: Using a decision tree, the third category information of the similar objects is determined based on the fourth feature of each similar object; Saving the target 3D model to the model library includes: The third category information of the target object is added to the target 3D model, and the target 3D model is saved to the model library.
10. An image detection device, characterized in that, include: A detection module is used to detect at least one object in the acquired image to be detected, and to extract a first feature of each object; The model module is used to acquire the target 3D model of the target object and extract the second feature of the target 3D model; A filtering module is used to determine first category information of each of the objects based on a first feature of each object using a decision tree; and to determine the object to be detected among the at least one object based on a second category information pre-input and / or a third category information in the target 3D model, wherein the first category information of the object to be detected matches the second category information and / or the third category information; A determining module is configured to determine a target object among the at least one target object based on the second feature and a first feature of each target object, wherein the first feature of the target object matches the second feature.
11. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is used to store computer instructions that can be executed on the processor. The processor is used to execute the computer instructions based on the image detection method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Class-level 6D pose and size estimation method and device
CN113012122A