An object registration method and apparatus
By extracting and registering feature information of multiple input images, the problem of increasing training time and decreasing recognition effect when new objects are registered in the prior art is solved, and fast and accurate object registration and detection performance are achieved.
Patent Information
- Application Number
- CN202011607387.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-12-29
AI Technical Summary
When the prior art adds new objects to be identified in the multi-object pose estimation network, the training time increases linearly, affecting the recognition effect of registered objects, resulting in a decrease in detection accuracy and success rate.
Object registration is performed by extracting feature information of multiple input images, including descriptors of local feature points and descriptors of global feature. The method includes obtaining the real image of the object and the synthetic image that can be slightly rendered under multiple poses, extracting feature information, and registering it corresponding to the identification of the object.
The registration time of object is shortened and does not affect the recognition performance of other registered objects. Even if multiple new object registration is added, the registration time will not be significantly increased, ensuring the detection performance of registered objects.
Smart Images

Figure CN114758334B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer vision, and in particular, to an object registration method and apparatus. Background Art
[0002] Computer vision is an integral part of various intelligent / autonomous systems in various application fields, such as manufacturing, inspection, document analysis, medical diagnosis, and military. It is a discipline about how to use cameras / video cameras and computers to obtain the data and information of the objects to be photographed that we need. Figuratively speaking, it is to install eyes (cameras / video cameras) and a brain (algorithms) on the computer to replace the human eye to identify, track, and measure targets, etc., so that the computer can perceive the environment. Since perception can be regarded as extracting information from sensory signals, computer vision can also be regarded as the science of studying how to make artificial systems "perceive" from images or multi-dimensional data. Generally speaking, computer vision uses various imaging systems to replace the visual organs to obtain input information, and then uses the computer to replace the brain to process and interpret these input information. The ultimate research goal of computer vision is to enable the computer to observe and understand the world through vision like a human being and have the ability to adapt to the environment independently.
[0003] The pose detection and tracking of objects (people or things) is a key technology in the field of computer vision, which can endow machines with the ability to perceive the three-dimensional spatial position and semantics of objects in the real environment and has a wide range of applications in fields such as robotics, autonomous driving, and augmented reality.
[0004] In practical applications, usually a multi-object pose estimation network is first constructed to identify the poses of recognizable objects from the input images; then, using the method of machine learning, with the three-dimensional (3D) network models of multiple objects to be recognized provided by the user, the multi-object pose estimation network is trained, and the objects to be recognized are registered into the multi-object pose estimation network to enable the multi-object pose estimation network to recognize the already registered objects. When applied online, the picture is input into the multi-object pose estimation network, and the multi-object pose estimation network identifies the poses of the objects in the picture.
[0005] When it is necessary to add the registration of an object to be recognized in the multi-object pose estimation network, the user provides a three-dimensional model of the newly added object to be recognized, and the multi-object pose estimation network is retrained using the newly added three-dimensional model and the original three-dimensional model together, resulting in a linear increase in the training time and affecting the recognition effect of the multi-object pose estimation network on the already trained recognizable objects, resulting in a decrease in the detection accuracy and success rate. Summary of the Invention
[0006] The object registration method and device provided by this application solve the problem of how to improve the accuracy of the pose of a terminal device.
[0007] To achieve the above object, this application adopts the following technical solutions:
[0008] In a first aspect, an object registration method is provided. The method may include: obtaining a plurality of first input images including a first object, where the plurality of first input images include a real image of the first object and / or a plurality of first synthesized images of the first object; the plurality of first synthesized images are obtained by differentiable rendering of a three-dimensional model of the first object in a plurality of first poses; the plurality of first poses are different; respectively extracting feature information of the plurality of first input images, where the feature information is used to indicate the features of the first object in the first input image where it is located; corresponding the feature information extracted from each first input image with the identifier of the first object to perform registration of the first object.
[0009] Through the object registration method provided by this application, feature information of an object in an image is extracted for object registration. The registration time is short, and it does not affect the recognition performance of other already registered objects. Even if a large number of new objects are registered, the registration time will not increase significantly, and the detection performance of the already registered objects is also guaranteed.
[0010] In a possible implementation manner, the above feature information may include descriptors of local feature points and descriptors of global features.
[0011] In another possible implementation manner, respectively extracting feature information of the plurality of first input images may be specifically implemented as: respectively inputting the plurality of first input images into a first network for pose recognition to obtain the pose of the first object in each first input image; according to the obtained pose of the first object, projecting the three-dimensional model of the first object onto each first input image respectively to obtain a projection area in each first input image; respectively extracting feature information in the projection area of each first input image. Wherein, the first network is used to recognize the pose of the first object in the image. By the pose of the first object in the image, the area of the first object in the image is determined, and feature information is extracted in this area, which improves the efficiency and accuracy of feature information extraction.
[0012] In another possible implementation manner, respectively extracting feature information of the plurality of first input images may be specifically implemented as: respectively inputting the plurality of first input images into a first network for black-and-white image extraction to obtain the black-and-white image of the first object in each first input image; respectively extracting feature information in the black-and-white image of the first object in each first input image. Wherein, the first network is used to extract the black-and-white image of the first object in the image. The black-and-white image of the first object in the image is determined as the area of the first object in the image, and feature information is extracted in this area, which improves the efficiency and accuracy of feature information extraction.
[0013] In another possible implementation, the object registration method provided in this application may further include a process of optimizing the differentiable renderer, and this process may include: inputting N real images of a first object into a first network respectively to obtain the second poses of the first object in each of the real images output by the first network; N is greater than or equal to 1; the first network is used to identify the pose of the first object in the image; according to the three-dimensional model of the first object, using the differentiable renderer, N second synthesized images are differentiably rendered at each second pose; obtaining that a real image of a second pose corresponds to the second synthesized image rendered at the second pose; respectively intercepting the regions at the same positions of the first object in the second synthesized image corresponding to the real object in each real image as the foreground images of each real image; constructing a first loss function according to the first difference information between the foreground images of the N real images and their corresponding second synthesized images; wherein, the first difference information is used to indicate the difference between the foreground image and its corresponding second synthesized image; updating the differentiable renderer according to the first loss function so that the synthesized image output by the differentiable renderer approximates the real image of the object. By optimizing the differentiable renderer, the rendering authenticity of the differentiable renderer is improved, and the difference between the synthesized image obtained by differentiable rendering and the real image is reduced.
[0014] In another possible implementation, the first difference information may include one or more of the following information: feature map difference, pixel color difference, difference between extracted feature descriptors.
[0015] In another possible implementation, the first loss function may be the sum of the calculated values of the multiple first difference information between the foreground images of the N real images and their corresponding second synthesized images.
[0016] In another possible implementation, the object registration method provided by this application may further include a method for training an object pose detection network, which may specifically include: obtaining a plurality of second input images including a first object, where the plurality of second input images include real images of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from a three-dimensional model of the first object in a plurality of third poses; the plurality of third poses are different; inputting the plurality of second input images into a second network for pose recognition respectively to obtain the fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image; obtaining a fourth synthesized image of the first object in each fourth pose through differentiable rendering according to the three-dimensional model of the first object; it is obtained that a second input image of a fourth pose corresponds to a fourth synthesized image rendered in the fourth pose; constructing a second loss function according to the second difference information between each fourth synthesized image and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image; updating the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is smaller than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image. By training the object pose detection network, the accuracy of the pose recognition network is improved, and the difference between the output of the pose recognition network and the real pose of the object in the image is reduced.
[0017] In another possible implementation, the second loss function Loss 2 may satisfy the following expression: X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image. This implementation provides an expression of a specific second loss function, enabling the training of the object pose detection network and improving the accuracy of the pose detection network.
[0018] In another possible implementation, the second difference information may include one or more of the following: the difference between the intersection of union (IOU) of the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose of the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image. This implementation provides possible realizations of the second difference information and enriches the content of the second difference information.
[0019] In a second aspect, a display method is provided, which may include: obtaining a first image; if the first image includes one or more recognizable objects, outputting first information for prompting that recognizable objects are detected in the first image; obtaining the poses of each recognizable object in the first image through a pose detection network corresponding to each recognizable object included in the first image; displaying virtual content corresponding to each recognizable object according to the pose of each recognizable object; if no recognizable object is included in the first image, outputting second information for prompting that no recognizable object is detected, and adjusting the viewing angle to obtain a second image, where the second image is different from the first image.
[0020] Through the display method provided in this application, a prompt is output to the user as to whether the image includes recognizable objects, enabling the user to intuitively obtain whether the image includes recognizable objects and improving the user experience.
[0021] In a possible implementation, the display method provided in this application may further include: extracting feature information from the first image, where the feature information is used to indicate recognizable features in the first image; determining whether there is feature information in the feature library whose matching distance with the extracted feature information meets a preset condition; where one or more pieces of feature information of different objects are stored in the feature library; if there is feature information in the feature library whose matching distance with the feature information meets the preset condition, determining that the first image includes one or more recognizable objects; if there is no feature information in the feature library whose matching distance with the feature information meets the preset condition, determining that the first image does not include any recognizable objects. By comparing the feature information of the image with the feature library, it is possible to simply and quickly determine whether the image includes recognizable objects.
[0022] In a possible implementation, the preset condition may include being less than or equal to a preset threshold.
[0023] In another possible implementation manner, the display method provided by this application may further include: obtaining one or more first local feature points in a first image, where the matching distance between the descriptor of the first local feature point and the descriptor of the local feature point in the feature library is less than or equal to a first threshold; the feature library stores descriptors of local feature points of different objects; determining one or more regions of interest (ROIs) in the first image according to the first local feature points; one ROI includes one object; extracting global features in each ROI; if there are one or more first global features in the global features in each ROI, determining that the first image includes an identifiable object corresponding to the first global feature; where the matching distance between the descriptor of the first global feature and the descriptor of the global feature in the feature library is less than or equal to a second threshold; the feature library also stores descriptors of global features of different objects; if there are no first global features in the global features in each ROI, determining that the first image does not include any identifiable object. By first comparing the local feature points of the image with the feature library to determine the ROI region, and then extracting the global features in the ROI region and comparing them with the feature library, the efficiency of determining whether the image includes an identifiable object can be improved, and the accuracy of determining whether the image includes an identifiable object can also be improved.
[0024] In another possible implementation manner, the display method provided by this application may also be used in combination with the object registration method provided in the foregoing first aspect. The display method provided by this application may further include: obtaining a plurality of first input images including a first object, where the plurality of first input images include a real image of the first object and / or a plurality of first synthesized images of the first object; the plurality of first synthesized images are obtained by differentiable rendering of the three-dimensional model of the first object in a plurality of first poses; the plurality of first poses are different; respectively extracting feature information of the plurality of first input images, where the feature information is used to indicate the features of the first object in the first input image where it is located; storing the feature information extracted from each first input image corresponding to the identifier of the first object in the feature library for registration of the first object. By extracting the feature information of the object in the image for object registration, the registration time is short, and it does not affect the recognition performance of other registered objects. Even if a lot of new objects are registered, the registration time will not increase too much, and the detection performance of the registered objects is also guaranteed.
[0025] In another possible implementation manner, the above feature information may include a descriptor of a local feature point and a descriptor of a global feature.
[0026] It should be noted that the specific implementation of the object registration method may refer to the foregoing first aspect and will not be elaborated here.
[0027] In another possible implementation, the display method provided in this application may further include a process of optimizing a differentiable renderer, and this process may include: inputting N real images of a first object into a first network respectively to obtain the second pose of the first object in each of the real images output by the first network; N is greater than or equal to 1; the first network is used to identify the pose of the first object in the image; according to the three-dimensional model of the first object, using the differentiable renderer to perform differentiable rendering to obtain N second synthesized images at each second pose; obtaining that a real image of one second pose corresponds to the second synthesized image rendered at this one second pose; respectively intercepting the regions at the same positions of the first object in the second synthesized image corresponding to the real object in each real image as the foreground images of each real image; constructing a first loss function according to the first difference information between the foreground images of the N real images and their corresponding second synthesized images; wherein the first difference information is used to indicate the difference between the foreground image and its corresponding second synthesized image; updating the differentiable renderer according to the first loss function so that the synthesized image output by the differentiable renderer approaches the real image of the object. By optimizing the differentiable renderer, the rendering authenticity of the differentiable renderer is improved, and the difference between the synthesized image obtained by differentiable rendering and the real image is reduced.
[0028] In another possible implementation, the first difference information may include one or more of the following information: feature map difference, pixel color difference, difference between extracted feature descriptors.
[0029] In another possible implementation, the first loss function may be the sum of the calculated values of the multiple first difference information between the foreground images of the N real images and their corresponding second synthesized images.
[0030] In another possible implementation, the object registration method provided by this application may further include a method for training an object pose detection network, which may specifically include: obtaining a plurality of second input images including a first object, where the plurality of second input images include real images of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from a three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different; inputting the plurality of second input images into a second network for pose recognition respectively to obtain the fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image; obtaining a fourth synthesized image of the first object at each fourth pose through differentiable rendering according to the three-dimensional model of the first object; obtaining that a second input image at a fourth pose corresponds to a fourth synthesized image rendered at the fourth pose; constructing a second loss function according to the second difference information between each fourth synthesized image and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image; updating the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is less than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image. By training the pose detection network, the accuracy of the pose detection network recognition is improved, and the difference between the output of the pose detection network and the real pose of the object in the image is reduced.
[0031] In another possible implementation, the second loss function Loss 2 may satisfy the following expression: X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image. This implementation provides an expression of a specific second loss function to optimize the pose detection network and improve the accuracy of the pose detection network recognition.
[0032] In another possible implementation, the second difference information may include one or more of the following: the difference in IOU between the black-and-white image of the first object in the fourth synthesized image and the black-and-white image of the first object in its corresponding second input image, the difference between the fourth pose of the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image. This implementation provides possible realizations of the second difference information and enriches the content of the second difference information.
[0033] In a third aspect, the present application provides a method for training an object pose detection network. The method may specifically include: obtaining a plurality of second input images including a first object, where the plurality of second input images include a real image of the first object and / or a plurality of third synthetic images of the first object; the plurality of third synthetic images are rendered from a three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different; inputting the plurality of second input images into a second network for pose recognition respectively to obtain a fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image; obtaining a fourth synthetic image of the first object at each fourth pose by differentiable rendering according to the three-dimensional model of the first object; there is a correspondence between a second input image of a fourth pose and a fourth synthetic image rendered at the fourth pose; constructing a second loss function according to second difference information between each fourth synthetic image and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthetic image and its corresponding second input image; updating the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is less than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image.
[0034] By training the pose detection network, the accuracy of the pose detection network recognition is improved, and the difference between the output of the pose detection network and the real pose of the object in the image is reduced.
[0035] It should be noted that for the specific implementation of the third aspect, reference may be made to the specific implementation of training the object pose detection network described in the foregoing first aspect, and the same beneficial effects can be achieved, which will not be elaborated here.
[0036] In a fourth aspect, an object registration device is provided. The device includes: a first acquisition unit, an extraction unit, and a registration unit. Among them:
[0037] The first acquisition unit is configured to obtain a plurality of first input images including a first object, where the plurality of first input images include a real image of the first object and / or a plurality of first synthetic images of the first object; the plurality of first synthetic images are obtained by differentiable rendering of the three-dimensional model of the first object at a plurality of first poses; the plurality of first poses are different.
[0038] The extraction unit is configured to extract feature information of the plurality of first input images respectively, and the feature information is used to indicate the features of the first object in the first input image where it is located.
[0039] The registration unit is configured to correspond the feature information extracted by the extraction unit in each first input image with the identifier of the first object to perform registration of the first object.
[0040] Through the object registration device provided by this application, the feature information of the object in the image is extracted for object registration. The registration time is short, and it does not affect the recognition performance of other registered objects. Even if many new objects are registered, the registration time will not increase too much, and the detection performance of the registered objects is also guaranteed.
[0041] In a possible implementation, the feature information includes the descriptors of local feature points and the descriptors of global features.
[0042] In another possible implementation, the extraction unit can specifically be used to: input multiple first input images into the first network respectively for pose recognition to obtain the pose of the first object in each first input image; project the three-dimensional model of the first object onto each first input image according to the obtained pose of the first object to obtain the projection area in each first input image; extract feature information in the projection area in each first input image respectively. Among them, the first network is used to recognize the pose of the first object in the image. By the pose of the first object in the image, the area of the first object in the image is determined, and feature information is extracted in this area, which improves the efficiency and accuracy of feature information extraction.
[0043] In another possible implementation, the extraction unit can specifically be used to: input multiple first input images into the first network respectively for black-and-white image extraction to obtain the black-and-white image of the first object in each first input image; extract feature information in the black-and-white image of the first object in each first input image respectively. Among them, the first network is used to extract the black-and-white image of the first object in the image. The black-and-white image of the first object in the image is determined as the area of the first object in the image, and feature information is extracted in this area, which improves the efficiency and accuracy of feature information extraction.
[0044] In another possible implementation, the device may further include: a processing unit, a differentiable renderer, a screenshot unit, a construction unit, and an update unit. Among them:
[0045] The processing unit is used to input N real images of the first object into the first network respectively to obtain the second pose of the first object in each real image output by the first network. N is greater than or equal to 1. The first network is used to recognize the pose of the first object in the image.
[0046] The differentiable renderer is used to, according to the three-dimensional model of the first object, perform differentiable rendering to obtain N second synthesized images at each second pose; obtain that one real image of a second pose corresponds to one second synthesized image rendered at the second pose.
[0047] The interception unit is used to intercept the area at the same position of the first object in the second synthesized image corresponding to the real object in each real image respectively as the foreground image of each real image.
[0048] A construction unit is configured to construct a first loss function according to first difference information between a foreground image of N real images and a corresponding second synthesized image thereof. The first difference information is used to indicate the difference between the foreground image and the corresponding second synthesized image.
[0049] An update unit is configured to update the differentiable renderer according to the first loss function, so that the synthesized image output by the differentiable renderer approximates the real image of the object.
[0050] By optimizing the differentiable renderer, the rendering authenticity of the differentiable renderer is improved, and the difference between the synthesized image obtained by differentiable rendering and the real image is reduced.
[0051] In another possible implementation manner, the first difference information may include one or more of the following information: feature map difference, pixel color difference, difference between extracted feature descriptors.
[0052] In another possible implementation manner, the first loss function may be the sum of calculated values of multiple first difference information between the foreground images of N real images and their corresponding second synthesized images.
[0053] In another possible implementation manner, the apparatus may further include: a second acquisition unit, a processing unit, a differentiable renderer, a construction unit, and an update unit. Wherein:
[0054] The second acquisition unit is configured to acquire multiple second input images including a first object, and the multiple second input images include a real image of the first object and / or multiple third synthesized images of the first object; the multiple third synthesized images are rendered from a three-dimensional model of the first object at multiple third poses. The multiple third poses are different.
[0055] The processing unit is configured to input the multiple second input images into a second network for pose recognition respectively, and obtain a fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image.
[0056] The differentiable renderer is configured to perform differentiable rendering on the three-dimensional model of the first object to obtain a fourth synthesized image of the first object at each fourth pose. The second input image of one fourth pose corresponds to the fourth synthesized image rendered at one fourth pose.
[0057] The construction unit is configured to construct a second loss function according to second difference information between each fourth synthesized image and its corresponding second input image. The second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image.
[0058] An update unit is configured to update a second network according to a second loss function to obtain a first network. The difference between the pose of a first object in an image recognized by the first network and the true pose of the first object in the image is smaller than the difference between the pose of the first object in the image recognized by the second network and the true pose of the first object in the image.
[0059] By training the pose detection network, the accuracy of the pose detection network in recognition is improved, and the difference between the output of the pose detection network and the true pose of the object in the image is reduced.
[0060] In another possible implementation, the second loss function Loss 2 may satisfy the following expression: X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image. This implementation provides a specific expression of the second loss function to optimize the pose detection network and improve the accuracy of the pose detection network in recognition.
[0061] In another possible implementation, the second difference information may include one or more of the following: the difference in IOU between the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose of the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image. This implementation provides possible realizations of the second difference information and enriches the content of the second difference information.
[0062] It should be noted that the object registration device provided in the fourth aspect is used to implement the object registration method provided in the first aspect above. Its specific implementation may refer to the specific implementation of the first aspect above and will not be elaborated here.
[0063] In a fifth aspect, a display device is provided. The device includes a first acquisition unit, an output unit, and a processing unit; wherein:
[0064] The first acquisition unit is configured to acquire a first image.
[0065] The output unit is configured to, if the first image includes one or more recognizable objects, output first information, where the first information is used to prompt that recognizable objects are detected in the first image. If the first image does not include any recognizable objects, output second information, where the second information is used to prompt that no recognizable objects are detected, and adjust the viewing angle so that the first acquisition unit acquires a second image, where the second image is different from the first image.
[0066] A processing unit, configured to, if one or more recognizable objects are included in the first image, obtain the poses of each recognizable object in the first image through a pose detection network corresponding to each recognizable object included in the first image; and display virtual content corresponding to each recognizable object according to the pose of each recognizable object.
[0067] Through the display device provided by this application, a prompt is output to the user as to whether a recognizable object is included in the image, enabling the user to intuitively obtain whether a recognizable object is included in the image and improving the user experience.
[0068] In a possible implementation manner, the device may further include: an extraction unit, a determination unit, and a first determination unit. Wherein:
[0069] The extraction unit is configured to extract feature information in the first image, and the feature information is used to indicate recognizable features in the first image.
[0070] The determination unit is configured to determine whether there is feature information in the feature library whose matching distance with the feature information meets a preset condition. Wherein, one or more pieces of feature information of different objects are stored in the feature library.
[0071] The first determination unit is configured to, if there is feature information in the feature library whose matching distance with the feature information meets a preset condition, determine that one or more recognizable objects are included in the first image; if there is no feature information in the feature library whose matching distance with the feature information meets a preset condition, determine that no recognizable object is included in the first image.
[0072] By comparing the feature information of the image with the feature library, it is possible to simply and quickly determine whether a recognizable object is included in the image.
[0073] In another possible implementation manner, the preset condition may include being less than or equal to a preset threshold.
[0074] In another possible implementation manner, the device may further include: a second acquisition unit, a second determination unit, and a first determination unit. Wherein:
[0075] The second acquisition unit is configured to acquire one or more first local feature points in the first image, and the matching distance between the descriptor of the first local feature point and the descriptor of the local feature point in the feature library is less than or equal to a first threshold; the descriptors of the local feature points of different objects are stored in the feature library.
[0076] The second determination unit is configured to determine one or more ROIs in the first image according to the first local feature points; one object is included in one ROI.
[0077] Correspondingly, the extraction unit may further be configured to extract global features in each ROI.
[0078] A first determination unit, configured to determine that the first image includes recognizable objects corresponding to the first global features if there are one or more first global features in the global features of each ROI; and determine that the first image does not include any recognizable objects if there are no first global features in the global features of each ROI. Wherein, the matching distance between the descriptor of the first global feature and the descriptor of the global feature in the feature library is less than or equal to a second threshold. The feature library also stores descriptors of global features of different objects.
[0079] First, compare the local feature points of the image with the feature library to determine the ROI region, and then extract the global features in the ROI region and compare them with the feature library, which can improve the efficiency of determining whether the image includes recognizable objects and also improve the accuracy of determining whether the image includes recognizable objects.
[0080] In another possible implementation, the apparatus may further include: a third acquisition unit, an extraction unit, and a registration unit. Wherein:
[0081] The third acquisition unit is configured to acquire a plurality of first input images including a first object, and the plurality of first input images include a real image of the first object and / or a plurality of first synthetic images of the first object; the plurality of first synthetic images are obtained by differentiable rendering of the three-dimensional model of the first object in a plurality of first poses; and the plurality of first poses are different.
[0082] The extraction unit is configured to extract the feature information of the plurality of first input images respectively, and the feature information is used to indicate the features of the first object in the first input image where it is located.
[0083] The registration unit is configured to store the feature information extracted from each first input image corresponding to the identifier of the first object in the feature library for registration of the first object.
[0084] By extracting the feature information of the objects in the image for object registration, the registration time is short, and it does not affect the recognition performance of other registered objects. Even if a lot of new objects are registered, it will not greatly increase the registration time, and it also ensures the detection performance of the registered objects.
[0085] In another possible implementation, the above feature information may include descriptors of local feature points and descriptors of global features.
[0086] In another possible implementation, the apparatus may further include: a fourth acquisition unit, a processing unit, a differentiable renderer, a construction unit, and an update unit. Wherein:
[0087] A fourth acquisition unit, configured to acquire a plurality of second input images including a first object, where the plurality of second input images include a real image of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from a three-dimensional model of the first object in a plurality of third poses; the plurality of third poses are different.
[0088] A processing unit, configured to respectively input the plurality of second input images into a second network for pose recognition, so as to obtain a fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image.
[0089] A differentiable renderer, configured to perform differentiable rendering according to the three-dimensional model of the first object to obtain a fourth synthesized image of the first object in each fourth pose; the second input image of one fourth pose corresponds to the fourth synthesized image rendered in one fourth pose.
[0090] A construction unit, configured to construct a second loss function according to second difference information between each fourth synthesized image and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image.
[0091] An update unit, configured to update the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is smaller than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image.
[0092] In another possible implementation, the second loss function Loss 2 may satisfy the following expression: X is greater than or equal to 1, λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image. This implementation provides an expression of a specific second loss function to train the pose detection network and improves the accuracy of the pose detection network recognition.
[0093] In another possible implementation, the second difference information may include one or more of the following: the difference in IOU between the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose for obtaining the fourth synthesized image and the pose obtained by the first network for the fourth synthesized image, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image. This implementation provides possible realizations of the second difference information and enriches the content of the second difference information.
[0094] It should be noted that the display device provided in the fifth aspect is used to implement the display method provided in the second aspect above. Its specific implementation can refer to the specific implementation of the second aspect above, and will not be elaborated here.
[0095] In a sixth aspect, a device for training an object pose detection network is provided. The device may include: an acquisition unit, a processing unit, a differentiable renderer, a construction unit, and an update unit. Among them:
[0096] The acquisition unit is configured to acquire a plurality of second input images including a first object. The plurality of second input images include a ground truth image of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from a three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different.
[0097] The processing unit is configured to respectively input the plurality of second input images into a second network for pose recognition to obtain a fourth pose of the first object in each second input image output by the second network; the second network is used to recognize the pose of the first object in the image.
[0098] The differentiable renderer is configured to, according to the three-dimensional model of the first object, perform differentiable rendering to obtain a fourth synthesized image of the first object at each fourth pose; the second input image of one fourth pose corresponds to the fourth synthesized image rendered at one fourth pose.
[0099] The construction unit is configured to construct a second loss function according to second difference information between each fourth synthesized image and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image.
[0100] The update unit is configured to update the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the ground truth pose of the first object in the image is smaller than the difference between the pose of the first object recognized by the second network in the image and the ground truth pose of the first object in the image.
[0101] By training the pose detection network, the accuracy of the pose detection network in recognition is improved, and the difference between the output of the pose detection network and the ground truth pose of the object in the image is reduced.
[0102] It should be noted that the device for optimizing the pose recognition network provided in the sixth aspect is used to implement the method for optimizing the pose recognition network provided in the third aspect above. Its specific implementation can refer to the specific implementation of the third aspect above, and will not be elaborated here.
[0103] In a seventh aspect, the present application provides an electronic device, which can implement the functions in the method examples described in the first aspect, the second aspect, or the third aspect above. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The electronic device can exist in the form of a chip product.
[0104] In a possible implementation, the electronic device may include a processor and a transmission interface. Among them, the transmission interface is used to receive and send data. The processor is configured to call program instructions stored in the memory so that the electronic device executes the functions in the method examples described in the first aspect, the second aspect, or the third aspect above.
[0105] In an eighth aspect, a computer-readable storage medium is provided, including instructions that, when running on a computer, cause the computer to execute the object registration method, or the display method, or the method of training an object pose detection network described in any of the above aspects or any possible implementation.
[0106] In a ninth aspect, a computer program product is provided that, when running on a computer, causes the computer to execute the object registration method, or the display method, or the method of training an object pose detection network described in any of the above aspects or any possible implementation.
[0107] In a tenth aspect, a chip system is provided. The chip system includes a processor and may also include a memory for implementing the functions in the above method. The chip system may be composed of chips or may include chips and other discrete devices.
[0108] The solutions provided in the fourth aspect to the tenth aspect above are used to implement the methods provided in the first aspect, the second aspect, or the third aspect above, and thus can achieve the same beneficial effects as the first aspect, the second aspect, or the third aspect, and will not be elaborated here.
[0109] It should be noted that, on the premise that the solutions do not conflict, any possible implementation in each of the above aspects can be combined. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] Figure 1 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application;
[0111] Figure 2 It is a schematic software structure diagram of a terminal device provided by an embodiment of the present application;
[0112] Figure 3 It is a schematic diagram of the virtual-real fusion effect of the explanation information and the real object provided by an embodiment of the present application;
[0113] Figure 4 Schematic diagram of a pose detection process provided by an embodiment of this application;
[0114] Figure 5a Schematic diagram of a method flow of incremental learning provided by an embodiment of this application;
[0115] Figure 5b Schematic diagram of a system architecture provided by an embodiment of this application;
[0116] Figure 6 Schematic diagram of the structure of a convolutional neural network provided by an embodiment of this application;
[0117] Figure 7a Schematic diagram of a chip hardware structure provided by an embodiment of this application;
[0118] Figure 7b Schematic diagram of the overall system framework of the solution provided by an embodiment of this application;
[0119] Figure 8 Schematic diagram of a process of an object registration method provided in Embodiment 1 of this application;
[0120] Figure 9 Schematic diagram of a spherical surface provided by an embodiment of this application;
[0121] Figure 10 Schematic diagram of a process of another object registration method provided in Embodiment 1 of this application;
[0122] Figure 11 Schematic diagram of an image area provided by an embodiment of this application;
[0123] Figure 12 Schematic diagram of a process of an optimization method provided in Embodiment 2 of this application;
[0124] Figure 13 Schematic diagram of a process of another optimization method provided in Embodiment 2 of this application;
[0125] Figure 14 Schematic diagram of a process of yet another object registration method provided by an embodiment of this application;
[0126] Figure 15 Schematic diagram of a process of a method for training an object pose detection network provided in Embodiment 3 of this application;
[0127] Figure 16 Schematic diagram of a process of another method for training an object pose detection network provided in Embodiment 3 of this application;
[0128] Figure 17Schematic flowchart of a display method provided in the fourth embodiment of the present application;
[0129] Figure 18 Schematic flowchart of a method for determining whether a first image contains a recognizable object provided in the fourth embodiment of the present application;
[0130] Figure 19 Schematic diagram of a mobile phone interface provided in the embodiment of the present application;
[0131] Figure 20a Another schematic diagram of a mobile phone interface provided in the embodiment of the present application;
[0132] Figure 20b Another schematic diagram of a mobile phone interface provided in the embodiment of the present application;
[0133] Figure 21 Schematic diagram of the structure of an object registration device provided in the embodiment of the present application;
[0134] Figure 22 Schematic diagram of the structure of an optimization device provided in the embodiment of the present application;
[0135] Figure 23 Schematic diagram of the structure of a device for training an object pose detection network provided in the embodiment of the present application;
[0136] Figure 24 Schematic diagram of the structure of a display device provided in the embodiment of the present application;
[0137] Figure 25 Schematic diagram of the structure of the device provided in the embodiment of the present application. Detailed implementation manners
[0138] In the embodiments of the present application, in order to facilitate a clear description of the technical solutions of the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit to being different. There is no sequence or size order between the technical features described by the "first" and "second".
[0139] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0140] In the embodiments of the present application, at least one can also be described as one or more. Multiple can be two, three, four, or more, and the present application does not make any restrictions.
[0141] In addition, the network architectures and scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art will know that with the evolution of network architectures and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0142] Before describing the embodiments of the present application, the nouns involved in the present application are first uniformly explained here, and no further explanations will be given later.
[0143] An image, which can also be called a picture, refers to a visual picture. The image referred to in the present application can be a static image, a video frame in a video stream, or others, without limitation.
[0144] An object can be a person or thing that exists. For example, things can be buildings, commodities, plants, animals, etc., and will not be listed one by one here.
[0145] Pose refers to the attitude of an object in the camera coordinate system. Pose can include 6DoF pose, that is, the translational and rotational postures of the object relative to the camera.
[0146] Pose detection refers to detecting and recognizing the pose of an object in an image.
[0147] The real image of an object refers to an image that contains a static visual picture of the object and the drawing of the background area. The real image of an object can be in the red, green, blue (RGB) format or the red, green, blue, depth (RGBD) format.
[0148] Rendering is the process of converting the 3D model of an object into a 2D image through a renderer. Usually, scenes and entities are represented in three dimensions, which can be closer to the real world and facilitate manipulation and transformation, while most graphic display devices are two-dimensional rasterized displays and dot matrix printers. A raster display can be regarded as a pixel matrix, and any graphic displayed on a raster display is actually a set of pixels with one or more colors and grayscales. Representing a three-dimensional entity scene through raster and dot matrix is image rendering - that is, rasterization.
[0149] Conventional rendering refers to a rendering method in which rasterization is not differentiable.
[0150] Differentiable rendering refers to a rendering method in which rasterization is differentiable. Since the rendering process is differentiable, a loss function can be constructed based on the difference between the rendered image and the real image, and the parameters of differentiable rendering can be updated to improve the authenticity of the differentiable rendering result.
[0151] The synthetic image of an object refers to an image that only contains the object obtained by rendering the 3D model of the object in a certain desired pose. The synthetic image of an object in a certain pose is equivalent to the image obtained by taking a photo of the object in that pose.
[0152] The image corresponding to the synthetic image refers to the pose for rendering the synthetic image, which is obtained by inputting the corresponding image into a neural network. It should be understood that the synthetic image and the source image of the pose for rendering the synthetic image correspond to each other, and will not be elaborated one by one hereinafter. Among them, the image corresponding to the synthetic image can be the real image of the object or other synthetic images of the object.
[0153] The black-and-white image of an object refers to a black-and-white pixel image that only contains the object and does not contain the background. Specifically, the black-and-white image of an object can be represented by a binary image, where the pixel values of the area containing the object are 1 and the pixel values of other areas are 0.
[0154] Local feature points are local expressions of image features, which reflect the local characteristics of the image. Local feature points are the points on the image that have obvious distinctiveness from other pixels, including but not limited to corner points, key points, etc. In image processing, local feature points mainly refer to points or blocks with scale invariance. Scale invariance means that for the same object or scene, when multiple pictures are taken from different angles, the same places can be recognized as the same. Local feature points can include SIFT feature points, SURF feature points, DAISY feature points, etc. Usually, methods such as FAST and DOG can be used to extract local feature points of an image. The descriptor of a local feature point is a high-dimensional vector that characterizes the local image information of the feature point.
[0155] Global features refer to features that can represent the entire image. Global features are relative to local image features and are used to describe the overall features such as the color and shape of an image or object. For example, global features can include color features, texture features, shape features, etc. Usually, the method of bag-of-words tree can be used to extract global features of an image. The descriptor of a global feature is a high-dimensional vector that characterizes the image information of the entire image or a larger area.
[0156] The calculated value can refer to the mathematical calculated value of multiple data, and the mathematical calculation can be taking the average, or taking the maximum, or taking the minimum, or others.
[0157] For the sake of clear and concise description of the following embodiments, a brief introduction to the related technologies is given first:
[0158] In recent years, the functions of terminal devices have become increasingly rich, bringing a better user experience. For example, a terminal device can implement a virtual reality (VR) function, enabling a user to be in a virtual world and experience the virtual world. Another example is that a terminal device can implement an augmented reality (AR) function, combining virtual objects with the real scene and enabling the user to interact with the virtual objects.
[0159] Among them, the terminal device can be a smart phone, a tablet computer, a wearable device, an AR / VR device, etc. The specific form of the terminal device is not limited in this application. A wearable device can also be called a wearable intelligent device, which is the general term for devices developed by applying wearable technology to the intelligent design of daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that can be directly worn on the body or integrated into the user's clothes or accessories. A wearable device is not just a hardware device, but also realizes powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable intelligent devices include those with complete functions, large sizes, and can realize complete or partial functions without relying on a smart phone, such as smart watches or smart glasses, etc., and those that only focus on a certain type of application function and need to cooperate with other devices such as smart phones, such as various smart bracelets and smart jewelry for physical sign monitoring.
[0160] In this application, the structure of the terminal device can be as Figure 1 shown. As Figure 1 shown, the terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0161] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the terminal device 100. In some other embodiments, the terminal device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0162] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. For example, in this application, when the processor 110 determines that the first image meets the abnormal condition, it may control to turn on other cameras.
[0163] Among them, the controller may be the nerve center and command center of the terminal device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0164] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0165] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0166] The MIPI interface may be used to connect the processor 110 to peripheral devices such as the display screen 194, the camera 193, etc. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the terminal device 100.
[0167] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, an MIPI interface, etc.
[0168] The USB interface 130 is an interface that conforms to the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device 100, and can also be used for data transmission between the terminal device 100 and peripheral devices. It can also be used to connect headphones to play audio through the headphones. This interface can also be used to connect other terminal devices, such as AR devices, etc.
[0169] It can be understood that the interface connection relationships between the modules illustrated in this embodiment are only illustrative descriptions and do not constitute a structural limitation on the terminal device 100. In some other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0170] The power management module 141 is used to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 may also be provided in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may also be provided in the same device.
[0171] The wireless communication function of the terminal device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0172] The terminal device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0173] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oled, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.
[0174] A series of graphical user interfaces (GUIs) can be displayed on the display screen 194 of the terminal device 100, and these GUIs are all the main screens of the terminal device 100. Generally speaking, the size of the display screen 194 of the terminal device 100 is fixed, and only a limited number of controls can be displayed in the display screen 194 of the terminal device 100. A control is a GUI element, which is a software component included in an application and controls all the data processed by the application and the interaction operations related to these data. Users can interact with the control through direct manipulation to read or edit relevant information of the application. Generally speaking, controls can include visible interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, Widgets, etc.
[0175] The terminal device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, the application processor, etc.
[0176] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193.
[0177] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the terminal device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0178] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0179] The video codec is used for compressing or decompressing digital videos. The terminal device 100 can support one or more video codecs. In this way, the terminal device 100 can play or record videos in multiple encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0180] The NPU is a neural-network (NN) computing processor. By learning from the structure of biological neural networks, such as the transmission pattern between neurons in the human brain, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the terminal device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0181] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0182] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121. For example, in this embodiment, the processor 110 can obtain the pose of the terminal device 100 by executing the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the terminal device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0183] The terminal device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone interface 170D, and the application processor, etc. Such as music playback, recording, etc.
[0184] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0185] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device 100 can listen to music or hands-free calls through the speaker 170A.
[0186] The receiver 170B, also known as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device 100 answers a call or a voice message, the voice can be listened to by bringing the receiver 170B close to the human ear.
[0187] The microphone 170C, also known as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 100 can be provided with at least one microphone 170C. In some other embodiments, the terminal device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the terminal device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0188] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0189] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In some embodiments, the pressure sensor 180A may be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor may include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The terminal device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation acts on the display screen 194, the terminal device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The terminal device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities may correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0190] The gyroscope sensor 180B can be used to determine the motion posture of the terminal device 100. In some embodiments, the angular velocity of the terminal device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the shaking angle of the terminal device 100, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to offset the shaking of the terminal device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.
[0191] The barometric pressure sensor 180C is used to measure barometric pressure. In some embodiments, the terminal device 100 calculates the altitude according to the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0192] The magnetic sensor 180D includes a Hall sensor. The terminal device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In some embodiments, when the terminal device 100 is a flip phone, the terminal device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Furthermore, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic flip unlocking are set.
[0193] The acceleration sensor 180E can detect the magnitude of the acceleration of the terminal device 100 in various directions (generally three axes). When the terminal device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the terminal device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0194] A distance sensor 180F is used to measure distance. The terminal device 100 can measure distance through infrared or laser. In some embodiments, when shooting a scene, the terminal device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0195] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode may be an infrared light-emitting diode. The terminal device 100 emits infrared light outward through the light-emitting diode. The terminal device 100 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal device 100. When insufficient reflected light is detected, the terminal device 100 can determine that there is no object near the terminal device 100. The terminal device 100 can use the proximity light sensor 180G to detect that the user holds the terminal device 100 close to the ear for a call, so as to automatically turn off the screen to achieve the purpose of power saving. The proximity light sensor 180G can also be used for automatic unlocking and locking of the holster mode and pocket mode.
[0196] An ambient light sensor 180L is used to sense the ambient light brightness. The terminal device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the terminal device 100 is in the pocket to prevent accidental touch.
[0197] A fingerprint sensor 180H is used to collect fingerprints. The terminal device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.
[0198] A temperature sensor 180J is used to detect temperature. In some embodiments, the terminal device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the terminal device 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the terminal device 100 heats the battery 142 to avoid abnormal shutdown of the terminal device 100 caused by low temperature. In still other embodiments, when the temperature is lower than yet another threshold, the terminal device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.
[0199] The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect touch operations acting thereon or in its vicinity. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the terminal device 100, at a different position from that of the display screen 194.
[0200] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can analyze the voice signals based on the vibration signals of the vibrating bone mass acquired by the bone conduction sensor 180M to implement the voice function. The application processor can analyze the heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 180M to implement the heart rate detection function.
[0201] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button or a touch button. The terminal device 100 can receive button inputs and generate key signal inputs related to the user settings and function control of the terminal device 100.
[0202] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations acting on different regions of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving information, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0203] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0204] In addition, an operating system runs on the above components. For example, the iOS operating system developed by Apple Inc., the Android open-source operating system developed by Google Inc., the Windows operating system developed by Microsoft Corporation, etc. Application programs can be installed and run on this operating system.
[0205] The operating system of the terminal device 100 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, taking the Android system with a layered architecture as an example, the software structure of the terminal device 100 will be exemplarily described.
[0206] Figure 2 It is the software structure block diagram of the terminal device 100 in the embodiments of the present application.
[0207] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, which are the application layer, the application framework layer, Android runtime and system libraries, and the kernel layer.
[0208] The application layer may include a series of application packages. As Figure 2 shown, the application packages may include applications such as a camera, a gallery, a calendar, a call, a map, a navigation, a WLAN, a Bluetooth, music, a video, a short message, etc. For example, when taking a photo, the camera application can access the camera interface management service provided by the application framework layer.
[0209] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. As Figure 2 shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc. For example, in the embodiments of the present application, when taking a photo, the application framework layer can provide APIs related to the photo-taking function for the application layer and provide a camera interface management service for the application layer to implement the photo-taking function.
[0210] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0211] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, a phone book, etc.
[0212] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures.
[0213] The telephone manager is used to provide the communication function of the terminal device 100. For example, the management of call status (including answering, hanging up, etc.).
[0214] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.
[0215] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears on the screen in the form of a dialogue window. For example, it prompts text information in the status bar, emits a prompt tone, the terminal device vibrates, the indicator light flashes, etc.
[0216] Android Runtime includes a core library and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system.
[0217] The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0218] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0219] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.
[0220] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0221] The media library supports the playback and recording of various common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: Moving Picture Experts Group (MPEG) 4, H.264, MP3, Advanced Audio Coding (AAC), Adaptive Multi Rate (AMR), Joint Photographic Experts Group (JPEG), Portable Network Graphic Format (PNG), etc.
[0222] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0223] The two-dimensional (2D) graphics engine is a drawing engine for 2D drawing.
[0224] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0225] It should be noted that although the embodiments of this application are described by taking the system as an example, its basic principle also applies to terminal devices based on or operating systems such as.
[0226] Next, the technical solutions in this application will be described in conjunction with the accompanying drawings.
[0227] The 3D object registration method provided by the embodiments of this application can be applied to scenarios such as AR, VR, and virtual content that needs to display objects. Specifically, the 3D object registration method of the embodiments of this application can be applied to the AR scenario. Next, in conjunction with Figure 1 and the AR scenario, the working processes of the software and hardware of the terminal device 100 will be exemplarily described.
[0228] The touch sensor 180K receives a touch operation and reports it to the processor 110, causing the processor to start the AR application in response to the touch operation and display the user interface of the AR application on the display screen 194. For example, when the touch sensor 180K receives a touch operation on the AR icon, it reports the touch operation on the AR icon to the processor 110, causing the processor 110 to start the AR application corresponding to the AR icon in response to the touch operation and display the user interface of AR on the display screen 194. In addition, in the embodiments of the present application, the terminal can also be caused to start AR and display the user interface of AR on the display screen 194 in other ways. For example, when the terminal is blacked out, displays a lock screen interface, or displays a certain user interface after unlocking, it can start AR in response to a user's voice command or shortcut operation, etc., and display the user interface of AR on the display screen 194.
[0229] In the APP in the terminal device 100 for detecting and tracking the pose of an object, a network for detecting the pose of the object is configured. When the user starts the APP for detecting and tracking the pose of an object in the terminal device 100, the terminal device 100 captures an image in the field of view through the camera 193, and the network for detecting the pose of the object identifies the recognizable object included in the image, obtains the pose of the recognizable object, and then, according to the obtained pose, superimposes and displays the virtual content corresponding to the recognizable object on the captured image through the display screen 194.
[0230] Exemplarily, during the process of shopping in a mall, visiting a museum or an exhibition hall, etc., after the terminal device 100 captures a scene image, it identifies the recognizable object in the image, and superimposes the explanation information of the recognizable object on different positions of the three-dimensional object accurately in the three-dimensional space according to the pose of the recognizable object, and presents it to the user. As Figure 3 Schematically shows the virtual-real fusion effect of the explanation information and the real object, and can intuitively, vividly and vividly present the information of the real object in the image to the user.
[0231] In the current scenario of displaying virtual content of an object, usually offline, according to the three-dimensional model of the object, the network for detecting the pose of the object configured in the terminal device 100 is trained, so that the network supports the pose recognition of multiple objects, that is, the object to be recognized is registered in the network. In the actual application process of 3D object pose detection, it is necessary to add and delete recognizable objects efficiently and quickly. Currently, usually based on machine learning methods, recognizable objects are added. For each newly added object, all recognizable objects need to be retrained again, which will lead to a linear increase in training time and affect the recognition effect of the objects that have been trained well.
[0232] Figure 4 Schematically shows an existing pose detection process, such asFigure 4 As shown, an image containing objects is input into a multi-object pose estimation network. The multi-object pose estimation network outputs the poses and categories of the recognizable objects contained in the image, and optimizes the output poses. The multi-object pose estimation network used in this process is generated by offline training. When a user needs to add a new recognizable object, it is necessary to retrain together with all the recognizable objects supported by the original network to obtain a new multi-object pose estimation network, so as to support the pose detection of the newly added recognizable object. In this way, retraining for the newly added object will cause a sharp increase in the training time, and the newly added object will affect the pose detection effect of the already trained objects, resulting in a decrease in the detection accuracy and success rate.
[0233] To solve the problem of the increased training time caused by retraining all recognizable objects for the newly added object, the industry has proposed an incremental learning method. The process of this method can be as Figure 5a shown. The user submits the 3D model of the object expected to be recognized, and trains a multi-object pose estimation network M0 based on the submitted 3D model; when a new recognizable object is added, the user newly submits the 3D model of the object expected to be recognized, and based on the trained network M0, performs incremental training using a small amount of data of the already trained objects to obtain a new network M1; when another new recognizable object is added, the user continues to submit the 3D model of the object expected to be recognized, and based on the trained network M1, performs incremental training using a small amount of data of the already trained objects to obtain a new network M2, and so on. In this solution, when a new object to be recognized is added to the training set, only a small amount of data of the already trained objects is used, and incremental learning is performed on the basis of the existing model, which can greatly reduce the retraining time. However, due to the catastrophic forgetting problem faced by incremental learning, that is, the training of the new model only refers to a small amount of data of the already recognized objects. As the number of newly added objects increases, the performance of the already trained objects will drop sharply.
[0234] Based on this, the present application provides a 3D object registration method, which specifically includes: configuring a single-object pose detection network, constructing a loss function by using the difference between the real image of the 3D object and the synthetic images obtained by differentiable rendering at multiple poses to train the single-object pose detection network, so as to obtain a pose detection network for extracting the pose of the 3D object, and then using the trained pose detection network to extract the features of the 3D object in the real image of the 3D object and the synthetic images obtained by differentiable rendering at multiple poses, and recording the features and the identifier of the 3D object to complete the registration of the 3D object. Since the 3D object registration method provided by the present application uses a single-object pose detection network, even if a new recognizable object is added, the training time is short and it will not affect the recognition effect of other recognizable objects; in addition, by using the difference between the real image of the 3D object and the synthetic images obtained by differentiable rendering at multiple poses to construct a loss function, the accuracy of the single-object pose detection network is improved.
[0235] The method provided by this application will be described below from the perspective of model training and model application:
[0236] The method for obtaining the pose of an object provided by the embodiments of this application involves computer vision processing and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as the image of the object in this application), and finally obtains a trained single-object pose detection network. Moreover, the object registration method provided by the embodiments of this application can use the above-trained single-object pose detection network to input input data (such as the image including the object to be recognized in this application) into the single-object pose detection network corresponding to the object to be recognized that has been trained, and obtain output data (such as the pose of the recognizable object in the image in this application). It should be noted that the training method of the single-object pose detection network and the object registration method provided by the embodiments of this application are inventions based on the same concept, and can also be understood as two parts of a system or two stages of an overall process: such as the model training stage and the model application stage.
[0237] Since the embodiments of this application involve the application of a large number of neural networks, for the convenience of understanding, relevant terms and related concepts such as neural networks involved in the embodiments of this application will be introduced below.
[0238] (1) Neural network (NN)
[0239] A neural network is a machine learning model and a machine learning technology that simulates the neural network of the human brain to achieve artificial intelligence-like capabilities. The input and output of the neural network can be configured according to actual needs, and the neural network can be trained with sample data to minimize the error between its output and the true output corresponding to the sample data. A neural network can be composed of neural units, and a neural unit can refer to an operation unit with x s and intercept 1 as inputs, and the output of this operation unit can be:
[0240]
[0241] where s = 1, 2,..., n, n is a natural number greater than 1, and W s is x sThe weight is \(w\), \(b\) is the bias of the neuron. \(f\) is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0242] (2) Deep neural network
[0243] A deep neural network (DNN), also known as a multi - layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the \(i\) - th layer must be connected to any neuron in the \((i + 1)\) - th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply speaking, it is the following linear relationship expression: Among them, \(\mathbf{x}\) is the input vector, \(\mathbf{y}\) is the output vector, \(b\) is the offset vector, \(W\) is the weight matrix (also known as the coefficient), and \(\alpha()\) is the activation function. Each layer is just a simple operation on the input vector \(\mathbf{x}\) to get the output vector Since the DNN has many layers, the number of coefficients \(W\) and offset vectors \(b\) is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient \(W\) as an example: Suppose in a three - layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient \(W\) is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the \(k\) - th neuron in the \((L - 1)\) - th layer to the \(j\) - th neuron in the \(L\) - th layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0244] (3) Convolutional Neural Network
[0245] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can only be connected to some adjacent layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weight here is the convolution kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle is that the statistical information of a certain part of an image is the same as that of other parts. That is to say, the image information learned in a certain part can also be used in another part. Therefore, for all positions on the image, the same learned image information can be used. In the same convolutional layer, multiple convolution kernels can be used to extract different image information. Generally, the more convolution kernels there are, the richer the image information reflected by the convolution operation.
[0246] The convolution kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolution kernel can learn reasonable weights. In addition, the direct benefit brought by sharing weights is to reduce the connections between layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0247] (4) Recurrent neural networks (RNN) are used to process sequential data. In traditional neural network models, it goes from the input layer to the hidden layer and then to the output layer, with full connections between layers, while there are no connections between individual nodes within each layer. Although this ordinary neural network has solved many problems, it is still powerless in many aspects. For example, when you want to predict the next word in a sentence, you generally need to use the previous words because the words before and after in a sentence are not independent. The reason RNN is called a recurrent neural network is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes within the hidden layer itself are no longer unconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNN can process sequential data of any length. The training of RNN is the same as that of traditional CNN or DNN. The error backpropagation algorithm is also used, but there is one difference: that is, if the RNN is unfolded, the parameters such as W are shared; while the traditional neural network mentioned above is not like this. And in using the gradient descent algorithm, the output of each step depends not only on the network of the current step but also on the states of the networks in several previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).
[0248] Since there are already convolutional neural networks, why do we still need recurrent neural networks? The reason is very simple. In convolutional neural networks, there is a premise assumption that elements are independent of each other, and the input and output are also independent, such as cats and dogs. However, in the real world, many elements are interconnected, such as the change of stocks over time, or for example, a person says: "I like traveling, and the favorite place is Yunnan. I must go there if I have the chance in the future." Here, for filling in the blank, humans should all know that it is "Yunnan". Because humans will make inferences based on the context, but how can we make machines do this? This is where RNN came into being. RNN aims to enable machines to have the ability to remember like humans. Therefore, the output of RNN needs to depend on the current input information and historical memory information.
[0249] (6) Loss function
[0250] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0251] (7) Pixel value
[0252] The pixel value of an image can be an RGB (Red, Green, Blue) color value, and the pixel value can be a long integer representing a color. For example, the pixel value is 256*Red + 100*Green + 76*Blue, where Blue represents the blue component, Green represents the green component, and Red represents the red component. Among the respective color components, the smaller the value, the lower the brightness, and the larger the value, the higher the brightness. For a grayscale image, the pixel value can be a grayscale value.
[0253] The following introduces the system architecture provided by the embodiments of this application.
[0254] See the appendix Figure 5b, an embodiment of the present invention provides a system architecture 500. As shown in the system architecture 500, a data acquisition device 560 is used to acquire training data. In the embodiments of the present application, the training data includes: real images and / or synthetic images of the object to be recognized; and the training data is stored in a database 530. A training device 520 trains a target model / rule 501 based on the training data maintained in the database 530. How the training device 520 obtains the target model / rule 501 based on the training data will be described in more detail in Embodiment 2 and Embodiment 3. The target model / rule 501 may be a single-object pose detection network (the first network for extracting the pose of an object in an image) described in the embodiments of the present application. That is, when an image is input into the target model / rule 501, the pose of the recognizable object included in the image can be obtained; or, the target model / rule 501 may be a differentiable renderer described in the embodiments of the present application. That is, when the three-dimensional model of the object and the preset pose are input into the target model / rule 501, the synthetic image of the object in the preset pose can be obtained. The target model / rule 501 in the embodiments of the present application may specifically be a single-object pose detection network or a differentiable renderer. In the embodiments provided by the present application, the single-object pose detection network is obtained by training a single-object basic pose detection network. It should be noted that in actual applications, the training data maintained in the database 530 may not necessarily come from the acquisition of the data acquisition device 560, and it may also be received from other devices. Additionally, it should be noted that the training device 520 does not necessarily train the target model / rule 501 completely based on the training data maintained in the database 530. It may also obtain training data from the cloud or other places for model training. The above description should not be construed as a limitation on the embodiments of the present application.
[0255] The target model / rule 501 trained according to the training device 520 can be applied to different systems or devices, such as being applied to Figure 5b the execution device 510 shown. The execution device 510 may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, AR / VR, a vehicle-mounted terminal, etc., or may also be a server or the cloud, etc. In the appendix Figure 5b , the execution device 510 is configured with an I / O interface 512 for data interaction with external devices. A user can input data to the I / O interface 512 through a client device 540. The input data in the embodiments of the present application may include: the three-dimensional model of the object to be recognized, the real image of the object to be recognized, and the synthetic image rendered by the three-dimensional model of the object to be recognized in different poses.
[0256] During the process of the computing module 511 of the execution device 510 performing processing related to computing and the like, the execution device 510 can call data, code, etc. in the data storage system 550 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 550.
[0257] Finally, the I / O interface 512 returns the processing result, such as the virtual content and pose of the recognizable object in the obtained image, to the client device 540, so as to provide the virtual content displayed according to the pose to the user, realizing the experience of combining virtual and real.
[0258] It should be noted that the training device 520 can generate corresponding target models / rules 501 based on different training data for different targets or tasks, and the corresponding target models / rules 501 can be used to achieve the above-mentioned targets or complete the above-mentioned tasks, so as to provide the required results for the user.
[0259] In Figure 5b the shown case, the user can manually give input data, and this manual giving can be operated through the interface provided by the I / O interface 512. In another case, the client device 540 can automatically send input data to the I / O interface 512. If the client device 540 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 510 on the client device 540, and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 540 can also be used as a data acquisition end to acquire the input data input to the I / O interface 512 and the output result output from the I / O interface 512 as shown in Figure 5b as new sample data and store them in the database 530. Of course, it can also be collected without passing through the client device 540, but directly by the I / O interface 512 taking the input data input to the I / O interface 512 and the output result output from the I / O interface 512 as shown in the figure as new sample data and storing them in the database 530.
[0260] It should be noted that Figure 5b is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 5b the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510.
[0261] The methods and devices provided by the embodiments of the present application can also be used to expand the training database, such as Figure 5bThe I / O interface 512 of the execution device 510 shown can send the processed image of the execution device (such as the composite image rendered by the object to be recognized in different poses) and the real image of the object to be recognized input by the user to the database 530 together as a training data pair, so that the training data maintained by the database 530 is richer, thereby providing richer training data for the training work of the training device 520.
[0262] Such as Figure 5b As shown, the target model / rule 501 is trained according to the training device 520. The target model / rule 501 can be a single object pose recognition network (the first network and the second network described in the embodiments of the present application) in the embodiments of the present application. In the single object pose recognition networks provided in the embodiments of the present application, they can all be convolutional neural networks, recurrent neural networks, or others.
[0263] As described in the basic concept introduction above, a convolutional neural network is a deep neural network with a convolutional structure and is a deep learning architecture. A deep learning architecture refers to performing multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network can respond to the input image.
[0264] Such as Figure 6 As shown, the convolutional neural network (CNN) 600 can include an input layer 610, a convolutional layer / pooling layer 620 (where the pooling layer is optional), and a neural network layer 630.
[0265] Convolutional layer / pooling layer 620:
[0266] Convolutional layer:
[0267] Such as Figure 6 As shown, the convolutional layer / pooling layer 620 can include layers such as examples 621-626. For example: in one implementation, layer 621 is a convolutional layer, layer 622 is a pooling layer, layer 623 is a convolutional layer, layer 624 is a pooling layer, 625 is a convolutional layer, and 626 is a pooling layer; in another implementation, 621 and 622 are convolutional layers, 623 is a pooling layer, 624 and 625 are convolutional layers, and 626 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0268] Next, taking the convolutional layer 621 as an example, the internal working principle of one convolutional layer will be introduced.
[0269] The convolutional layer 621 may include a number of convolutional operators, also known as kernels, which act as filters for extracting specific information from the input image matrix in image processing. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on an image, the weight matrix typically processes the input image pixel by pixel (or two pixels at a time, depending on the value of the stride) along the horizontal direction, thus completing the task of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image, and during the convolution operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix produces a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices of the same size (rows × columns), i.e., multiple matrices of the same type, are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, where the dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these multiple weight matrices of the same size also have the same size. Then, the multiple feature maps of the same size that are extracted are combined to form the output of the convolution operation.
[0270] The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can be used to extract information from the input image, enabling the convolutional neural network 600 to make correct predictions.
[0271] When the convolutional neural network 600 has multiple convolutional layers, the initial convolutional layer (such as 621) often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network 600 increases, the later convolutional layers (such as 626) extract increasingly complex features, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0272] Pooling layer:
[0273] Since it is often necessary to reduce the number of training parameters, a pooling layer is often introduced periodically after the convolutional layer. Figure 6Each of the layers 621 - 626 shown in 620 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. During image processing, the sole purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of pixel values in the image within a specific range as the result of average pooling. The max pooling operator can take the pixel with the maximum value within the specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer, and each pixel point in the image output by the pooling layer represents the average value or the maximum value of the corresponding sub-region of the image input to the pooling layer.
[0274] Neural network layer 630:
[0275] After being processed by the convolutional / pooling layer 620, the convolutional neural network 600 is still not sufficient to output the required output information. Because as mentioned before, the convolutional / pooling layer 620 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network 600 needs to use the neural network layer 630 to generate one or a set of outputs of the number of required classes. Therefore, the neural network layer 630 can include multiple hidden layers (such as Figure 6 631, 632 to 63n shown) and an output layer 640. The parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, and so on...
[0276] After the multiple hidden layers in the neural network layer 630, that is, the last layer of the entire convolutional neural network 600 is the output layer 640. The output layer 640 has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 600 (such as Figure 6 the propagation from 610 to 640 is the forward propagation) is completed, the backpropagation (such as Figure 6 the propagation from 640 to 610 is the backpropagation) will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network 600, that is, the error between the result output by the convolutional neural network 600 through the output layer and the ideal result.
[0277] It should be noted that as Figure 6The convolutional neural network 600 shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network may also exist in the form of other network models.
[0278] The following introduces a chip hardware structure provided by an embodiment of the present application.
[0279] Figure 7a A chip hardware structure provided by an embodiment of the present invention, the chip includes a neural network processor (NPU) 70. The chip can be disposed in an execution device 510 as shown in Figure 5b to complete the computing work of the computing module 511. The chip can also be disposed in a training device 520 as shown in Figure 5b to complete the training work of the training device 520 and output a target model / rule 501. Algorithms of each layer in the convolutional neural network as shown in Figure 6 can all be implemented in the chip as shown in Figure 7a shown.
[0280] As shown in Figure 7a shown, the NPU 70 is mounted on the main central processing unit (CPU) (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 70, and the controller 704 controls the arithmetic circuit 703 to extract data from the memory (weight memory or input memory) and perform operations.
[0281] In some implementations, the arithmetic circuit 703 includes multiple processing units (process engine, PE) inside. In some implementations, the arithmetic circuit 703 is a two-dimensional systolic array. The arithmetic circuit 703 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 703 is a general matrix processor.
[0282] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit 703 fetches the corresponding data of the matrix B from the weight memory 702 and caches it on each PE in the arithmetic circuit. The arithmetic circuit 703 fetches the data of the matrix A from the input memory 701 and performs matrix operations with the matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 708 (accumulator).
[0283] The vector calculation unit 707 can further process the output of the arithmetic circuit 703, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, and so on. For example, the vector calculation unit 707 can be used for network calculations in non-convolutional / non-fully connected layers (FC) of a neural network, such as pooling, batch normalization, local response normalization, etc.
[0284] In some implementations, the vector calculation unit 707 stores the processed output vector into the unified buffer 706. For example, the vector calculation unit 707 can apply a non-linear function to the output of the arithmetic circuit 703, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 707 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as an activation input to the arithmetic circuit 703, for example, for use in subsequent layers in a neural network.
[0285] For example, as Figure 6 shown, the algorithms for each layer in the convolutional neural network can be executed by 703 or 707. Figure 5b The algorithms of the computing module 511 and the training device 520 in
[0286] The unified memory 706 is used to store input data and output data.
[0287] The weight data directly transfers the input data in the external memory to the input memory 701 and / or the unified memory 706, stores the weight data in the external memory into the weight memory 702, and stores the data in the unified memory 706 into the external memory through the direct memory access controller (DMAC) 705.
[0288] The bus interface unit (BIU) 710 is used to interact between the main CPU, the DMAC, and the instruction fetch memory 709 through the bus.
[0289] The instruction fetch buffer 709 connected to the controller 704 is used to store the instructions used by the controller 704.
[0290] The controller 704 is used to call the instructions cached in the instruction memory 709 to control the working process of the arithmetic accelerator.
[0291] Exemplarily, the data here can be illustrative data, which can be Figure 6 the input or output data of each layer in the convolutional neural network shown, or, it can be Figure 5b the input or output data of the computing module 511 and the training device 520 in
[0292] Generally, the unified memory 706, the input memory 701, the weight memory 702, and the fetch memory 709 are all on-chip memories, and the external memory is the memory external to the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0293] Optionally, Figure 5b and Figure 6 the program algorithms in
[0294] Exemplarily, Figure 7b illustrates the overall system framework of the solution provided by this application. As Figure 7b shown, the framework includes two parts: offline registration and online detection.
[0295] Among them, in the offline registration part, the user inputs the three-dimensional model and real pictures of the object, and then trains the basic single-object pose detection network according to the three-dimensional model and real pictures to obtain the single-object pose detection network of the object. The single-object pose detection network of the object is used to detect the pose of the object included in the image, and the detection accuracy is better than that of the basic single-object pose detection network. Further, according to the three-dimensional model and real pictures, using the already obtained single-object pose detection network of the object, the features of the object can be extracted for incremental object registration, and the features of the object are registered with the category of the object (which can be represented by an identifier) to obtain a multi-object category classifier. The multi-object category classifier can be used to identify the categories of recognizable objects included in the image. For the specific operations of the offline registration part, reference can be made to the specific implementation of the object registration method provided in the embodiments of this application below (for example Figure 8 the object registration method illustrated), which will not be elaborated here.
[0296] In the online detection part, after obtaining the input image, a multi-object category classifier obtained from the offline registration part is used for multi-feature fusion classification to obtain the classification result (category) of the recognizable objects included in the input image. A single-object pose detection network of the object obtained from the offline registration part is used to obtain the pose of the recognizable objects included in the input image. According to the pose of the recognizable objects in the input image, virtual content corresponding to the category of the recognizable objects is presented. For the specific operations of the online detection part, reference can be made to the specific implementation of the display method provided in the embodiments of the present application below (for example Figure 17 the schematic display method), which will not be elaborated here.
[0297] Embodiment 1 of the present application provides an object registration method for registering a first object, where the first object is any object to be recognized. The registration process of each object is the same. Embodiment 1 of the present application takes the registration of the first object as an example for description, and others will not be elaborated one by one.
[0298] The object registration method provided in Embodiment 1 of the present application can be executed by an execution device 510 as shown in Figure 5b . The real image in the object registration method can be the input data given by a client device 540 as shown in Figure 5b . The calculation module 511 in the execution device 510 can be used to execute S801 to S803.
[0299] Optionally, the object registration method provided in Embodiment 1 of the present application can be processed by a CPU, or jointly processed by a CPU and a GPU. It is also possible not to use a GPU, but to use other processors suitable for neural network computing. The present application does not make any restrictions.
[0300] As shown in Figure 8 , the object registration method provided in Embodiment 1 of the present application may include:
[0301] S801. Obtain a plurality of first input images including the first object.
[0302] Among them, the plurality of first input images include the real image of the first object and / or a plurality of first synthetic images of the first object.
[0303] In a possible implementation manner, the first input image may include a plurality of real images of the first object.
[0304] Since the number of real images of the object input by the user is limited, when extracting feature information by executing S802 using the real images of the first object, there is a deficiency that the visual features of the object to be recognized under different illuminations, different angles, and distances cannot be fully characterized. Therefore, based on the real images, synthetic images can be added to further increase the effectiveness of feature extraction. Therefore, the first input image may include the real image of the first object and the first synthetic image. The first synthetic image is obtained by differentiable rendering, which can reduce the difference between the first synthetic image and the real image.
[0305] In another possible implementation, the first input image may include a plurality of first synthetic images of the first object. The plurality of first synthetic images may be obtained by differentiable rendering of the three-dimensional model of the first object in a plurality of first poses. The plurality of first poses are different.
[0306] Among them, the plurality of first poses being different means that the plurality of first poses correspond to different shooting angles of the camera.
[0307] Exemplarily, the plurality of first poses may be taken on a sphere. The plurality of first poses may be Figure 9 a plurality of poses obtained by uniform sampling on the sphere shown. The density of the sampled poses on the sphere is not limited in the embodiments of the present application and can be selected according to actual needs. Of course, the first pose may also be a plurality of different poses input by the user, which is not limited.
[0308] Specifically, in the case where the first input image includes the first synthetic image in S801, according to the three-dimensional model of the first object with texture information input by the user, a differentiable renderer (which may also be referred to as a differentiable rendering engine or a differentiable rendering network) can be used to synthesize 2D images (first synthetic images) of the 3D model obtained by the camera in a plurality of first poses.
[0309] Exemplarily, in the case where the first input image includes the first synthetic image in S801, as Figure 10 shown in the specific process of the object registration method, S801 may specifically include S801a and S801b.
[0310] S801a: Perform differentiable rendering on the three-dimensional model of the first object to obtain a plurality of first synthetic images.
[0311] S801b: Obtain the first input image according to the first synthetic image.
[0312] Among them, the first input image obtained in S801b may include the real image of the first object and a plurality of first synthetic images, or the first input image obtained in S801b may only include a plurality of first synthetic images.
[0313] It should be noted that the process of differentiable rendering has been described in detail in the foregoing content and will not be elaborated here.
[0314] S802. Extract the feature information of multiple first input images respectively, where the feature information is used to indicate the features of the first object in the first input image where it is located.
[0315] Among them, the feature information may include descriptors of local feature points and descriptors of global features.
[0316] A descriptor is used to describe a feature and is a data structure for characterizing a feature. The dimension of a descriptor can be multi-dimensional. There can be various types of descriptors, such as SIFT, SURF, MSER, etc. The embodiments of the present application do not limit the type of the descriptor. The descriptor can be in the form of a multi-dimensional vector.
[0317] For example, the descriptor of a local feature point can be a high-dimensional vector representing the local image information of the feature point; the descriptor of a global feature can be a high-dimensional vector representing the image information of the entire image or a larger area.
[0318] It should be noted that the definitions and extraction methods of local feature points and global features have been described in the foregoing content and will not be elaborated here.
[0319] Further optionally, the feature information may further include the positions of local feature points.
[0320] In a possible implementation manner, in S802, the feature information may be extracted in all regions within each first input image.
[0321] In another possible implementation manner, in S802, the region of the first object may be determined in each first input image, and then the feature information may be extracted in the region of the first object.
[0322] Among them, the region of the first object may be the region in the first input image that only includes the first object, or the region of the first object may be the region in the first input image that includes the first object, and this region is included in the first input image.
[0323] Exemplarily, each first input image may be input into the first network in parts, and the region of the first object in each first input image may be determined according to the output of the first network. The first network is used to identify the pose of the first object in the image, or the first network is used to extract the black and white image of the first object in the image.
[0324] Optionally, determining the region of the first object in the first input image and then extracting the feature information in the region of the first object may include but are not limited to the following two possible implementation manners:
[0325] The first implementation method: The first network is used to identify the pose of the first object in the image. In S802, multiple first input images can be respectively input into the first network for pose recognition to obtain the pose of the first object in each first input image. According to the obtained pose of the first object, the three-dimensional model of the first object is respectively projected onto each first input image to obtain the projection area (as the area of the first object) in each first input image. Feature information is extracted in the projection area of each first input image respectively.
[0326] The second implementation method: The first network is used to extract the black-and-white image of the first object in the image. In S802, multiple first input images can be respectively input into the first network for pose recognition to obtain the black-and-white image (as the area of the first object) of the first object in each first input image. Feature information is extracted in the black-and-white image of the first object in each of the first input images respectively.
[0327] Exemplarily, in the area of the first object in each first input image obtained in S802, the pixel positions of visually significant feature points and the corresponding descriptors can be extracted as the feature information of the local feature points, and each descriptor is a multi-dimensional vector. In the area of the first object in each second input image obtained in S802, the feature information of all visual feature points is extracted, and a multi-dimensional vector is output as the feature information of the global feature.
[0328] It should be noted that the present application embodiment does not limit the algorithm for extracting features, and can be selected according to actual needs.
[0329] Exemplarily, as Figure 11 shown in the image area, according to the result obtained by inputting the first input image into the first network, it can be determined that the area where the first object is located in this first input image is Figure 11 the area surrounded by the dashed line box shown in. Feature information of visually significant feature points can be extracted in the area surrounded by the dashed line box shown in Figure 11 shown, and feature information of all visual feature points can be extracted in the area surrounded by the dashed line box shown in Figure 11 shown.
[0330] S803: Corresponding the feature information extracted from each first input image with the identifier of the first object, and registering the first object.
[0331] Specifically, in S803, the feature information of the first object is recorded corresponding to the identifier of the first object, and the registration of the first object is completed. The registration content of multiple objects can be referred to as a multi-object classifier. The multi-object classifier records the corresponding relationship between the features and identifiers of different objects. In practical applications, by extracting the feature information of the object in the image and comparing it with the features recorded in the multi-object classifier, the identifier of the object in the image can be determined, and the recognition of the object in the image can be completed.
[0332] Among them, the identifier of the first object can be used to indicate the category of the first object. The embodiment of the present application does not limit the form of this identifier.
[0333] Specifically, in S803, the feature information of the first object in each first input image extracted in S802 can be stored in a certain structure corresponding to the identifier of the first object, and the registration of the first object is completed.
[0334] The feature information and identifiers of multiple objects stored in the structure are called a multi-object classifier for efficient search.
[0335] Through the object registration method provided by the present application, the features of the object in the image are extracted for object registration, the registration time is short, and it does not affect the recognition performance of other registered objects. Even if many new objects are registered, the registration time will not increase significantly, and the detection performance of the registered objects is also guaranteed.
[0336] It should be noted that the above object registration method can obtain the first network and complete the registration of the object offline. When applied online, the pose of the recognizable object can be obtained through the first network, and the identifier for object registration can be determined according to the features of the recognizable object. Then, according to the pose of the recognizable object, the virtual content corresponding to the identifier of the recognizable object is displayed, thereby completing the presentation effect of the virtual-real result.
[0337] Embodiment 2 of the present application provides an optimization method, which can be combined with Figure 8 the object registration method shown, and the method for training the object pose detection network provided in Embodiment 3 of the present application, or can be used independently. The embodiment of the present application does not limit its usage scenario. The optimization method provided in Embodiment 2 of the present application is as Figure 12 shown, including:
[0338] S1201: Input N real images of the first object into the first network respectively, and obtain the second pose of the first object in each real image output by the first network.
[0339] Among them, N is greater than or equal to 1. The first network is used to recognize the pose of the first object in the image.
[0340] S1202. According to the three-dimensional model of the first object, use a differentiable renderer to perform differentiable rendering to obtain N second synthesized images at each second pose.
[0341] Among them, in S1202, a differentiable renderer is used to perform differentiable rendering to obtain N second synthesized images at each second pose obtained in S1201.
[0342] Specifically, a real image of a second pose corresponds to a second synthesized image rendered at the second pose.
[0343] S1203. Respectively intercept the regions at the same positions of the first object in the corresponding second synthesized images in each real image as the foreground images of each real image.
[0344] Specifically, in S1203, the foreground images of each real image input to the first network in S1201 are intercepted.
[0345] It should be understood that the region images at the same positions can refer to the region images with the same coordinates based on a certain point of the first object in the second synthesized image. In other words, the region images at the same positions can refer to the projection region images obtained by projecting the black-and-white image of the second synthesized image onto the real image based on a certain point of the first object in the second synthesized image.
[0346] In a possible implementation, the black-and-white image of the first object in the second synthesized image can be projected onto the real image corresponding to the second synthesized image, and the projection region is used as the foreground image of the real image.
[0347] In another possible implementation, the black-and-white image of the first object in the second synthesized image can be obtained first. The black-and-white image is a binary image. Multiply the black-and-white image of the first object in the second synthesized image by the real image corresponding to the second synthesized image, and use the remaining region obtained after multiplication as the foreground image of the real image.
[0348] S1204. Construct a first loss function according to the first difference information between the foreground images of N real images and their corresponding second synthesized images.
[0349] Among them, the first difference information is used to indicate the difference between the foreground image and its corresponding second synthesized image.
[0350] Optionally, the first difference information may include one or more of: the difference in feature maps, the difference in pixel colors, and the difference in extracted feature descriptors.
[0351] Among them, the difference between the feature maps of two images can be the perceptual loss, that is, the difference between the feature maps encoded by the deep learning network. Specifically, the deep learning network can use the pre-trained Visual Geometry Group (VGG) 16 network. For a given image, the feature map encoded by the VGG16 network is a tensor of C*H*W, where C is the number of channels of the feature map, and H and W are the length and width of the feature map. The difference between the feature maps is the distance between the two tensors. Specifically, it can be the L1 norm, L2 norm or other norms of the difference between the tensors.
[0352] The pixel color difference can be: the numerical calculation difference of the pixel colors, specifically, it can be the L2 norm of the difference between the pixels of two images.
[0353] The difference between the extracted feature descriptors: refers to the distance between the vectors representing the descriptors. For example, the distance can include but is not limited to: Euclidean distance, Hamming distance, etc. Specifically, the descriptor can be an N-dimensional floating-point vector, and the distance is the Euclidean distance between two N-dimensional vectors, and the L2 norm of the difference between two N-dimensional vectors; or the descriptor can be an M-dimensional binary vector, and the distance is the L1 norm of the difference between the two vectors.
[0354] The L2 norm mentioned above refers to taking the square root of the sum of the squares of each element of the vector, and the L1 norm refers to the sum of the absolute values of each element in the vector.
[0355] It should be noted that if N is greater than 1, the first difference information when constructing the first loss function in S1204 can be the calculated value of the difference information between multiple foreground images and their corresponding second synthesized images.
[0356] Among them, the second synthesized image corresponding to the foreground image refers to the true image to which the foreground image belongs. After obtaining the second pose through the first network, the second synthesized image is obtained through differentiable rendering.
[0357] Exemplarily, for the true image I (one or more) of the first object input by the user, the 6DOF pose of the first object in the true image is detected by the first network. Then, based on the detected 6DOF pose of the first object in the true image, the differentiable renderer can perform differentiable rendering to obtain the second synthesized image R (the same number as the true image I) according to the initial 3D model input by the user and the corresponding texture and lighting information. According to each second synthesized image R, the mask of the first object in each R is obtained (the same number as the true image I), and the foreground image F of the first object in the true image I is intercepted based on the mask (one mask is used to intercept the foreground image of its corresponding true image), and the first loss function Loss of F and R is constructed. 1 As follows: Loss 1 = L p + Li +L f 。
[0358] Among them, L p is the calculated value of the difference between the feature map of each foreground image and the corresponding second synthesized image, and L i is the calculated value of the pixel color difference between each foreground image and the corresponding second synthesized image, and L f is the calculated value of the difference between the feature descriptors of each foreground image and the corresponding second synthesized image.
[0359] It should be noted that the feature descriptors extracted from an image can be one or more. When the feature descriptors extracted from an image are multiple, the difference between the feature descriptors of two images can be the calculated value of the difference between the feature descriptors at the same position.
[0360] S1205. Update the differentiable renderer according to the first loss function, so that the synthesized image output by the differentiable renderer approximates the real image of the object.
[0361] Exemplarily, in S1205, the texture, lighting or other parameters in the differentiable renderer can be updated according to the first loss function constructed in S1204, so as to minimize the difference between the synthesized image output by the differentiable renderer and the real image.
[0362] Specifically, S1205 can adjust and update the texture, lighting or other parameters in the differentiable renderer according to a preset rule, and then repeat Figure 12 the schematic optimization process until the difference between the synthesized image output by the differentiable renderer and the real image is minimized.
[0363] Among them, the preset rule can be to adjust different parameters and adjustment values correspondingly when the first loss function is pre-configured to meet different conditions. After determining the first loss function in S1204, adjust according to the satisfied conditions and the corresponding adjustment parameters and adjustment values.
[0364] Exemplarily, Figure 12 the schematic optimization process can also be as Figure 13 shown. The real image of the object is input into the first network, the second pose output by the first network is input into the differentiable renderer, the differentiable renderer renders and outputs a second synthesized image and a black-and-white image of the object in the second synthesized image according to the three-dimensional model of the object, constructs a first loss function according to the second synthesized image and the real image, and updates the differentiable renderer using the first loss function.
[0365] Through the optimization method provided by this application, the differentiable renderer (which can also be referred to as a differentiable rendering engine or a differentiable rendering network) can be optimized to improve the rendering authenticity of the differentiable renderer and reduce the difference between the synthesized image obtained by differentiable rendering and the real image.
[0366] The optimization method provided in the second embodiment of this application can be specifically executed by a training device 520 as shown in Figure 5b The second synthesized image in the optimization method can be training data maintained in a database 530 as shown in Figure 5b Optionally, some or all of S1201 to S1203 in the optimization method provided in the second embodiment can be executed in the training device 520, or can be pre-executed by other functional modules before the training device 520, that is, the training data received or obtained from the database 530 is preprocessed first, and through the processes described in S1201 to S1203, a foreground image and a second synthesized image are obtained as the input of the training device 520, and S1204 to S1205 are executed by the training device 520.
[0367] Optionally, the optimization method provided in the second embodiment of this application can be processed by a CPU, or can be jointly processed by a CPU and a GPU, or a GPU can be not used and other processors suitable for neural network computing can be used, and this application does not make any restrictions.
[0368] Exemplarily, before obtaining the synthesized image, the object registration method provided by this application can adopt the Figure 12 schematic optimization method to optimize the differentiable renderer and reduce the difference between the synthesized image obtained by differentiable rendering and the real image. Figure 14 Another object registration method is schematically shown, and this method may include:
[0369] S1401. Optimize the differentiable renderer.
[0370] Among them, the specific operation of S1401 can refer to the Figure 12 schematic process and will not be elaborated here.
[0371] S1402. Input the three-dimensional model of the first object into the optimized differentiable renderer for differentiable rendering to obtain a first synthesized image.
[0372] S1403. Obtain a first input image according to the first synthesized image.
[0373] Among them, the specific implementation of S1403 can refer to S801b and will not be elaborated here.
[0374] S1404. Extract the feature information of each first input image respectively.
[0375] Among them, the specific implementation of S1404 refers to the aforementioned S802, which will not be elaborated here.
[0376] S1405: Register the feature information extracted from each first input image.
[0377] Among them, the specific implementation of S1405 refers to the aforementioned S803, which will not be elaborated here.
[0378] Embodiment 3 of the present application provides a method for training an object pose detection network. This method can train the aforementioned first network to improve the accuracy of the pose of the first object output by the first network. Exemplarily, the method for training the object pose detection network can be combined with Figure 8 the object registration method shown schematically, and the optimization method provided in Embodiment 2 of the present application, or can be used independently. The application scenarios of the present application are not limited to this.
[0379] Specifically, Embodiment 3 of the present application provides a method for training an object pose detection network. Through the real image and the synthetic image of the first object, the prediction ability (generalization) of the object pose detection network for the object pose in the real image is further optimized.
[0380] The method for training the object pose detection network provided in Embodiment 3 of the present application is as Figure 15 shown and includes:
[0381] S1501: Obtain a plurality of second input images including the first object.
[0382] Among them, the plurality of second input images include the real image of the first object and / or a plurality of third synthetic images of the first object.
[0383] In a possible implementation manner, the second input image may include a plurality of real images of the first object.
[0384] In another possible implementation manner, the second input image may include a plurality of real images of the first object and a plurality of third synthetic images of the first object. The plurality of third synthetic images are rendered from the three-dimensional model of the first object in a plurality of third poses. The plurality of third poses are different.
[0385] Exemplarily, the plurality of third poses may be taken on a spherical surface. The plurality of third poses may be Figure 9 a plurality of poses evenly sampled on the spherical surface shown. The present application does not limit the density of the sampled poses on the spherical surface, and can be selected according to actual needs.
[0386] In another possible implementation manner, the second input image may include a plurality of third synthetic images of the first object.
[0387] Specifically, in S1501, according to the three-dimensional model with texture information of the first object input by the user, through conventional rendering, differentiable rendering, or other rendering methods, 2D images of the first object at different angles can be synthesized at multiple third poses and recorded as the third synthesized images.
[0388] S1502: Input the multiple second input images into the second network for pose recognition respectively to obtain the fourth pose of the first object in each of the second input images output by the second network.
[0389] Among them, the second network can be used to recognize the pose of the first object in the image.
[0390] In a possible implementation, the second network can be a basic single-object pose detection network. The basic single-object pose detection network is a general neural network that only recognizes a single object in the image and is an initial configured model.
[0391] Specifically, the input of the basic single-object pose detection network is an image, and the output is the pose of the single object recognizable by the network in the image. Further optionally, the output of the basic single-object pose detection network can also include the black and white image (mask) of the single object recognizable by the network.
[0392] Among them, the mask of an object refers to the black and white form of the area in the image that only contains the object.
[0393] In another possible implementation, the second network can be a neural network trained according to the image of the first object on the basis of the basic single-object pose detection network.
[0394] Exemplarily, since the third synthesized images are rendered and generated at known third poses, the pose of the first object relative to the camera in each third synthesized image is known. Each third synthesized image can be input into the basic single-object pose detection network to obtain the predicted pose of the first object in each third synthesized image; then, according to the predicted pose and the actual pose of the first object in each third synthesized image, calculate the loss in the iterative process of the basic single-object pose detection network, construct a loss function, and further train the basic single-object pose detection network until convergence as the second network. It should be noted that the training process of the neural network in this application will not be elaborated.
[0395] In yet another possible implementation, the second network can be the currently used first network.
[0396] S1503: Obtain the fourth synthesized images of the first object at each fourth pose through differentiable rendering according to the three-dimensional model of the first object.
[0397] Among them, the second input image of a fourth pose corresponds to the fourth composite image rendered from the fourth pose.
[0398] S1504. Construct a second loss function according to the second difference information between each of the fourth composite images and its corresponding second input image.
[0399] Among them, the second difference information is used to indicate the difference between the fourth composite image and its corresponding second input image.
[0400] Exemplarily, the second difference information includes one or more of the following: the difference between the intersection over union (IOU) of the black and white image of the first object in the fourth composite image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose of the fourth composite image and the pose obtained by the fourth composite image passing through the first network, and the similarity between the fourth composite image and the regional image at the same position as the first object in the fourth composite image in its corresponding second input image.
[0401] Among them, the difference between the IOU of two black and white images: refers to the ratio of the area of the intersection of two masks to the area of the union.
[0402] The difference between two poses: refers to the calculation of the difference in the mathematical expressions of two poses. Specifically, the difference is the sum of the differences in translation and rotation. Given two poses as R 1 , T 1 and R 2 , T 2 , then the difference in translation is the Euclidean distance between the two vectors T 1 and T 2 ; the difference in rotation is where Tr represents the trace of the matrix, arccos is the inverse cosine function, and R 1 T is the transpose of R 1 .
[0403] The similarity between two images can be the pixel color difference or the feature map difference between the two images.
[0404] Among them, the regional image of the first object in the second input image can be obtained by cropping the second input image according to the black and white image of the first object in its corresponding fourth composite image. For example, the black and white image of the first object in the fourth composite image can be a binary image. Multiplying the binary image by the corresponding second input image can crop the regional image of the first object in the second input image.
[0405] Optionally, when the number in the second input image is M (M is greater than or equal to 2), there are also M fourth composite images, and there are M pieces of second difference information between the fourth composite image and its corresponding second input image, L i can be the calculated value of the second difference information between the fourth composite image and its corresponding second input image.
[0406] Specifically, in S1504, the differences between the poses of the second network on different second input images are calculated and constructed as the second loss function, so that the generalization of the second network on real images can be improved by optimizing these differences.
[0407] In a possible implementation, the second loss function Loss 2 satisfies the following expression:
[0408] where X is greater than or equal to 1; λ i is the weight, and L i is the calculated value used to represent the second difference information between the fourth composite image and its corresponding second input image. X is less than or equal to the number of types of the second difference information.
[0409] In a possible implementation, the first loss function Loss 2 can be Loss 2 = λ 1 L 1 + λ 2 L 2 + λ 3 L 3 .
[0410] where, L 1 is the calculated value of the difference in IoU between the black-and-white image of the first object in the second input image output by S1502 and the black-and-white image of the first object in the fourth composite image obtained in S1503; L 2 is the calculated value of the difference between the pose of the first object in the second input image output by S1502 and the pose of the first object detected in the fourth composite image using the second network; L 3 is the calculated value of the visual perception similarity of the part of the fourth composite image and its corresponding second input image that is the same region as the first object in the fourth composite image.
[0411] It should be understood that the preset weight λ i when constructing the second loss function can be configured according to actual needs, and the embodiments of the present application do not limit this.
[0412] S1505. Update the second network according to the second loss function to obtain the first network.
[0413] Among them, the difference between the pose of the first object in the image recognized by the first network and the true pose of the first object in the image is smaller than the difference between the pose of the first object in the image recognized by the second network and the true pose of the first object in the image.
[0414] It should be noted that the process of updating the second network according to the second loss function in S1505 to obtain the first network can refer to the training process of the neural network, which will not be elaborated in this embodiment of the present application.
[0415] Exemplarily, Figure 15 The process of the method for training the object pose detection network can also be as Figure 16 shown. The second input image composed of the real image of the object without annotation and the third synthesized image is input into the second network, and the fourth pose is output. The fourth pose and the three-dimensional model of the object are input into the differentiable renderer, and the fourth synthesized image and the corresponding mask are output. According to the fourth synthesized image and the corresponding mask, and the result of the second input image passing through the second network, a second loss function is constructed, and the second network is trained using the second loss function to obtain the first network.
[0416] Through the method for training the object pose detection network provided by the present application, using a single object pose detection network, even if a new recognizable object is added, the training time is short and it will not affect the recognition effect of other recognizable objects; in addition, by constructing a loss function using the difference between the real image of the object and the differentiable rendered images at multiple poses, the accuracy of the single object pose detection network is improved.
[0417] The method for training the object pose detection network provided in Embodiment 3 of the present application can be specifically executed by a training device 520 as shown in Figure 5b The second synthesized image in this optimization method can be the training data maintained in a database 530 as shown in Figure 5b . Optionally, some or all of S1501 to S1503 in the method for training the object pose detection network provided in Embodiment 3 can be executed in the training device 520, or can be pre-executed by other functional modules before the training device 520, that is, the training data received or obtained from the database 530 is preprocessed first, such as the process described in S1501 to S1503, to obtain the second input image and the fourth synthesized image as the input of the training device 520, and S1504 to S1505 are executed by the training device 520.
[0418] Optionally, the optimization method provided in Embodiment 2 of the present application can be processed by the CPU, or can be jointly processed by the CPU and the GPU, or without using the GPU, but using other processors suitable for neural network computing, which is not limited in the present application.
[0419] Embodiment 4 of the present application further provides a display method, which is applied to a terminal device. This display method can be used in combination with the foregoing object registration method, optimization method, and method for training an object pose detection network, or can be used alone. The embodiments of the present application do not specifically limit this.
[0420] As Figure 17 shown, the display method provided by the embodiments of the present application may include:
[0421] S1701. The terminal device acquires a first image.
[0422] Wherein, the first image refers to any image acquired by the terminal device.
[0423] In a possible implementation manner, the terminal device may acquire an image in the viewfinder by shooting with a camera device as the first image.
[0424] Exemplarily, the terminal device may be a mobile phone. A user may start an APP in the mobile phone and input an instruction to acquire an image in the APP interface, and the mobile phone starts the camera to acquire the first image.
[0425] Exemplarily, the terminal device may be smart glasses. After the user wears the smart glasses, an image in the field of view is captured through the viewfinder as the first image.
[0426] In another possible implementation manner, the terminal device may load an image stored locally as the first image.
[0427] Certainly, the embodiments of the present application do not limit the manner in which the terminal acquires the first image in S1701.
[0428] S1702. The terminal device determines whether the first image includes a recognizable object.
[0429] Specifically, the terminal determines whether the first image contains a recognizable object according to a network for recognizing objects configured offline.
[0430] Optionally, the network for recognizing objects may be the current pose detection network, or the first network or the second network described in the foregoing embodiments.
[0431] In a possible implementation manner, in S1702, the terminal device may determine whether the first image contains a recognizable object according to the current pose for recognizing an object and the identified network.
[0432] Exemplarily, the terminal device may input the first image into a network for recognizing the pose and identification of an object. If the network outputs the pose and identification of one or more objects, it is determined that the first image includes a recognizable object; otherwise, the first image does not include a recognizable object.
[0433] In another possible implementation, in S1702, the terminal device may determine whether the first image includes recognizable objects based on the first network described in the foregoing embodiments of the present application and the multi-object category classifier obtained by the registration object.
[0434] Exemplarily, the terminal device may extract the features of the first image and match them with the features in the multi-object category classifier described in the foregoing embodiments of the present application. If there are matching features, it is determined that the first image includes recognizable objects; otherwise, the first image does not include recognizable objects.
[0435] Exemplarily, in S1702, the feature information in the first image may be extracted, and the feature information is used to indicate the recognizable features in the first image; then it is determined whether there is feature information in the feature library whose matching distance with the extracted feature information meets the preset condition; if there is feature information in the feature library whose matching distance with the feature information meets the preset condition, it is determined that the first image includes one or more recognizable objects; if there is no feature information in the feature library whose matching distance with the feature information meets the preset condition, it is determined that the first image does not include any recognizable objects. Wherein, the feature library stores one or more pieces of feature information of different objects.
[0436] In one possible implementation, the preset condition may include being less than or equal to a preset threshold. The value of the preset threshold can be configured according to actual needs, and the embodiments of the present application do not limit this.
[0437] Optionally, if the terminal device in S1702 determines that the first image includes one or more recognizable objects, then S1703 and S1704 are executed. If the terminal device in S1702 determines that the first image does not include any recognizable objects, then S1705 is executed.
[0438] S1703: Output the first information, where the first information is used to prompt that recognizable objects are detected in the first image.
[0439] Optionally, the first information may be text information, voice information, or other forms, and the embodiments of the present application do not limit this.
[0440] Exemplarily, the content of the first information may be "Recognizable objects have been detected. Please keep the current angle", and the content of the first information may be superimposed and displayed on the first image in text form, or the content of the first information may be played through the speaker of the terminal device.
[0441] S1704: Obtain the poses of each recognizable object in the first image through the first network corresponding to each recognizable object included in the first image; display the virtual content corresponding to each recognizable object according to the pose of each recognizable object.
[0442] Among them, the virtual content corresponding to the object can be configured according to actual needs, and the embodiments of the present application do not limit it. For example, it can be the introduction information of the exhibition items, or it can also be the attribute information of the product, or others.
[0443] S1705. Output the second information, where the second information is used to prompt that no recognizable object is detected, and adjust the viewing angle to obtain a second image.
[0444] Among them, the second image is different from the first image.
[0445] Optionally, the second information can be text information, voice information, or other forms, and the embodiments of the present application do not limit this.
[0446] Exemplarily, the content of the second information can be "No recognizable object is detected. Please adjust the device angle". The content of the second information can be superimposed and displayed on the first image in text form, or the content of the second information can be played through the speaker of the terminal device.
[0447] Through the display method provided by the present application, a prompt is output to the user as to whether the image includes a recognizable object, so that the user can intuitively obtain whether the image includes a recognizable object, improving the user experience.
[0448] Exemplarily, in S1702, the process by which the terminal device determines whether the first image contains a recognizable object according to the first network described in the foregoing embodiments of the present application and the multi-object category classifier obtained by registering the object can be as Figure 18 shown, and this process may include:
[0449] S1801. The terminal device extracts local feature points in the first image.
[0450] It should be noted that the embodiments of the present application do not limit the specific scheme for extracting local feature points in S1801, but the method by which the terminal device extracts local feature points in S1801 needs to be consistent with that in S803 to ensure the accuracy of the scheme.
[0451] S1802. The terminal device obtains one or more first local feature points in the first image.
[0452] Among them, the matching distance between the descriptor of the first local feature point and the descriptor of the local feature point in the feature library is less than or equal to the first threshold. The value of the first threshold can be configured according to actual needs, and the embodiments of the present application do not limit this.
[0453] Among them, the matching distance between descriptors refers to the distance between the vectors representing the descriptors. For example, Euclidean distance, Hamming distance, etc.
[0454] Specifically, the feature library stores descriptors of local feature points of different objects. For example, the feature library can be the multi-object category classifier described in the foregoing embodiments.
[0455] S1803. The terminal device determines one or more ROIs in the first image according to the first local feature points.
[0456] Wherein, one ROI includes one object.
[0457] Specifically, the terminal device can classify the first local feature points into different objects by combining the color information and depth information of the image, and determine the area of the first local feature points of the same object that are concentratedly distributed as one ROI, so as to obtain one or more ROIs.
[0458] S1804. The terminal device extracts the global features in each ROI.
[0459] It should be noted that the specific solution for extracting the global features in S1804 in the embodiments of the present application is not limited, but the method for the terminal device to extract the global features in S1804 needs to be consistent with that in S803 to ensure the accuracy of the solution.
[0460] S1805. The terminal device determines whether the first image contains recognizable objects according to the global features in each ROI.
[0461] Exemplarily, if there are one or more first global features in the global features in each ROI, it is determined that the first image includes the recognizable object corresponding to the first global feature. Wherein, the matching distance between the descriptor of the first global feature and the descriptor of the global feature in the feature library is less than or equal to the second threshold. The feature library also stores descriptors of global features of different objects. If there is no first global feature in the global features in each ROI, it is determined that the first image does not include any recognizable object.
[0462] Wherein, the value of the second threshold can be configured according to actual needs, and the embodiments of the present application do not limit this.
[0463] The display method provided by the present application is described below through specific examples.
[0464] Suppose a certain merchant provides an APP. When a user shops, the APP captures an image through the mobile phone camera, recognizes the products included in the image, and displays the product introductions corresponding to the products. The three-dimensional objects of each product are registered offline according to the solution provided by the embodiments of the present application. When a certain user shops at this merchant, the scene captured by the APP in the mobile phone Figure 1 Such as Figure 19 As shown in the schematic mobile phone interface, the APP extracts the scene through Figure 18 The method shown in the schematic diagram extracts the sceneFigure 1 features in it to identify the scenario Figure 1 contains an identifiable object (smart speaker), and outputs "An identifiable object has been detected. Please keep the current angle" as shown Figure 20a in the mobile phone interface shown. Next, the APP calls the second network corresponding to the smart speaker to obtain the 6DOF pose of the smart speaker in this scenario Figure 1 and then displays the virtual content corresponding to the smart speaker "Hello everyone, I'm Xiaoyi, a Huawei smart speaker that can play music, tell stories, tell jokes, and also know countless encyclopedic knowledge" according to this 6DOF pose. The virtual and real combined display content is as shown Figure 20b in the mobile phone interface shown.
[0465] The above mainly introduces the solution provided by the embodiment of the present invention from the perspective of the working principle of the device. It can be understood that in order for electronic devices and the like to implement the above functions, they include the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, combining the units and algorithm steps of each example described in the embodiments disclosed in this article, the present invention can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0466] The embodiments of the present invention can divide the device for executing the method of the present application and the like into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0467] In the case of dividing each functional module corresponding to each function, Figure 21 it schematically shows an object registration device 210 provided by an embodiment of the present application, which is used to implement the functions in the first embodiment above. As shown Figure 21 in, the object registration device 210 may include: a first acquisition unit 2101, an extraction unit 2102, and a registration unit 2103. The first acquisition unit 2101 is used to execute the process S801 in Figure 8 , or S801a and S801b in Figure 10 , or S1402 and S1403 in Figure 14 ; the extraction unit 2102 is used to execute Figure 8 orFigure 10 the process S802 in Figure 14 or S1404 in Figure 8 or Figure 10 the process S803 in Figure 14 or S1405 in
[0468] In the case of dividing each function into corresponding function modules, Figure 22 FIG. shows an optimization device provided by an embodiment of the present application, which is used to implement the functions in the second embodiment above. As Figure 22 shown, the optimization device 220 may include: a processing unit 2201, a differentiable renderer 2202, a screenshot unit 2203, a construction unit 2204, and an update unit 2205. The processing unit 2201 is used to execute Figure 12 the process S1201 in Figure 12 the differentiable renderer 2202 is used to execute Figure 12 the process S1202 in Figure 12 the screenshot unit 2203 is used to execute Figure 12 the process S1203 in
[0469] In the case of dividing each function into corresponding function modules, Figure 23 FIG. shows a device 230 for training an object pose detection network provided by an embodiment of the present application, which is used to implement the functions in the third embodiment above. As Figure 23 shown, the device 230 for training an object pose detection network may include: a second acquisition unit 2301, a processing unit 2302, a differentiable renderer 2303, a construction unit 2304, and an update unit 2305. The second acquisition unit 2301 is used to execute Figure 15 the process S1501 in Figure 15 the processing unit 2302 is used to execute Figure 15 the process S1502 in Figure 15 the differentiable renderer 2303 is used to execute Figure 15 the process S1503 in
[0470] In the case of dividing each functional module corresponding to each function, Figure 24 FIG. 240 shows a display device 240 provided by an embodiment of the present application, which is used to implement the functions in the fourth embodiment above. As Figure 24 shown, the display device 240 may include: a first acquisition unit 2401, an output unit 1402, and a processing unit 1403. The first acquisition unit 2401 is used to execute Figure 17 the process S1701 in Figure 17 ; the output unit 1402 is used to execute Figure 17 the process S1703 or S1705 in
[0471] Figure 25 ; the processing unit 1403 is used to execute
[0472] the process S1704 in Figure 25 . Wherein, all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.
[0473] FIG. 250 provides a schematic hardware structure diagram of a device 250. The device 250 may be an object registration device provided by an embodiment of the present application, which is used to execute the object registration method provided by the first embodiment of the present application. Alternatively, the device 250 may be an optimization device provided by an embodiment of the present application, which is used to execute the optimization method provided by the second embodiment of the present application. Alternatively, the device 250 may be a device for training an object pose detection network provided by an embodiment of the present application, which is used to execute the method for training an object pose detection network provided by the third embodiment of the present application. Alternatively, the device 250 may be a display device provided by an embodiment of the present application, which is used to execute the display method provided by the fourth embodiment of the present application.
[0474] The processor 2502 may adopt a general - purpose central processing unit (CPU), a microprocessor, an application - specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits to execute relevant programs to implement the functions required by the units in the object registration device, optimization device, device for training an object pose detection network, and display device in the embodiments of the present application, or to execute the method provided in any one of Embodiments 1 to 4 of the method embodiments of the present application.
[0475] The processor 2502 may also be an integrated circuit chip with signal - processing capabilities. In the implementation process, each step of the method for training an object pose detection network in the present application may be completed by the integrated logic circuit in the hardware of the processor 2502 or instructions in software form. The above - mentioned processor 2502 may also be a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general - purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware decoding processor, or completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read - only memory, programmable read - only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 2501, and the processor 2502 reads the information in the memory 2501 and combines its hardware to complete the functions required by the units included in the object registration device, optimization device, device for training an object pose detection network, and display device in the embodiments of the present application, or to execute the method provided in any one of Embodiments 1 to 4 of the present application.
[0476] The communication interface 2503 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the device 250 and other devices or communication networks. For example, training data (such as the real images and synthetic images described in the method embodiments of the present application) may be obtained through the communication interface 2503.
[0477] The bus 2504 may include a path for transmitting information between various components of the device 250 (e.g., the memory 2501, the processor 2502, the communication interface 2503).
[0478] It should be understood that the first acquisition unit 2101, the extraction unit 2102, and the registration unit 2103 in the object registration device 210 are equivalent to the processor 2502 in the device 250. The processing unit 2201, the differentiable renderer 2202, the screenshot unit 2203, the construction unit 2204, and the update unit 2205 in the optimization device 220 are equivalent to the processor 2502 in the device 250. The second acquisition unit 2301, the processing unit 2302, the differentiable renderer 2303, the construction unit 2304, and the update unit 2305 in the device 230 for training the object pose detection network are equivalent to the processor 2502 in the device 250. The first acquisition unit 2401 and the processing unit 1403 in the display device 240 are equivalent to the processor 2502 in the device 250, and the output unit 1402 is equivalent to the communication interface 2503 in the device 250.
[0479] It should be noted that although Figure 25 the shown device 250 only shows the memory, the processor, and the communication interface, in the specific implementation process, those skilled in the art should understand that the device 250 also includes other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the device 250 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the device 250 may also only include the devices necessary for implementing the embodiments of the present application, and do not necessarily include Figure 25 all the devices shown in
[0480] It can be understood that the device 250 is equivalent to Figure 5b the training device 520 or the execution device 510 in
[0481] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid state drive (SSD).
[0482] In various embodiments of the present application, if there is no special description and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0483] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. In the written description of the present application, the character " / " generally means that the associated objects before and after are in an "or" relationship; in the formulas of the present application, the character " / " means that the associated objects before and after are in a "division" relationship.
[0484] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic.
Claims
1. An object registration method, characterized in that, the method includes: Obtaining a plurality of first input images including a first object, the plurality of first input images including a real image of the first object and / or a plurality of first synthetic images of the first object; the plurality of first synthetic images are obtained by differentiable rendering of a three-dimensional model of the first object in a plurality of first poses; the plurality of first poses are different; Respectively extracting feature information of the plurality of first input images, the feature information being used to indicate the features of the first object in the first input image where it is located; Corresponding the feature information extracted from each of the first input images with the identifier of the first object, and performing registration of the first object to obtain a multi-object category classifier; the multi-object category classifier records the corresponding relationship between the features and identifiers of different objects.
2. The method according to claim 1, characterized in that, the feature information includes descriptors of local feature points and descriptors of global features.
3. The method according to claim 1 or 2, characterized in that, Respectively extracting feature information of the plurality of first input images includes: Inputting the plurality of first input images into a first network respectively for pose recognition to obtain the pose of the first object in each of the first input images; wherein, the first network is used to recognize the pose of the first object in the image; Projecting the three-dimensional model of the first object onto each of the first input images according to the obtained pose of the first object to obtain a projection area in each of the first input images; Respectively extracting the feature information in the projection area of each of the first input images.
4. The method according to claim 1 or 2, characterized in that, Respectively extracting feature information of the plurality of first input images includes: Inputting the plurality of first input images into a first network respectively for black-and-white image extraction to obtain the black-and-white image of the first object in each of the first input images; wherein, the first network is used to extract the black-and-white image of the first object in the image; Respectively extracting the feature information in the black-and-white image of the first object in each of the first input images.
5. The method according to any one of claims 1-4, characterized in that, the method further includes: Inputting N real images of the first object into the first network respectively to obtain the second pose of the first object in each real image output by the first network; N is greater than or equal to 1; the first network is used to recognize the pose of the first object in the image; According to the three-dimensional model of the first object, using a differentiable renderer to perform differentiable rendering at each of the second poses to obtain N second synthetic images; obtaining that the real image of one second pose corresponds to the second synthetic image rendered at the one second pose; Respectively intercepting the area at the same position of the first object in the second synthetic image corresponding to the real object in each of the real images as the foreground image of each of the real images; Construct a first loss function according to the first difference information between the foreground image of N said real images and its corresponding second synthesized image; wherein, the first difference information is used to indicate the difference between the foreground image and its corresponding second synthesized image. Update the differentiable renderer according to the first loss function, so that the synthesized image output by the differentiable renderer approximates the real image of the object.
6. The method according to claim 5, wherein, The first difference information includes one or more of the following information: feature map difference, pixel color difference, difference between extracted feature descriptors.
7. The method according to any one of claims 1-6, wherein, The method further includes: Obtain a plurality of second input images including a first object, the plurality of second input images including the real image of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from the three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different; Input the plurality of second input images into a second network for pose recognition respectively, and obtain the fourth pose of the first object in each of the second input images output by the second network; the second network is used to recognize the pose of the first object in the image; According to the three-dimensional model of the first object, perform differentiable rendering to obtain a fourth synthesized image of the first object at each of the fourth poses; obtain that the second input image of one fourth pose corresponds to the fourth synthesized image rendered at the one fourth pose; Construct a second loss function according to the second difference information between each of the fourth synthesized images and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image; Update the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is less than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image.
8. The method according to claim 7, wherein, The second loss function Loss 2 satisfies the following expression: where X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image.
9. The method according to claim 7 or 8, wherein, The second difference information includes one or more of the following: the difference in the intersection over union (IOU) between the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose of obtaining the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image.
10. A display method, wherein, The method includes: Obtain a first image; Determine whether one or more recognizable objects are included in the first image according to the feature library; wherein, one or more feature information of different objects are stored in the feature library; the feature information is used to indicate the features of the object in the image; the feature library is a multi-object category classifier obtained for registered objects, and the registered objects include: extracting the feature information of multiple input images including the object, and registering the object by corresponding the identifier of the object with the feature information; the multiple input images include real images and / or multiple first synthetic images, and the multiple first synthetic images are obtained by differentiable rendering of the three-dimensional model of the object in multiple first poses; the multiple first poses are different; If one or more recognizable objects are included in the first image, output first information, where the first information is used to prompt that a recognizable object is detected in the first image; obtain the pose of each recognizable object in the first image through the pose detection network corresponding to each recognizable object included in the first image; and display the virtual content corresponding to each recognizable object according to the pose of each recognizable object; If no recognizable object is included in the first image, output second information, where the second information is used to prompt that no recognizable object is detected, and adjust the perspective to obtain a second image, where the second image is different from the first image.
11. The method according to claim 10, wherein, the method further includes: extracting the feature information in the first image, where the feature information in the first image is used to indicate the recognizable features in the first image; judging whether there is feature information in the feature library whose matching distance with the feature information in the first image meets a preset condition; if there is feature information in the feature library whose matching distance with the feature information in the first image meets a preset condition, determine that one or more recognizable objects are included in the first image; if there is no feature information in the feature library whose matching distance with the feature information in the first image meets a preset condition, determine that no recognizable object is included in the first image.
12. The method according to claim 10, wherein, the method further includes: obtaining one or more first local feature points in the first image, where the matching distance between the descriptor of the first local feature point and the descriptor of the local feature point in the feature library is less than or equal to a first threshold; the descriptor of the local feature point of different objects is stored in the feature library; determining one or more regions of interest (ROIs) in the first image according to the first local feature points; one object is included in one ROI; extracting the global feature in each ROI; if there is one or more first global features in the global features in each ROI, determine that the first image includes the recognizable object corresponding to the first global feature; wherein, the matching distance between the descriptor of the first global feature and the descriptor of the global feature in the feature library is less than or equal to a second threshold; the descriptor of the global feature of different objects is also stored in the feature library; If the first global feature does not exist in the global features of each of the ROIs, determine that the first image does not include any recognizable object.
13. The method according to any one of claims 10-12, wherein, the method further includes: obtaining a plurality of first input images including a first object, the plurality of first input images including a real image of the first object and / or a plurality of first synthetic images of the first object; the plurality of first synthetic images are obtained by differentiable rendering of a three-dimensional model of the first object at a plurality of first poses; the plurality of first poses are different; extracting feature information of the plurality of first input images respectively, the feature information being used to indicate the features of the first object in the first input image where it is located; storing the feature information extracted from each of the first input images corresponding to the identifier of the first object in a feature library, performing registration of the first object, and obtaining a multi-object category classifier; the multi-object category classifier records the corresponding relationship between the feature information of different objects and the identifiers.
14. The method according to claim 11 or 13, wherein, the feature information includes descriptors of local feature points and descriptors of global features.
15. The method according to any one of claims 10-14, wherein, the method further includes: obtaining a plurality of second input images including a first object, the plurality of second input images including a real image of the first object and / or a plurality of third synthetic images of the first object; the plurality of third synthetic images are obtained by rendering a three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different; inputting the plurality of second input images into a second network for pose recognition respectively, and obtaining a fourth pose of the first object in each of the second input images output by the second network; the second network is used to recognize the pose of the first object in the image; obtaining a fourth synthetic image of the first object at each of the fourth poses by differentiable rendering according to the three-dimensional model of the first object; obtaining that a second input image of one fourth pose corresponds to the fourth synthetic image rendered at the one fourth pose; constructing a second loss function according to second difference information between each of the fourth synthetic images and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthetic image and its corresponding second input image; updating the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is smaller than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image.
16. The method according to claim 15, wherein, The second loss function Loss 2 satisfies the following expression: where X is greater than or equal to 1, and λ i is a weight, and L i is a calculated value representing the second difference information between the fourth synthesized image and its corresponding second input image.
17. The method according to claim 15 or 16, wherein, The second difference information includes one or more of the following: the difference in the intersection over union (IOU) between the black-and-white image of the first object in the fourth synthesized image and the black-and-white image of the first object in the corresponding second input image, the difference between the fourth pose of the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in the corresponding second input image.
18. An object registration device, characterized in that, the device includes: A first acquisition unit, configured to acquire a plurality of first input images including a first object, where the plurality of first input images include the real image of the first object and / or a plurality of first synthesized images of the first object; the plurality of first synthesized images are obtained by differentiable rendering of the three-dimensional model of the first object at a plurality of first poses; the plurality of first poses are different; An extraction unit, configured to extract the feature information of the plurality of first input images respectively, where the feature information is used to indicate the features of the first object in the first input image where it is located; A registration unit, configured to correspond the feature information extracted by the extraction unit in each of the first input images with the identifier of the first object, and perform registration of the first object to obtain a multi-object category classifier; the multi-object category classifier records the corresponding relationship between the features and identifiers of different objects.
19. The device according to claim 18, characterized in that, the feature information includes the descriptor of local feature points and the descriptor of global features.
20. The device according to claim 18 or 19, characterized in that, the extraction unit is specifically configured to: Input the plurality of first input images into a first network respectively for pose recognition to obtain the pose of the first object in each of the first input images; where the first network is used to recognize the pose of the first object in the image; Project the three-dimensional model of the first object onto each of the first input images according to the obtained pose of the first object to obtain the projection area in each of the first input images; Extract the feature information in the projection area in each of the first input images respectively.
21. The device according to claim 18 or 19, characterized in that, the extraction unit is specifically configured to: Input the plurality of first input images into a first network respectively for black-and-white image extraction to obtain the black-and-white image of the first object in each of the first input images; where the first network is used to extract the black-and-white image of the first object in the image; Extract the feature information in the black-and-white image of the first object in each of the first input images respectively.
22. The device according to any one of claims 18-21, characterized in that, the device further includes: A processing unit for respectively inputting N real images of the first object into a first network to obtain the second pose of the first object in each real image output by the first network; N is greater than or equal to 1; the first network is used to identify the pose of the first object in the image; A differentiable renderer for, according to the three-dimensional model of the first object, differentiably rendering N second synthesized images at each of the second poses; obtaining that a real image of one second pose corresponds to the second synthesized image rendered at the one second pose; A truncation unit for respectively truncating, in each of the real images, a region at the same position of the first object in the second synthesized image corresponding to the real object as the foreground image of each of the real images; A construction unit for constructing a first loss function according to the first difference information between the foreground images of the N real images and their corresponding second synthesized images; wherein the first difference information is used to indicate the difference between the foreground image and its corresponding second synthesized image; An update unit for updating the differentiable renderer according to the first loss function, so that the synthesized image output by the differentiable renderer approximates the real image of the object.
23. The apparatus according to claim 22, wherein, the first difference information includes one or more of the following information: feature map difference, pixel color difference, difference between extracted feature descriptors.
24. The apparatus according to any one of claims 18-23, wherein, the apparatus further includes: A second acquisition unit for acquiring a plurality of second input images including the first object, the plurality of second input images including real images of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from the three-dimensional model of the first object at a plurality of third poses; the plurality of third poses are different; A processing unit for respectively inputting the plurality of second input images into a second network for pose recognition to obtain the fourth pose of the first object in each of the second input images output by the second network; the second network is used to identify the pose of the first object in the image; A differentiable renderer for, according to the three-dimensional model of the first object, differentiably rendering to obtain a fourth synthesized image of the first object at each of the fourth poses; obtaining that a second input image of one fourth pose corresponds to the fourth synthesized image rendered at the one fourth pose; A construction unit for constructing a second loss function according to the second difference information between each of the fourth synthesized images and their corresponding second input images; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image; An update unit for updating the second network according to the second loss function to obtain a first network; the difference between the pose of the first object identified by the first network in the image and the real pose of the first object in the image is smaller than the difference between the pose of the first object identified by the second network in the image and the real pose of the first object in the image.
25. The apparatus according to claim 24, wherein, The second loss function Loss 2 satisfies the following expression: where X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth synthesized image and its corresponding second input image.
26. The device according to claim 24 or 25, wherein, the second difference information includes one or more of the following: the difference in the intersection over union (IOU) between the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in the corresponding second input image, the difference between the fourth pose of the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in the corresponding second input image.
27. A display device, wherein, the device includes a first acquisition unit, a determination unit, an output unit, and a processing unit; wherein: the first acquisition unit is configured to acquire a first image; the determination unit is configured to determine, according to a feature library, whether one or more recognizable objects are included in the first image; wherein, one or more feature information of different objects are stored in the feature library; the feature information is used to indicate the features of the object in the image; the feature library is a multi-object category classifier obtained for registered objects, and the registered objects include: extracting the feature information of multiple input images including the object, and registering the object by corresponding the identifier of the object with the feature information; the multiple input images include real images and / or multiple first synthesized images, and the multiple first synthesized images are obtained by differentiable rendering of the three-dimensional model of the object at multiple first poses; the multiple first poses are different; the output unit is configured to, if one or more recognizable objects are included in the first image, output first information, where the first information is used to prompt that recognizable objects are detected in the first image; if no recognizable object is included in the first image, output second information, where the second information is used to prompt that no recognizable object is detected, and adjust the viewing angle so that the first acquisition unit acquires a second image, and the second image is different from the first image; the processing unit is configured to, if one or more recognizable objects are included in the first image, obtain the pose of each recognizable object in the first image through a pose detection network corresponding to each recognizable object included in the first image; and display virtual content corresponding to each recognizable object according to the pose of each recognizable object.
28. The device according to claim 27, wherein, the device further includes: an extraction unit configured to extract the feature information in the first image, where the feature information is used to indicate the recognizable features in the first image; a judgment unit configured to judge whether there is feature information in the feature library whose matching distance with the feature information meets a preset condition; wherein, one or more feature information of different objects are stored in the feature library; A first determination unit, configured to determine that the first image includes one or more recognizable objects if there is feature information in the feature library whose matching distance with the feature information meets a preset condition; and determine that the first image does not include any recognizable objects if there is no feature information in the feature library whose matching distance with the feature information meets the preset condition.
29. The apparatus according to claim 27, wherein, the apparatus further includes: A second acquisition unit, configured to acquire one or more first local feature points in the first image, and the matching distance between the descriptor of the first local feature point and the descriptor of the local feature point in the feature library is less than or equal to a first threshold; the feature library stores descriptors of local feature points of different objects; A second determination unit, configured to determine one or more regions of interest (ROIs) in the first image according to the first local feature points; one object is included in one ROI; An extraction unit, configured to extract global features in each of the ROIs; A first determination unit, configured to determine that the first image includes the recognizable object corresponding to the first global feature if there is one or more first global features in the global features of each of the ROIs; and determine that the first image does not include any recognizable objects if there is no first global feature in the global features of each of the ROIs; wherein, the matching distance between the descriptor of the first global feature and the descriptor of the global feature in the feature library is less than or equal to a second threshold; the feature library also stores descriptors of global features of different objects.
30. The apparatus according to any one of claims 27-29, wherein, the apparatus further includes: A third acquisition unit, configured to acquire a plurality of first input images including a first object, the plurality of first input images including a real image of the first object and / or a plurality of first synthetic images of the first object; the plurality of first synthetic images are obtained by differentiable rendering of a three-dimensional model of the first object in a plurality of first poses; the plurality of first poses are different; An extraction unit, configured to extract feature information of the plurality of first input images respectively, the feature information being used to indicate the features of the first object in the first input image where it is located; A registration unit, configured to store the feature information extracted from each of the first input images corresponding to the identifier of the first object in the feature library, perform registration of the first object, and obtain a multi-object category classifier; the multi-object category classifier records the corresponding relationship between the feature information and the identifier of different objects.
31. The apparatus according to claim 28 or 30, wherein, the feature information includes a descriptor of a local feature point and a descriptor of a global feature.
32. The apparatus according to any one of claims 27-31, wherein, the apparatus further includes: A fourth acquisition unit, configured to acquire a plurality of second input images including a first object, where the plurality of second input images include a real image of the first object and / or a plurality of third synthesized images of the first object; the plurality of third synthesized images are rendered from a three-dimensional model of the first object in a plurality of third poses; the plurality of third poses are different; A processing unit, configured to respectively input the plurality of second input images into a second network for pose recognition, so as to obtain a fourth pose of the first object in each of the second input images output by the second network; the second network is used to recognize the pose of the first object in the image; A differentiable renderer, configured to perform differentiable rendering according to a three-dimensional model of the first object to obtain a fourth synthesized image of the first object in each of the fourth poses; the second input image of one of the fourth poses corresponds to the fourth synthesized image rendered in the one of the fourth poses; A construction unit, configured to construct a second loss function according to second difference information between each of the fourth synthesized images and its corresponding second input image; the second difference information is used to indicate the difference between the fourth synthesized image and its corresponding second input image; An update unit, configured to update the second network according to the second loss function to obtain a first network; the difference between the pose of the first object recognized by the first network in the image and the real pose of the first object in the image is smaller than the difference between the pose of the first object recognized by the second network in the image and the real pose of the first object in the image.
33. The apparatus according to claim 32, wherein, The second loss function Loss 2 satisfies the following expression: where X is greater than or equal to 1, and λ i is a weight, and L i is used to represent the calculated value of the second difference information between the fourth composite image and its corresponding second input image.
34. The apparatus according to claim 32 or 33, wherein, The second difference information includes one or more of the following: the difference between the intersection over union (IOU) of the black and white image of the first object in the fourth synthesized image and the black and white image of the first object in its corresponding second input image, the difference between the fourth pose for obtaining the fourth synthesized image and the pose obtained by passing the fourth synthesized image through the first network, and the similarity between the fourth synthesized image and the regional image at the same position as the first object in the fourth synthesized image in its corresponding second input image.
35. An electronic device, wherein, the electronic device includes: a processor and a memory; the memory is connected to the processor; the memory is used to store computer instructions, and when the processor executes the computer instructions, the electronic device is caused to execute the object registration method according to any one of claims 1-9, or the electronic device is caused to execute the display method according to any one of claims 10-17.
36. A computer-readable storage medium, wherein, it includes instructions that, when running on a computer, cause the computer to execute the object registration method according to any one of claims 1-9, or cause the computer to execute the display method according to any one of claims 10-17.
37. A computer program product, wherein, When it runs on a computer, it causes the computer to execute the object registration method described in any one of claims 1-9, or causes the computer to execute the display method described in any one of claims 10-17.
Citation Information
Patent Citations
Virtual content display method and device, electronic equipment and computer readable medium
CN111510701A
Network training method and device and attitude prediction method and device
CN111783986A