Method and device for generating three-dimensional virtual model
By a method of obtaining subject images from target images and obtaining component images from subject images, a lightweight feature extractor and mask information acquisition model is used to solve the problems of complex background interference and artificial errors in the prior art, and the quality and accuracy of the three-dimensional virtual model are improved.
Patent Information
- Application Number
- CN202111148991.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-28
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2041-09-28
AI Technical Summary
When generating a three-dimensional virtual model of furniture, it is difficult to effectively remove complex content interference in the target image other than the subject object, resulting in low image quality of component images, affecting the matching degree and accuracy of the three-dimensional virtual model, and manual participation is prone to introduce errors.
By first obtaining the subject image of the subject object from the target image and then obtaining the component image from the subject image, the model is obtained using a lightweight feature extractor and mask information to automatically intercept the subject and component images to avoid manual intervention.
It improves the quality of component images and the matching degree of the three-dimensional virtual model, reduces labor costs, reduces the impact of human errors, and improves the accuracy and efficiency of model generation.
Smart Images

Figure CN114140574B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and device for generating a three-dimensional virtual model. Background Art
[0002] In the home furnishing scene, there is usually a need to generate a three-dimensional virtual model of a furniture object based on a two-dimensional image of the furniture object. In this way, buyers can fully understand the furniture based on the three-dimensional virtual model of the furniture object, thereby improving the promotion effect of the furniture.
[0003] When a user (such as a seller, etc.) needs to obtain a three-dimensional virtual model of a furniture object, the user can manually capture an image including the furniture object, then input the captured image into the electronic device, and control the electronic device to generate a three-dimensional virtual model of the furniture object based on the image.
[0004] In order to make the generated three-dimensional virtual model of the furniture object match the actual furniture object, it is necessary to obtain images of various component objects in the furniture object in the image. Summary of the Invention
[0005] This application describes a method and apparatus for generating a three-dimensional virtual model.
[0006] In a first aspect, the present application shows a method for generating a three-dimensional virtual model, which includes: when it is necessary to generate a three-dimensional virtual model of a subject object, obtaining a target image, wherein the target image includes the subject object, and the subject object includes at least one component object; obtaining a subject image of the subject object based on the target image; obtaining a component image of the component object based on the subject image; and generating a three-dimensional virtual model of the subject object based on the component image.
[0007] In a second aspect, the present application shows a method for generating a three-dimensional virtual model, the method comprising: acquiring a captured target image, the target image including the furniture object, and the furniture object including at least one component object; acquiring a furniture image of the furniture object based on the target image; acquiring a component image of the component object based on the furniture image; and generating a three-dimensional virtual model of the furniture object based on the component image.
[0008] In a third aspect, the present application shows a device for generating a three-dimensional virtual model, which includes: a first acquisition module, used to acquire a target image when a three-dimensional virtual model of a subject object needs to be generated, wherein the target image includes the subject object, and the subject object includes at least one component object; a second acquisition module, used to acquire a subject image of the subject object based on the target image; a third acquisition module, used to acquire a component image of the component object based on the subject image; and a first generation module, used to generate a three-dimensional virtual model of the subject object based on the component image.
[0009] In a fourth aspect, the present application shows a device for generating a three-dimensional virtual model, the device comprising: a fourth acquisition module for acquiring a captured target image, the target image including the furniture object, and the furniture object including at least one component object; a fifth acquisition module for acquiring a furniture image of the furniture object based on the target image; a sixth acquisition module for acquiring a component image of the component object based on the furniture image; and a second generation module for generating a three-dimensional virtual model of the furniture object based on the component image.
[0010] In a fifth aspect, the present application shows an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the method shown in any of the aforementioned aspects.
[0011] In a sixth aspect, the present application shows a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the method shown in any of the aforementioned aspects.
[0012] In a seventh aspect, the present application shows a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the method shown in any of the aforementioned aspects.
[0013] Compared with the prior art, this application has the following advantages:
[0014] In the present application, when a three-dimensional virtual model of a subject object needs to be generated, if a target image including the subject object is acquired and the subject object includes at least one component object, a main image of the subject object can be first acquired based on the target image, and then component images of the component objects can be acquired based on the main image of the subject object, and then the three-dimensional virtual model of the subject object can be generated based on the component images. In this way, the component images of the component objects are directly acquired from the main image of the subject object, rather than directly from the target image.
[0015] The subject image only includes the content of the subject object, and does not include the content other than the subject object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the subject image of the subject object". Since there is no interference, compared with directly obtaining the component image of the component object of the subject object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the subject image of the subject object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the subject object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the subject object and the actual subject object can be improved.
[0016] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the main object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and thus can avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the main object being affected by real-time human errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a scenario diagram of this application.
[0018] Figure 2 This is a flowchart of the steps of a method for generating a three-dimensional virtual model in the present application.
[0019] Figure 3 This is a flowchart of the steps of a method for obtaining a subject image in the present application.
[0020] Figure 4 This is a flowchart of the steps of a method for obtaining a component image in the present application.
[0021] Figure 5 This is a flowchart of the steps of a method for generating a three-dimensional virtual model in the present application.
[0022] Figure 6 This is a structural block diagram of a device for generating a three-dimensional virtual model in the present application.
[0023] Figure 7 This is a structural block diagram of a device for generating a three-dimensional virtual model in the present application.
[0024] Figure 8 This is a structural block diagram of a device of the present application. DETAILED DESCRIPTION
[0025] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0026] Reference Figure 1 The solution of the present application is summarized in a schematic diagram. The solution of the present application can be applied to electronic devices, which may include terminals or servers. Terminals may include mobile phones, tablet computers, laptop computers, or desktop computers.
[0027] In which, when it is necessary to generate a three-dimensional virtual model of the subject object, the target image including the subject object can be input into a lightweight feature extractor to extract the target image features of the target image using the lightweight feature extractor, and then the target image features of the target image can be input into a subject mask information acquisition model to obtain the subject mask information of the subject object in the target image using the subject mask information acquisition model, and then the subject image of the subject object can be intercepted in the target image according to the subject mask information of the subject object in the target image.
[0028] In addition, the subject image of the subject object can be input into a lightweight feature extractor to extract the subject image features of the subject image using the lightweight feature extractor, and then the subject image features of the subject image can be input into a component mask information acquisition model to obtain the component mask information of the component object in the subject image using the component mask information acquisition model, and then the component image of the component object can be intercepted in the subject image according to the component mask information, and then a three-dimensional virtual model of the subject object can be generated according to the component image of the component object.
[0029] It can be seen that the component image of the component object is directly obtained from the main image of the main object, rather than directly obtained from the target image.
[0030] The subject image only includes the content of the subject object, and does not include the content other than the subject object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the subject image of the subject object". Since there is no interference, compared with directly obtaining the component image of the component object of the subject object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the subject image of the subject object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the subject object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the subject object and the actual subject object can be improved.
[0031] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the main object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and thus can avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the main object being affected by real-time human errors.
[0032] Reference Figure 2 , shows a flowchart of the steps of a method for generating a three-dimensional virtual model of the present application, which may specifically include the following steps:
[0033] In step S101 , a target image is acquired, where the target image includes a main object, and the main object includes at least one component object.
[0034] In the present application, the target image may be a two-dimensional image, for example, a two-dimensional RGB (Red Green Blue) image, or a PNG (Portable Network Graphics) or GIF (Graphics Interchange Format) image.
[0035] The target image includes at least one main object, which may be a furniture object, such as a stool object, a table object, a nightstand object, a bed object, a TV cabinet object, or a wardrobe object. Of course, other types of objects may also be used. For example, in a 3D reconstruction scene of a building, the main object may be a building object, or in a 3D reconstruction scene of a person, the main object may be a person object, etc. This application does not limit this.
[0036] Each main object includes at least one component object. In one possible case, the main object may also include more than two component objects. For example, when the main object includes a wardrobe object, the wardrobe object includes different types of component objects such as hardware objects, cabinet door objects, side panel objects, and base objects.
[0037] In one possible scenario, a user may need to obtain a three-dimensional virtual model of a subject object for display or promotion of the subject object.
[0038] If a user needs to obtain a 3D virtual model of a subject object, the user can capture a target image including the subject object or download the target image including the subject object from the Internet. The user can then input the target image including the subject object into an electronic device, which then generates a 3D virtual model of the subject object based on the target image. For example, the electronic device can obtain a texture map corresponding to the texture material of the subject object based on the target image, generate a 3D white model of the subject object, and then map the texture map corresponding to the texture material of the subject object onto the 3D white model to obtain the 3D virtual model of the subject object.
[0039] Among them, the texture materials of the outer surfaces of different component objects in different main objects on the market are different. The texture material can be reflected at least by the texture style and the texture color.
[0040] Therefore, in order to make the texture material presented on the outer surface of the generated three-dimensional virtual model of the main object as similar as possible to the texture material actually presented on the outer surface of each component object in the main object, after the electronic device obtains the target image including the main object input by the user, when the electronic device obtains the texture map corresponding to the texture material of the main object, it needs to obtain the texture map corresponding to the texture material of the component object in the main object in the target image. Then, when mapping the texture map corresponding to the texture material of the main object on the three-dimensional white model, the texture map corresponding to the texture material of the component object in the main object can be mapped to the position of the component object on the three-dimensional white model.
[0041] In order to enable the electronic device to obtain the texture map corresponding to the texture material of the component object in the main object in the target image, the component image of the component object in the main object can be obtained according to the target image, and then the texture map corresponding to the texture material of the component object can be obtained according to the component image.
[0042] In order to obtain a component image of a component object in a main object, in the present application, the electronic device may acquire a target image and then execute step S102.
[0043] In step S102, a main image of the main object is acquired according to the target image.
[0044] Among them, the main image of the main object can be intercepted from the target image, for example, the main mask information of the main object in the target image can be obtained, and then the main image can be intercepted in the target image with the help of the main mask information. For details, please refer to the following Figure 3 The embodiment shown will not be described in detail here.
[0045] In step S103 , a component image of the component object is acquired based on the main image of the main object.
[0046] Among them, the component image of the component object can be intercepted from the main image of the main object. For example, the component mask information of the component object in the main image can be obtained, and then the component image can be intercepted in the main image with the help of the component mask information. For details, please refer to the following Figure 4 The embodiment shown will not be described in detail here.
[0047] In step S104 , a three-dimensional virtual model of the subject object is generated based on the component image of the component object.
[0048] In an optional embodiment, after obtaining the component image of the component object in the main image of the main object in the target image, a three-dimensional virtual model of the main object can be generated based on the component image of the component object. For example, the material characteristics of the texture material of the component object can be obtained based on the component image of the component object, and then the texture map corresponding to the texture material of the component object can be obtained based on the material characteristics of the texture material of the component object to generate a three-dimensional white model of the main object, and then the texture map corresponding to the texture material of the component object is mapped to the position of the component object on the three-dimensional white model of the main object, thereby obtaining the three-dimensional model of the main object, that is, completing the three-dimensional modeling of the main object.
[0049] Wherein, when it is necessary to obtain a component image of a component object in a main object according to a target image, in one approach, the target image may be directly analyzed to obtain the component image of the component object.
[0050] However, usually, the target image includes a lot of content and a complex background. There are many contents in the target image other than the main object. In this case, it will interfere with the process of "obtaining the component image of the component object by analyzing the target image". Due to the interference, the quality (such as accuracy) of "obtaining the component image of the component object directly by analyzing the target image" is often low. For example, the degree of match between the obtained component image of the component object and the edge contour of the component object in the target image is low.
[0051] In the present application, when it is necessary to generate a three-dimensional virtual model of a main object, if a target image including the main object is obtained and the main object includes at least one component object, the main image of the main object can be obtained based on the target image first, and then the component image of the component object can be obtained based on the main image of the main object, and then the three-dimensional virtual model of the main object can be generated based on the component image.
[0052] In this way, the component image of the component object is directly obtained from the main image of the main object, rather than directly obtained from the target image.
[0053] The subject image only includes the content of the subject object, and does not include the content other than the subject object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the subject image of the subject object". Since there is no interference, compared with directly obtaining the component image of the component object of the subject object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the subject image of the subject object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the subject object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the subject object and the actual subject object can be improved.
[0054] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the main object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and thus can avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the main object being affected by real-time human errors.
[0055] In one embodiment of the present application, see Figure 3 , step S102 includes:
[0056] In step S201 , subject mask information of a subject object is acquired according to a target image.
[0057] In one embodiment of the present application, a subject mask information acquisition model may be used to acquire subject mask information of a subject object. An example of a specific acquisition process is given below, which may include:
[0058] 2011. Using a subject mask information acquisition model to predict at least one candidate mask information of a subject object and a predicted probability value of each candidate mask information according to a target image.
[0059] In the present application, in order to obtain at least one candidate mask information of the main object in the target image and the predicted probability value of each candidate mask information, a main mask information acquisition model can be trained in advance, and then based on the main mask information acquisition model, at least one candidate mask information of the main object and the predicted probability value of each candidate mask information can be predicted according to the target image.
[0060] The subject mask information acquisition model can be obtained by training a sample dataset including "sample images and annotated mask information of sample subject objects in the sample images". An example of a specific training process is given below, which may include:
[0061] 11) Acquire at least one sample data set, where the sample data set includes: a sample image and annotation mask information of a sample subject object in the sample image.
[0062] There may be multiple sample data sets, and the sample images included in different sample data sets may be different images. The sample images in the sample data sets may include sample subject objects, which may include at least one sample component object, and the sample subject objects may include furniture objects, etc.
[0063] The sample image may include a two-dimensional RGB image, etc.
[0064] 12) Use the sample data set to train the network parameters in the model until the network parameters converge to obtain the main body mask information acquisition model.
[0065] The model may include CNN (Convolutional Neural Networks) or SOLOv2 (Instance Segmentation), etc. Of course, it may also include other types of models, which is not limited in this application.
[0066] In one embodiment of the present application, the network structure of a subject mask information acquisition model may include: a lightweight feature extraction network and a mask detection network. The lightweight extraction network is used to obtain image features of an image including a subject object. The mask detection network is used to obtain a subject image of the subject object in the image based on the image features. The input of the subject mask information acquisition model includes the input of the lightweight extraction network. The output of the lightweight extraction network is connected to the input of the mask detection network. The output of the subject mask information acquisition model includes the output of the mask detection network.
[0067] In one embodiment, the lightweight feature extraction network may include a lightweight feature extractor, and the lightweight feature extractor includes: Mobilenet (a lightweight convolutional neural network that can be used on mobile devices) or Shufflenet (a lightweight convolutional neural network that can be used on mobile devices), etc.
[0068] In another embodiment, the lightweight feature extraction network also includes a global feature descriptor. The global feature descriptor is used to expand the image features of the subject object image obtained by the lightweight feature extractor to obtain expanded image features at different scales of the image. This can improve the comprehensiveness and accuracy of the image features ultimately extracted from the image. For example, it can at least improve the comprehensiveness and accuracy of the image features extracted regarding details in the image.
[0069] The input end of the lightweight feature extraction network includes the input end of the lightweight feature extractor, and the output end of the lightweight feature extractor is connected to the input end of the global feature descriptor.
[0070] The output end of the lightweight feature extraction network includes the output end of the global feature descriptor. Alternatively, the output end of the lightweight feature extraction network includes the output end of the global feature descriptor and the output end of the lightweight feature extractor.
[0071] Furthermore, in another embodiment of the present application, the subject mask information acquisition model further includes a category detection network. The output end of the lightweight feature extractor is further connected to the input end of the category detection network.
[0072] In one embodiment, the output of the class detection network may be connected to the input of the mask detection network.
[0073] In another embodiment, the output end of the subject mask information acquisition model may also include the output end of the category detection network.
[0074] In this application, based on different actual needs, the network structure of the subject mask information acquisition model can be different, and the subject mask information acquisition models with different network structures can then be applied to different application scenarios. That is, the network structure of the subject mask information acquisition model suitable for different application scenarios is different.
[0075] During the training process, a sample image can be input into the model so that the model can process the sample image to obtain the mask information of the sample main object in the predicted sample image. After that, the loss function can be used to adjust the network parameters in the model based on the mask information of the sample main object in the predicted sample image and the annotated mask information of the sample image until the network parameters in the model converge. The training can be completed and the obtained main mask information acquisition model can be put into use online.
[0076] In this way, the target image can be input into the trained subject mask information acquisition model so that the subject mask information acquisition model processes the target image to obtain at least one candidate mask information of the subject object in the target image and the predicted probability value of each candidate mask information, and outputs at least one candidate mask information of the sample subject object in the target image and the predicted probability value of each candidate mask information. The electronic device can obtain at least one candidate mask information of the sample subject object in the target image and the predicted probability value of each candidate mask information output by the subject mask information acquisition model.
[0077] In one embodiment of the present application, when using a subject mask information acquisition model to predict at least one candidate mask information of a subject object and a predicted probability value for each candidate mask information based on a target image, a lightweight feature extraction network in the subject mask information acquisition model may be used to extract target image features of the target image. At least a mask detection network in the subject mask information acquisition model is used to predict at least one candidate mask information of the subject object and a predicted probability value for each candidate mask information based on the target image features.
[0078] In one embodiment, the lightweight feature extraction network in the subject mask information acquisition model includes a lightweight feature extractor, which enables rapid image feature extraction. The lightweight feature extractor has a simple structure and includes a small number of parameters, ensuring that the feature extractor meets lightweight requirements. This reduces video memory usage, making the subject mask information acquisition model adaptable to a variety of devices with different performance characteristics, thereby improving the adaptability of the subject mask information acquisition model. In one example, the subject mask information acquisition model can be adapted for use on low-performance devices.
[0079] The target image can be input into the lightweight feature extractor of the lightweight feature extraction network, so that the lightweight feature extractor of the lightweight feature extraction network processes the target image, obtains the target image features of the target image, and outputs the target image features of the target image.
[0080] In another embodiment of the application, the lightweight feature extraction network includes a lightweight feature extractor and a global feature descriptor.
[0081] Thus, when the lightweight feature extraction network in the subject mask information acquisition model is used to extract target image features of the target image, the lightweight feature extractor can be used to extract reference image features of the target image. For example, the target image can be input into the lightweight feature extractor of the lightweight feature extraction network, so that the lightweight feature extractor of the lightweight feature extraction network processes the target image, obtains reference image features of the target image, and outputs the reference image features of the target image.
[0082] The reference image features may be represented in the form of a matrix or a vector.
[0083] Then, the reference image feature can be expanded based on the global feature descriptor to obtain at least one expanded image feature.
[0084] In this application, the global feature descriptor includes at least one of the following: a cascaded pyramid network or a pooling network, etc.
[0085] The pooling network includes at least one of the following: a maximum pooling network, a minimum pooling network, a sum pooling network, and an average pooling network.
[0086] In one embodiment, when there is only one global feature descriptor, the reference image feature of the component object can be expanded using the global feature descriptor to obtain an expanded image feature.
[0087] In another embodiment, when there are two or more global feature descriptors, any one of the two or more global feature descriptors can be used to expand the reference image features of the subject image to obtain an expanded image feature. The above operation is similarly performed for each of the other two or more global feature descriptors, thereby obtaining two or more expanded image features.
[0088] In this embodiment, by expanding the reference image features of the target image based on at least one global feature descriptor, it is possible to obtain extended image features of the target image at different scales, which can improve the comprehensiveness and accuracy of the target image features finally extracted. For example, it can at least improve the comprehensiveness and accuracy of the image features extracted about the details in the target image, and provide a stable and reliable basis for the subsequent acquisition of the component image of the component object.
[0089] Then, a target image feature of the target image may be obtained based at least on the at least one extended image feature.
[0090] In one embodiment of the present application, if a reference image feature of a target image is expanded based on a global feature descriptor to obtain an extended image feature, the obtained extended image feature can be determined as the target image feature of the target image, or the obtained extended image feature can be fused with the reference image feature of the target image to obtain the target image feature of the target image.
[0091] In another embodiment of the present application, if the reference image features of the target image are expanded separately based on two or more different global feature descriptors to obtain two or more different extended image features, the two or more different extended image features obtained can be fused, and the fused features can be used as the target image features of the target image, or the two or more different extended image features obtained can be fused with the reference image features of the target image to obtain the target image features of the target image.
[0092] In one example, when the image features include vectors, a method of fusing two or more image features includes: connecting the two or more vectors end to end in sequence to obtain a large vector to achieve fusion.
[0093] In one embodiment of the present application, the target image features of the target image can be input into the mask detection network in the subject mask information acquisition model, so that the mask detection network in the subject mask information acquisition model processes the target image features to obtain at least one candidate mask information of the subject object and the predicted probability value of each candidate mask information, and outputs at least one candidate mask information of the subject object and the predicted probability value of each candidate mask information.
[0094] In another embodiment of the present application, the subject mask information acquisition model also includes a category detection network.
[0095] In this way, when at least the mask detection network in the subject mask information acquisition model is used to predict at least one candidate mask information of the subject object and the predicted probability value of each candidate mask information based on the target image features, the category detection network can be used to detect the subject category of the subject object. For example, the target image features of the target image can be input into the category detection network in the subject mask information acquisition model, so that the category detection network in the subject mask information acquisition model processes the target image features, obtains the subject category of the subject object, and outputs the subject category of the subject object, for example, outputs the subject category of the subject object to the mask detection network. In this way, the mask detection network can obtain the target image features of the target image and the subject category of the subject object. In this way, the mask detection network can be used to predict at least one candidate mask information of the subject object and the predicted probability value of each candidate mask information based on the target image features and the subject category.
[0096] Among them, for the mask detection network, compared with predicting at least one candidate mask information of the main object and the predicted probability value of each candidate mask information based on the target image features, predicting at least one candidate mask information of the main object and the predicted probability value of each candidate mask information based on the target image features and the main category can improve the accuracy of the predicted candidate mask information and the predicted probability value of each candidate mask information.
[0097] In the present application, when a piece of candidate mask information is obtained, the obtained piece of candidate mask information can be determined as the main mask information.
[0098] When more than two candidate mask information are obtained, step S2012 may be executed.
[0099] 2012. Obtain an intersection-over-union ratio between the candidate mask information with the maximum prediction probability value and at least part of the candidate mask information in at least one candidate mask information except the candidate mask information with the maximum prediction probability value.
[0100] In the present application, each candidate mask information occupies a partial position area in the target image. When calculating the intersection-and-union ratio between two candidate mask information, the area of the intersection of the partial position areas occupied by the two candidate mask information in the target image can be calculated, as well as the area of the union of the partial position areas occupied by the two candidate mask information in the target image. Then, the ratio between the area of the intersection and the area of the union is calculated as the intersection-and-union ratio between the two candidate mask information.
[0101] In one embodiment of the present application, an intersection-over-union ratio between the candidate mask information with the maximum prediction probability value and each candidate mask information in the at least one candidate mask information except the candidate mask information with the maximum prediction probability value may be obtained.
[0102] To further improve the efficiency of acquiring component images of component objects when there are a large number of candidate mask information, another embodiment of the present application can filter out candidate mask information with predicted probability values greater than a preset probability value from at least one candidate mask information, excluding the candidate mask information with the highest predicted probability value. The intersection-over-union ratio (IoU) is then calculated between the candidate mask information with the highest predicted probability value and each of the selected candidate mask information. This reduces the number of candidate mask information involved in calculating the IoU ratio, thereby reducing the computational effort and further improving the efficiency of acquiring component images of component objects.
[0103] 2013. Select, from at least part of the candidate mask information, candidate mask information whose intersection-over-union ratio with the candidate mask information having the largest prediction probability value is less than a preset intersection-over-union ratio.
[0104] The preset intersection-over-union ratio may include 0.7, 0.75, or 0.8, etc., and may be determined according to actual conditions, and is not limited in this application.
[0105] 2014. Obtain subject mask information according to the selected candidate mask information.
[0106] In step S202 , a subject image is captured from the target image according to the subject mask information.
[0107] In another embodiment of the present application, a target image may be displayed on a screen. When the target image is displayed, the subject edge contour of the subject object on the target image may be obtained based on the subject mask information. The subject edge contour and input controls may be displayed on the target image.
[0108] This allows users to view and verify whether the captured subject edge contour is correct. If there is any inaccuracy, the user can perform secondary editing on the subject edge contour based on the input control. Upon receiving the editing operation on the subject edge contour input according to the input control, the subject edge contour is corrected according to the editing operation to obtain the corrected subject edge contour; the subject mask information is corrected according to the corrected subject edge contour to obtain the corrected subject mask information.
[0109] Accordingly, when the subject image is intercepted in the target image according to the subject mask information, the subject image may be intercepted in the target image according to the corrected subject mask information to improve the accuracy of the obtained subject image.
[0110] In one embodiment of the present application, see Figure 4 , step S103 includes:
[0111] In step S301 , component mask information of a component object is acquired based on a subject image.
[0112] In one embodiment of the present application, a component mask information acquisition model may be used to acquire component mask information of a component object. An example of a specific acquisition process is given below, which may include:
[0113] 3011. Use the component mask information acquisition model to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the subject image.
[0114] In the present application, in order to obtain at least one candidate mask information of a component object in a subject image and a predicted probability value of each candidate mask information, a component mask information acquisition model can be trained in advance, and then based on the component mask information acquisition model, at least one candidate mask information of a component object and a predicted probability value of each candidate mask information can be predicted according to the subject image.
[0115] The component mask information acquisition model can be obtained by training a sample dataset including "sample images and annotated mask information of sample component objects in sample main objects in the sample images". An example of a specific training process is given below, which may include:
[0116] 31) Acquire at least one sample data set, the sample data set including: a sample image and annotation mask information of a sample component object in a sample main object in the sample image.
[0117] There may be multiple sample data sets, and the sample images included in different sample data sets may be different images. The sample images in the sample data sets may include sample subject objects, which may include at least one sample component object, and the sample subject objects may include furniture objects, etc.
[0118] The sample image may include a two-dimensional RGB image, etc.
[0119] 32) Use the sample data set to train the network parameters in the model until the network parameters converge to obtain the component mask information acquisition model.
[0120] The model may include CNN or SOLOv2, etc. Of course, it may also include other types of models, which is not limited in this application.
[0121] In one embodiment of the present application, the network structure of a component mask information acquisition model may include: a lightweight feature extraction network and a mask detection network. The lightweight feature extraction network is used to obtain subject image features of a subject image of a subject object, which includes a component object. The mask detection network is used to obtain component images of component objects in the subject image based on the subject image features. The input end of the component mask information acquisition model includes the input end of the lightweight extraction network. The output end of the lightweight extraction network is connected to the input end of the mask detection network. The output end of the component mask information acquisition model includes the output end of the mask detection network.
[0122] In one embodiment, the lightweight feature extraction network may include a lightweight feature extractor.
[0123] In another embodiment, the lightweight feature extraction network also includes a global feature descriptor.
[0124] The global feature descriptor is used to expand the subject image features of the subject image of the subject object obtained by the lightweight feature extractor to obtain extended image features of the subject image at different scales, which can improve the comprehensiveness and accuracy of the image features of the subject image finally extracted. For example, it can at least improve the comprehensiveness and accuracy of the image features extracted about the details in the subject image.
[0125] The input end of the lightweight feature extraction network includes the input end of the lightweight feature extractor, and the output end of the lightweight feature extractor is connected to the input end of the global feature descriptor.
[0126] The output end of the lightweight feature extraction network includes the output end of the global feature descriptor. Alternatively, the output end of the lightweight feature extraction network includes the output end of the global feature descriptor and the output end of the lightweight feature extractor.
[0127] Furthermore, in another embodiment of the present application, the component mask information acquisition model further includes a category detection network. The output end of the lightweight feature extractor is further connected to the input end of the category detection network.
[0128] In one embodiment, the output of the class detection network may be connected to the input of the mask detection network.
[0129] In another embodiment, the output end of the component mask information acquisition model may also include the output end of the category detection network.
[0130] In this application, based on different actual needs, the network structure of the component mask information acquisition model can be different, and the component mask information acquisition models with different network structures can then be applied to different application scenarios. That is, the network structure of the component mask information acquisition model suitable for different application scenarios is different.
[0131] During the training process, a sample image can be input into the model so that the model processes the sample image to obtain the mask information of the sample component object in the sample main object in the predicted sample image. After that, the loss function can be used to adjust the network parameters in the model based on the mask information of the sample component object in the sample main object in the predicted sample image and the labeled mask information of the sample component object in the sample main object in the sample image until the network parameters in the model converge. The training can be completed, and the obtained component mask information acquisition model can be put into use online.
[0132] In this way, the target image can be input into the trained component mask information acquisition model so that the component mask information acquisition model processes the target image to obtain at least one candidate mask information of the component object in the main object in the target image and the predicted probability value of each candidate mask information, and outputs at least one candidate mask information of the component object in the main object in the target image and the predicted probability value of each candidate mask information. The electronic device can obtain at least one candidate mask information of the component object in the main object in the target image and the predicted probability value of each candidate mask information output by the component mask information acquisition model.
[0133] In one embodiment of the present application, when using a component mask information acquisition model to predict at least one candidate mask information of a component object and a predicted probability value of each candidate mask information based on a subject image, a lightweight feature extraction network in the component mask information acquisition model can be used to extract subject image features of the subject image, and then at least a mask detection network in the component mask information acquisition model can be used to predict at least one candidate mask information of a component object and a predicted probability value of each candidate mask information based on the subject image features.
[0134] In one embodiment, the lightweight feature extraction network in the component mask information acquisition model includes a lightweight feature extractor, which enables rapid image feature extraction. The lightweight feature extractor has a simple structure and includes a small number of parameters, enabling the feature extractor to meet lightweight requirements. This reduces video memory usage, making the component mask information acquisition model adaptable to a variety of devices with different performance characteristics, thereby improving the adaptability of the component mask information acquisition model. In one example, the component mask information acquisition model can be adapted for use on low-performance devices.
[0135] The subject image can be input into the lightweight feature extractor of the lightweight feature extraction network, so that the lightweight feature extractor of the lightweight feature extraction network processes the subject image, obtains the subject image features of the subject image, and outputs the subject image features of the subject image.
[0136] In another embodiment of the application, the lightweight feature extraction network includes a lightweight feature extractor and a global feature descriptor.
[0137] Thus, when the lightweight feature extraction network in the component mask information acquisition model is used to extract subject image features of the subject image, the lightweight feature extractor can be used to extract reference image features of the subject image. For example, the subject image can be input into the lightweight feature extractor of the lightweight feature extraction network, so that the lightweight feature extractor of the lightweight feature extraction network processes the subject image, obtains reference image features of the subject image, and outputs the reference image features of the subject image.
[0138] The reference image features may be represented in the form of a matrix or a vector.
[0139] Then, the reference image feature can be expanded based on the global feature descriptor to obtain at least one expanded image feature.
[0140] In this application, the global feature descriptor includes at least one of the following: a cascaded pyramid network or a pooling network, etc.
[0141] The pooling network includes at least one of the following: a maximum pooling network, a minimum pooling network, a sum pooling network, and an average pooling network.
[0142] In one embodiment, when there is only one global feature descriptor, the reference image feature of the component object can be expanded using the global feature descriptor to obtain an expanded image feature.
[0143] In another embodiment, when there are two or more global feature descriptors, any one of the two or more global feature descriptors can be used to expand the reference image features of the target image to obtain an expanded image feature. The above operation is similarly performed for each of the other two or more global feature descriptors, thereby obtaining two or more expanded image features.
[0144] In this embodiment, by expanding the reference image features of the subject image based on at least one global feature descriptor, it is possible to obtain extended image features of the subject image at different scales, which can improve the comprehensiveness and accuracy of the subject image features of the subject object finally extracted. For example, it can at least improve the comprehensiveness and accuracy of the image features of the details extracted from the subject image, providing a stable and reliable basis for the subsequent acquisition of the component image of the component object.
[0145] Then, the subject image feature of the subject image may be obtained based on at least the at least one extended image feature.
[0146] In one embodiment of the present application, if the reference image feature of the subject image is expanded based on a global feature descriptor to obtain an extended image feature, the obtained extended image feature can be determined as the subject image feature of the subject image, or the obtained extended image feature can be fused with the reference image feature of the subject image to obtain the subject image feature of the subject image.
[0147] In another embodiment of the present application, if the reference image features of the subject image are expanded separately based on two or more different global feature descriptors to obtain two or more different extended image features, the two or more different extended image features obtained can be fused, and the fused features can be used as the subject image features of the subject image, or the two or more different extended image features obtained can be fused with the reference image features of the subject image to obtain the subject image features of the subject image.
[0148] In one embodiment of the present application, the subject image features of the subject image can be input into the mask detection network in the subject mask information acquisition model, so that the mask detection network in the subject mask information acquisition model processes the subject image features to obtain at least one candidate mask information of the component object and the predicted probability value of each candidate mask information, and outputs at least one candidate mask information of the component object and the predicted probability value of each candidate mask information.
[0149] In another embodiment of the present application, the subject mask information acquisition model also includes a category detection network.
[0150] In this way, when at least the mask detection network in the component mask information acquisition model is used to predict at least one candidate mask information of the component object and the predicted probability value of each candidate mask information based on the subject image features, the category detection network can be used to detect the component category of the component object. For example, the subject image features of the subject image can be input into the category detection network in the component mask information acquisition model, so that the category detection network in the component mask information acquisition model processes the subject image features to obtain the component category of the component object and outputs the component category of the component object, for example, outputting the component category of the component object to the mask detection network. In this way, the mask detection network can obtain the subject image features of the subject image and the component category of the component object in the subject image. In this way, the mask detection network can be used to predict at least one candidate mask information of the component object and the predicted probability value of each candidate mask information based on the subject image features and the component category.
[0151] Among them, for the mask detection network, compared with predicting at least one candidate mask information of a component object and the predicted probability value of each candidate mask information based on the main image features, predicting at least one candidate mask information of a component object and the predicted probability value of each candidate mask information based on the main image features and the component category can improve the accuracy of the predicted candidate mask information and the predicted probability value of each candidate mask information.
[0152] In the present application, when a piece of candidate mask information is obtained, the obtained piece of candidate mask information can be determined as component mask information.
[0153] When more than two candidate mask information are obtained, step S3012 may be executed.
[0154] 3012. Obtain intersection-over-union ratios between the candidate mask information with the maximum prediction probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the maximum prediction probability value.
[0155] In one embodiment of the present application, an intersection-over-union ratio between the candidate mask information with the maximum prediction probability value and each candidate mask information in the at least one candidate mask information except the candidate mask information with the maximum prediction probability value may be obtained.
[0156] To further improve the efficiency of acquiring component images of component objects when there are a large number of candidate mask information, another embodiment of the present application can filter out candidate mask information with predicted probability values greater than a preset probability value from at least one candidate mask information, excluding the candidate mask information with the highest predicted probability value. The intersection-over-union ratio (IoU) is then calculated between the candidate mask information with the highest predicted probability value and each of the selected candidate mask information. This reduces the number of candidate mask information involved in calculating the IoU ratio, thereby reducing the computational effort and further improving the efficiency of acquiring component images of component objects.
[0157] 3013. Select, from the at least a portion of the candidate mask information, a candidate mask information whose intersection-over-union ratio with the candidate mask information having the largest predicted probability value is less than a preset intersection-over-union ratio.
[0158] 3014. Obtain the component mask information according to the selected candidate mask information.
[0159] In step S302 , a component image is captured from the main image according to the component mask information.
[0160] In another embodiment of the present application, when a target image is displayed on a screen, the component edge contour of the component object on the main image on the target image can be obtained based on the component mask information, and the component edge contour and input control can be displayed on the main image on the target image.
[0161] This allows users to view and verify whether the captured component edge contour is correct. If there is any inaccuracy, the user can perform secondary editing on the component edge contour based on the input control. When receiving the editing operation on the component edge contour input according to the input control, the component edge contour is corrected according to the editing operation to obtain the corrected component edge contour; the component mask information is corrected according to the corrected component edge contour to obtain the corrected component mask information.
[0162] Accordingly, when the component image is intercepted in the main image according to the component mask information, the component image may be intercepted in the main image according to the corrected component mask information to improve the accuracy of the obtained component image.
[0163] Reference Figure 5 , shows a flowchart of the steps of a method for generating a three-dimensional virtual model of the present application, which may specifically include the following steps:
[0164] In step S401, a captured target image is acquired, where the target image includes a furniture object, and the furniture object includes at least one component object.
[0165] This step can refer to the description of the above embodiment and will not be described in detail here.
[0166] In step S402, a furniture image of the furniture object is acquired according to the target image.
[0167] This step can refer to the description of the above embodiment and will not be described in detail here.
[0168] In step S403, a component image of the component object is acquired based on the furniture image.
[0169] This step can refer to the description of the above embodiment and will not be described in detail here.
[0170] In step S404, a three-dimensional virtual model of the furniture object is generated according to the component image.
[0171] This step can refer to the description of the above embodiment and will not be described in detail here.
[0172] In the present application, if a target image including a furniture object is captured and the furniture object includes at least one component object, a furniture image of the furniture object can be first obtained based on the target image, and then a component image of the component object can be obtained based on the furniture image of the furniture object, and then a three-dimensional virtual model of the furniture object can be generated based on the component image.
[0173] In this way, the component image of the component object is directly obtained from the furniture image of the furniture object, rather than directly obtained from the target image.
[0174] The furniture image only includes the content of the furniture object, and does not include the content other than the furniture object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the furniture image of the furniture object". Since there is no interference, compared with directly obtaining the component image of the component object of the furniture object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the furniture image of the furniture object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the furniture object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the furniture object and the actual furniture object can be improved.
[0175] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the furniture object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and further avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the furniture object being affected by real-time human errors.
[0176] It should be noted that for the method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0177] Reference Figure 6 , shows a structural block diagram of a device for generating a three-dimensional virtual model of the present application, which may specifically include the following modules:
[0178] The first acquisition module 11 is used to acquire a target image when it is necessary to generate a three-dimensional virtual model of a main object, wherein the target image includes the main object, and the main object includes at least one component object; the second acquisition module 12 is used to acquire a main image of the main object based on the target image; the third acquisition module 13 is used to acquire a component image of the component object based on the main image; and the first generation module 14 is used to generate a three-dimensional virtual model of the main object based on the component image.
[0179] In an optional implementation, the second acquisition module includes: a first acquisition submodule, used to acquire the subject mask information of the subject object according to the target image; and a first interception submodule, used to intercept the subject image in the target image according to the subject mask information.
[0180] In an optional implementation, the second acquisition module also includes: a second acquisition submodule, which is used to obtain the subject edge contour of the subject object on the target image according to the subject mask information when the target image is displayed; a first display submodule, which is used to display the subject edge contour and input control on the target image; a first correction submodule, which is used to correct the subject edge contour according to the editing operation input according to the input control when an editing operation on the subject edge contour is received to obtain a corrected subject edge contour; and a second correction submodule, which is used to correct the subject mask information according to the corrected subject edge contour to obtain a corrected subject mask information.
[0181] Correspondingly, the first interception submodule includes: a first interception unit, configured to intercept the subject image in the target image according to the corrected subject mask information.
[0182] In an optional implementation, the first acquisition submodule includes: a first prediction unit, used to use a subject mask information acquisition model to predict at least one candidate mask information of the subject object and a predicted probability value of each candidate mask information according to the target image; a first acquisition unit, used to obtain the intersection-and-union ratio between the candidate mask information with the largest predicted probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest predicted probability value; a first selection unit, used to select, from the at least part of the candidate mask information, the candidate mask information whose intersection-and-union ratio with the candidate mask information with the largest predicted probability value is less than a preset intersection-and-union ratio; and a second acquisition unit, used to obtain the subject mask information based on the selected candidate mask information.
[0183] In an optional implementation, the first prediction unit includes: a first extraction subunit, used to extract target image features of the target image using a lightweight feature extraction network in the subject mask information acquisition model; a first acquisition subunit, used to predict at least one candidate mask information of the subject object and a prediction probability value of each candidate mask information based on the target image features using at least a mask detection network in the subject mask information acquisition model.
[0184] In an optional implementation, the lightweight feature extraction network includes a lightweight feature extractor and a global feature descriptor; the first extraction subunit is specifically used to: use the lightweight feature extractor to extract the reference image features of the target image; expand the reference image features based on the global feature descriptor to obtain at least one extended image feature; and obtain the target image features based on at least one extended image feature.
[0185] In an optional implementation, the subject mask information acquisition model also includes a category detection network; the first acquisition subunit is specifically used to: use the category detection network to detect the subject category of the subject object; use the mask detection network to predict at least one candidate mask information of the subject object and the predicted probability value of each candidate mask information based on the target image features and the subject category.
[0186] In an optional implementation, the first acquisition unit includes: a second acquisition sub-unit, configured to obtain intersection-and-union ratios (IORs) between the candidate mask information having the largest prediction probability value and each of the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest prediction probability value; or a third acquisition sub-unit, configured to screen candidate mask information having a prediction probability value greater than a preset probability value from the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest prediction probability value; and obtain IORs between the candidate mask information having the largest prediction probability value and each of the screened candidate mask information.
[0187] In an optional implementation, the third acquisition module includes: a third acquisition submodule, used to acquire component mask information of the component object based on the main image; and a second interception submodule, used to intercept the component image in the main image based on the component mask information.
[0188] In an optional implementation, the third acquisition module further includes: a fourth acquisition submodule for acquiring, when the target image is displayed, a component edge contour of the component object on the main image on the target image according to the component mask information; a second display submodule for displaying the component edge contour and an input control on the main image on the target image; a third correction submodule for, when receiving an edit operation on the component edge contour input according to the input control, correcting the component edge contour according to the edit operation to obtain a corrected component edge contour; and a fourth correction submodule for correcting the component mask information according to the corrected component edge contour to obtain the corrected component mask information;
[0189] Correspondingly, the second clipping submodule includes: a second clipping unit, configured to clip the component image in the main image according to the corrected component mask information.
[0190] In an optional implementation, the third acquisition submodule includes: a second prediction unit, used to use a component mask information acquisition model to predict at least one candidate mask information of the component object and a predicted probability value of each candidate mask information based on the main image; a third acquisition unit, used to obtain the intersection-and-union ratio between the candidate mask information with the largest predicted probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest predicted probability value; a second selection module, used to select, from the at least part of the candidate mask information, the candidate mask information whose intersection-and-union ratio with the candidate mask information with the largest predicted probability value is less than a preset intersection-and-union ratio; and a fourth acquisition unit, used to obtain the component mask information based on the selected candidate mask information.
[0191] In an optional implementation, the second prediction unit includes: a second extraction subunit, used to extract the main image features of the main image using the lightweight feature extraction network in the component mask information acquisition model; and a fourth acquisition subunit, used to predict at least one candidate mask information of the component object and the prediction probability value of each candidate mask information based on the main image features using at least the mask detection network in the component mask information acquisition model.
[0192] In an optional implementation, the lightweight feature extraction network includes a lightweight feature extractor and a global feature descriptor; the second extraction subunit is specifically used to: use the lightweight feature extractor to extract the reference image features of the subject image; expand the reference image features based on the global feature descriptor to obtain at least one extended image feature; and obtain the subject image features based on at least one extended image feature.
[0193] In an optional implementation, the component mask information acquisition model also includes a category detection network; the first acquisition subunit is specifically used to: use the category detection network to detect the component category of the component object; use the mask detection network to obtain at least one candidate mask information of the component object and the predicted probability value of each candidate mask information based on the subject image features and the component category.
[0194] In an optional implementation, the third acquisition unit includes: a fifth acquisition subunit, used to obtain the intersection-and-union ratios between the candidate mask information with the largest prediction probability value and each candidate mask information in the at least one candidate mask information except the candidate mask information with the largest prediction probability value; or, a sixth acquisition subunit, used to screen candidate mask information with a prediction probability value greater than a preset probability value from the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest prediction probability value; and obtain the intersection-and-union ratios between the candidate mask information with the largest prediction probability value and each of the screened candidate mask information.
[0195] In the present application, when a three-dimensional virtual model of a subject object needs to be generated, if a target image including the subject object is acquired and the subject object includes at least one component object, a main image of the subject object can be first acquired based on the target image, and then component images of the component objects can be acquired based on the main image of the subject object, and then the three-dimensional virtual model of the subject object can be generated based on the component images. In this way, the component images of the component objects are acquired directly from the main image of the subject object, rather than directly from the target image.
[0196] The subject image only includes the content of the subject object, and does not include the content other than the subject object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the subject image of the subject object". Since there is no interference, compared with directly obtaining the component image of the component object of the subject object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the subject image of the subject object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the subject object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the subject object and the actual subject object can be improved.
[0197] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the main object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and thus can avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the main object being affected by real-time human errors.
[0198] Reference Figure 7 , shows a structural block diagram of a device for generating a three-dimensional virtual model of the present application, which may specifically include the following modules:
[0199] The fourth acquisition module 21 is used to acquire a captured target image, wherein the target image includes the furniture object, and the furniture object includes at least one component object; the fifth acquisition module 22 is used to acquire a furniture image of the furniture object based on the target image; the sixth acquisition module 23 is used to acquire a component image of the component object based on the furniture image; and the second generation module 24 is used to generate a three-dimensional virtual model of the furniture object based on the component image.
[0200] In this application, if a target image including a furniture object is captured and the furniture object includes at least one component object, a furniture image of the furniture object can be first acquired based on the target image, and then a component image of the component object can be acquired based on the furniture image of the furniture object, and then a three-dimensional virtual model of the furniture object can be generated based on the component image. In this way, the component image of the component object is acquired directly from the furniture image of the furniture object, rather than directly from the target image.
[0201] The furniture image only includes the content of the furniture object, and does not include the content other than the furniture object in the target image. In this way, there is no other complex content to interfere with the process of "obtaining the component image of the component object based on the furniture image of the furniture object". Since there is no interference, compared with directly obtaining the component image of the component object of the furniture object from the target image, the quality of the component image of the component object obtained by obtaining the component image of the component object from the furniture image of the furniture object can be improved. For example, the matching degree between the obtained component image of the component object and the edge contour of the component object in the target image can be improved, etc., and then the quality of the generated three-dimensional virtual model of the furniture object can be improved. For example, the matching degree between the generated three-dimensional virtual model of the furniture object and the actual furniture object can be improved.
[0202] On the other hand, the process of obtaining the component image of the component object can be done without human participation, thereby reducing labor costs. Due to human physiological characteristics, people will inevitably make mistakes when participating in the process of obtaining the component image of the component object (for example, manually circling the edge contour of the component object of the furniture object in the target image, etc.). Since the present application can support the process of obtaining the component image of the component object without human participation, the present application can avoid the efficiency and quality (for example, accuracy, etc.) of obtaining the component image of the component object being affected by real-time human errors, and further avoid the efficiency and quality (for example, accuracy, etc.) of the generated three-dimensional virtual model of the furniture object being affected by real-time human errors.
[0203] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0204] An embodiment of the present application further provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.
[0205] The present application provides one or more machine-readable media having instructions stored thereon, which, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In the present application, the electronic device includes a server, a gateway, a sub-device, etc., and the sub-device is an IoT device or other device.
[0206] The embodiments of the present disclosure may be implemented as an apparatus configured as desired using any appropriate hardware, firmware, software, or any combination thereof, which may include a server (cluster), terminal devices such as IoT devices, and other electronic devices.
[0207] Figure 8 An exemplary apparatus 1300 that can be used to implement various embodiments described in this application is schematically illustrated.
[0208] For one embodiment, Figure 8 An exemplary apparatus 1300 is shown having one or more processors 1302, a control module (chip set) 1304 coupled to at least one of the processor(s) 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.
[0209] The processor 1302 may include one or more single-core or multi-core processors, and the processor 1302 may include any combination of general-purpose processors or dedicated processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, the apparatus 1300 can serve as a server device such as a gateway described in the embodiments of the present application.
[0210] In some embodiments, the apparatus 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage 1308) having instructions 1314 and one or more processors 1302 configured in conjunction with the one or more computer-readable media to execute the instructions 1314 to implement a module to perform the actions described in the present disclosure.
[0211] For one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1302 and / or any suitable device or component in communication with the control module 1304 .
[0212] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0213] The memory 1306 can be used, for example, to load and store data and / or instructions 1314 for the device 1300. For one embodiment, the memory 1306 can include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 1306 can include double data rate quad synchronous dynamic random access memory (DDR4 SDRAM).
[0214] For one embodiment, control module 1304 may include one or more input / output controllers to provide interfaces to NVM / storage device 1308 and input / output device(s) 1310 .
[0215] For example, NVM / storage 1308 may be used to store data and / or instructions 1314. NVM / storage 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).
[0216] NVM / storage device 1308 may include storage resources that are physically part of the device on which apparatus 1300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 1308 may be accessible over a network via input / output device(s) 1310.
[0217] (One or more) input / output devices 1310 may provide an interface for apparatus 1300 to communicate with any other appropriate device. Input / output devices 1310 may include a communication component, a phonetic component, a sensor component, etc. Network interface 1312 may provide an interface for apparatus 1300 to communicate via one or more networks. Apparatus 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.
[0218] For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers of the control module 1304 (e.g., a memory controller module). For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers of the control module 1304 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304 to form a system-on-chip (SoC).
[0219] In various embodiments, the apparatus 1300 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the apparatus 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the apparatus 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0220] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enable the electronic device to execute one or more methods described in the present application.
[0221] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0222] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0223] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the processes and / or boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable information processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0224] These computer program instructions may also be stored in a computer readable memory that can guide a computer or other programmable information processing terminal device to work in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0225] These computer program instructions can also be loaded onto a computer or other programmable information processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0226] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0227] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0228] The above is a detailed introduction to the method and device for obtaining component images provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for generating a three-dimensional virtual model, characterized in that: The method comprises: In a case where a three-dimensional virtual model of a subject object needs to be generated, a target image is acquired, wherein the target image includes the subject object, and the subject object includes at least one component object; Acquire a main image of the main object according to the target image; acquiring a component image of the component object according to the main image; generating a three-dimensional virtual model of the subject object according to the component image; The step of acquiring the component image of the component object according to the main body image includes: Acquire component mask information of the component object according to the subject image; intercepting the component image in the main image according to the component mask information; The acquiring component mask information of the component object according to the main image includes: Using a component mask information acquisition model to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the subject image; Obtaining intersection-over-union ratios between the candidate mask information with the maximum prediction probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the maximum prediction probability value; Selecting, from the at least part of the candidate mask information, candidate mask information whose intersection-over-union ratio with the candidate mask information having the largest predicted probability value is less than a preset intersection-over-union ratio; The component mask information is acquired according to the selected candidate mask information.
2. The method according to claim 1, characterized in that The method further comprises: When the target image is displayed, obtaining a component edge contour of the component object on the subject image on the target image according to the component mask information; displaying the component edge outline and input controls on the subject image on the target image; When receiving an editing operation on the edge contour of the component input according to the input control, correcting the edge contour of the component according to the editing operation to obtain a corrected edge contour of the component; Correcting the component mask information according to the corrected component edge contour to obtain corrected component mask information; Accordingly, intercepting the component image in the main image according to the component mask information includes: The component image is cut out from the main image according to the corrected component mask information.
3. The method according to claim 1, characterized in that The using the component mask information acquisition model to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the subject image includes: Extracting subject image features of the subject image using a lightweight feature extraction network in the component mask information acquisition model; At least a mask detection network in the component mask information acquisition model is used to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the subject image feature.
4. The method according to claim 3, characterized in that The lightweight feature extraction network includes a lightweight feature extractor and a global feature descriptor; The extracting the subject image features of the subject image using the lightweight feature extraction network in the component mask information acquisition model includes: extracting reference image features of the subject image using the lightweight feature extractor; Expanding the reference image feature based on the global feature descriptor to obtain at least one expanded image feature; The subject image feature is obtained based on at least the at least one extended image feature.
5. The method according to claim 3, characterized in that The component mask information acquisition model also includes a category detection network; The method of at least using the mask detection network in the component mask information acquisition model to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the subject image feature includes: Detecting the component category of the component object using the category detection network; The mask detection network is used to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the main image features and the component category.
6. A method for generating a three-dimensional virtual model, characterized in that: The method comprises: Acquire a captured target image, wherein the target image includes a furniture object, and the furniture object includes at least one component object; acquiring a furniture image of the furniture object according to the target image; acquiring a component image of the component object according to the furniture image; generating a three-dimensional virtual model of the furniture object based on the component image; The step of acquiring the component image of the component object according to the furniture image includes: acquiring component mask information of the component object according to the furniture image; intercepting the component image in the furniture image according to the component mask information; The acquiring component mask information of the component object according to the furniture image includes: Using a component mask information acquisition model to predict at least one candidate mask information of the component object and a prediction probability value of each candidate mask information according to the furniture image; Obtaining intersection-over-union ratios between the candidate mask information with the maximum prediction probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the maximum prediction probability value; Selecting, from the at least part of the candidate mask information, candidate mask information whose intersection-over-union ratio with the candidate mask information having the largest predicted probability value is less than a preset intersection-over-union ratio; The component mask information is acquired according to the selected candidate mask information.
7. A device for generating a three-dimensional virtual model, characterized in that: The device comprises: A first acquisition module is configured to acquire a target image when a three-dimensional virtual model of a subject object needs to be generated, wherein the target image includes the subject object, and the subject object includes at least one component object; A second acquisition module is used to acquire a main image of the main object according to the target image; a third acquisition module, configured to acquire a component image of the component object based on the main image; A first generating module, configured to generate a three-dimensional virtual model of the subject object according to the component image; The third acquisition module includes: a third acquisition submodule for acquiring component mask information of the component object according to the main image; a second interception submodule for intercepting the component image in the main image according to the component mask information; The third acquisition submodule includes: a second prediction unit, configured to use a component mask information acquisition model to predict at least one candidate mask information of the component object and a predicted probability value of each candidate mask information based on the main image; a third acquisition unit, configured to obtain an intersection-and-union (IoU) between the candidate mask information with the largest predicted probability value and at least a portion of the at least one candidate mask information other than the candidate mask information with the largest predicted probability value; a second selection module, configured to select, from the at least portion of the candidate mask information, candidate mask information whose IoU ratio with the candidate mask information with the largest predicted probability value is less than a preset IoU ratio; and a fourth acquisition unit, configured to obtain the component mask information based on the selected candidate mask information.
8. A device for generating a three-dimensional virtual model, characterized in that: The device comprises: a fourth acquisition module, configured to acquire a captured target image, wherein the target image includes a furniture object, and the furniture object includes at least one component object; a fifth acquisition module, configured to acquire a furniture image of the furniture object according to the target image; A sixth acquisition module, configured to acquire a component image of the component object based on the furniture image; a second generating module, configured to generate a three-dimensional virtual model of the furniture object based on the component image; The sixth acquisition module includes: a third acquisition submodule for acquiring component mask information of the component object according to the furniture image; a second interception submodule for intercepting the component image in the furniture image according to the component mask information; The third acquisition submodule includes: a second prediction unit, used to use a component mask information acquisition model to predict at least one candidate mask information of the component object and a predicted probability value of each candidate mask information based on the furniture image; a third acquisition unit, used to obtain the intersection-and-union ratio between the candidate mask information with the largest predicted probability value and at least part of the candidate mask information in the at least one candidate mask information except the candidate mask information with the largest predicted probability value; a second selection module, used to select, from the at least part of the candidate mask information, the candidate mask information whose intersection-and-union ratio with the candidate mask information with the largest predicted probability value is less than a preset intersection-and-union ratio; and a fourth acquisition unit, used to obtain the component mask information based on the selected candidate mask information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Computer three-dimensional model establishing method based on Kinect
CN103325142A
Image detection method and device, and computer readable storage medium
CN112258504A