Image processing method, computing device, storage medium, and computer program product
By processing target images through machine learning models, extracting and fusing spatial structure, image, and point cloud features, the problem of inaccurate modeling results in 3D building modeling is solved, achieving high-precision 3D modeling and a superior user experience.
Patent Information
- Application Number
- PCT/CN2025/087090
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-04-03
- Publication Date
- 2025-12-04
AI Technical Summary
In the process of 3D modeling of houses, existing technologies are prone to inaccurate modeling results, which affects the user's house viewing experience.
Image processing is performed using machine learning models. By extracting the spatial structure features, image features, point cloud features, and relationships between features of the target spatial object, a target spatial object model is generated.
It improves the accuracy and efficiency of 3D modeling, accurately reproduces the details of target spatial objects, and enhances the user's house viewing experience.
Smart Images

Figure CN2025087090_04122025_PF_FP_ABST
Abstract
Description
Image processing method, computing device, storage medium and computer program product
[0001] The present disclosure claims priority to Chinese Patent Application No. 202410684460.4, filed on May 29, 2024, with the Chinese Patent Office, entitled “Image processing method, computing device, storage medium and computer program product”, the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to the field of computer technology, in particular to an image processing method, a computing device, a storage medium and a computer program product. BACKGROUND
[0003] In today's rapidly developing digital era, virtual reality technology is gradually penetrating into various fields of our life, especially in the real estate industry, through virtual reality technology to look at the house has become a common way, through virtual reality and panoramic display, users can realize immersive house viewing and understand the details of the house without going to the site.
[0004] However, in the process of three-dimensional modeling of the house scene, the modeling result is prone to inaccuracy, therefore, an effective technical solution is needed to solve the above problem. SUMMARY
[0005] Therefore, the embodiments of the present disclosure provide four image processing methods. One or more embodiments of the present disclosure also relate to four image processing devices, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects in the prior art.
[0006] According to a first aspect of the embodiments of the present disclosure, a first image processing method is provided, comprising:
[0007] determining a target image containing a target space object and a target object in the target space object;
[0008] inputting the target image into an image processing model to obtain a target space object model containing the target space object and the target object;
[0009] The image processing model is a machine learning model, and the target space object model is determined by the image processing model according to the spatial structure characteristics of the target space object, the image characteristics of the target image, the point cloud characteristics of the target image, the object characteristics of the target object and the correlation between the characteristics.
[0010] According to a second aspect of the embodiments of the present disclosure, a first image processing apparatus is provided, comprising:
[0011] a determining module configured to determine a target image containing a target space object and a target object in the target space object;
[0012] an inputting module configured to input the target image into an image processing model to obtain a target space object model containing the target space object and the target object;
[0013] wherein the image processing model is a machine learning model, and the target space object model is determined by the image processing model according to spatial structure features of the target space object, image features of the target image, point cloud features of the target image, object features of the target object, and a correlation relationship between the features.
[0014] According to a third aspect of the embodiments of the present disclosure, a second image processing method is provided, comprising:
[0015] determining a target image containing a target space object and a target object in the target space object;
[0016] inputting the target image into an image processing model to obtain a target space object model containing the target space object and the target object;
[0017] wherein the image processing model is a machine learning model, and the image processing model comprises a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature analysis unit, the first feature extraction unit is configured to extract image features of the target image and spatial structure features of the target space object, the second feature extraction unit is configured to extract point cloud features of the target image and object features of the target object, the feature fusion unit is configured to fuse the spatial structure features, the image features, the point cloud features, and the object features, and the feature analysis unit is configured to obtain the target space object model containing the target space object and the target object.
[0018] According to a fourth aspect of the embodiments of the present disclosure, a second image processing apparatus is provided, comprising:
[0019] a determining module configured to determine a target image containing a target space object and a target object in the target space object;
[0020] an inputting module configured to input the target image into an image processing model to obtain a target space object model containing the target space object and the target object;
[0021] The image processing model is a machine learning model, comprising a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature parsing unit. The first feature extraction unit extracts image features of the target image and spatial structure features of the target spatial object. The second feature extraction unit extracts point cloud features of the target image and object features of the target object. The feature fusion unit fuses the spatial structure features, image features, point cloud features, and object features. The feature parsing unit obtains a target spatial object model containing the target spatial object and the target object.
[0022] According to a fifth aspect of the present disclosure, a third image processing method is provided, comprising:
[0023] Determine a target image containing the target house and target objects within the target house;
[0024] The target image is input into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0025] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0026] According to a sixth aspect of the present disclosure, a third image processing apparatus is provided, comprising:
[0027] The determination module is configured to determine a target image containing a target house and target objects within the target house;
[0028] The input module is configured to input the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0029] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0030] According to a seventh aspect of the present disclosure, a fourth image processing method is provided, comprising:
[0031] Receive a house search request sent by the client, wherein the house search request includes house information of the target house;
[0032] Based on the house information, a target image containing the target house and the target objects in the target house is determined;
[0033] The target image is input into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0034] The 3D house model is sent to the client and displayed through the client's interface;
[0035] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0036] According to an eighth aspect of the present disclosure, a fourth image processing apparatus is provided, comprising:
[0037] The receiving module is configured to receive a house query request sent by a client, wherein the house query request includes house information of the target house;
[0038] The determination module is configured to determine a target image containing the target house and target objects within the target house based on the house information;
[0039] The input module is configured to input the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0040] The sending module is configured to send the 3D house model to the client and display it through the client's display interface;
[0041] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0042] According to a ninth aspect of the present disclosure, a computing device is provided, comprising:
[0043] Memory and processor;
[0044] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0045] According to a tenth aspect of the present disclosure, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0046] According to an eleventh aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0047] One embodiment of this disclosure determines a target image containing a target spatial object and target objects within that object, and processes the target image using an image processing model to obtain a target spatial object model containing the target spatial object and target objects, thereby achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model is determined based on the spatial structural features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the relationships between these features. When performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target object, and the relationships between the spatial structure and the target objects are considered, thus ensuring the accuracy and efficiency of the 3D modeling of the target spatial object. This allows the final target spatial object model to reproduce the details of the target spatial object, further ensuring a better viewing experience for users when the target spatial object is a house that needs to be viewed. Attached Figure Description
[0048] Figure 1 is a schematic diagram of an application scenario of an image processing method provided in an embodiment of this disclosure;
[0049] Figure 2 is a flowchart of a first image processing method provided in an embodiment of this disclosure;
[0050] Figure 3 is a schematic diagram of the feature fusion unit of the image processing model in an image processing method provided in an embodiment of this disclosure;
[0051] Figure 4 is a schematic diagram of the relationship between sample objects and sample spatial objects in an image processing method provided by an embodiment of this disclosure;
[0052] Figure 5 is a flowchart of an image processing method provided in an embodiment of this disclosure;
[0053] Figure 6 is a schematic diagram of the structure of a first image processing apparatus provided in an embodiment of the present disclosure;
[0054] Figure 7 is a flowchart of a second image processing method provided in an embodiment of this disclosure;
[0055] Figure 8 is a schematic diagram of the structure of a second image processing apparatus provided in an embodiment of the present disclosure;
[0056] Figure 9 is a flowchart of a third image processing method provided in an embodiment of this disclosure;
[0057] Figure 10 is a schematic diagram of the structure of a third image processing apparatus provided in an embodiment of the present disclosure;
[0058] Figure 11 is a flowchart of a fourth image processing method provided in an embodiment of this disclosure;
[0059] Figure 12 is a schematic diagram of the structure of a fourth image processing apparatus provided in an embodiment of the present disclosure;
[0060] Figure 13 is a structural block diagram of a computing device provided in an embodiment of this disclosure. Detailed Implementation
[0061] Numerous specific details are set forth in the following description to provide a full understanding of this disclosure. However, this disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this disclosure. Therefore, this disclosure is not limited to the specific implementations disclosed below.
[0062] The terminology used in one or more embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this disclosure. The singular forms “a,” “the,” and “the” as used in one or more embodiments of this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0063] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first. Depending on the context, the word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when…”.
[0064] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0065] In one or more embodiments of this disclosure, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0066] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0067] First, the terms and concepts involved in one or more embodiments of this disclosure will be explained.
[0068] Virtual Reality (VR): Primarily based on computer technology, it utilizes and integrates the latest advancements in various high-tech fields such as 3D graphics, multimedia, simulation, display, and servo technologies to create a realistic 3D virtual world with multiple sensory experiences, including visual, tactile, and olfactory sensations, thereby giving people in the virtual world a sense of immersion.
[0069] Transformer network: a type of deep learning model.
[0070] Multi-Head Self-Attention (MHSA): This mechanism divides the input sequence into multiple heads, each focusing on a different part, and then merges them to obtain the final representation. This approach allows the model to better capture relationships and dependencies within the input sequence, thereby improving its accuracy and generalization ability. Each "head" independently learns different attention weights, and the outputs of these "heads" are then combined (usually concatenated and passed through a linear layer) to produce the final output representation. In this way, information from different subspaces of the input sequence can be simultaneously considered, thus enhancing the model's expressive power.
[0071] A feedforward network (FFN) is a neural network model composed of multiple neurons. Its main characteristic is that information is transmitted in only one direction, from the input layer through the hidden layers to the output layer, without any recurrent connections. This structure makes feedforward neural networks well-suited for processing static data, such as image classification and text recognition tasks.
[0072] Physical Violation: This can be understood as the physical collision between the walls of a house and objects inside the house.
[0073] Graph Convolutional Networks (GCNs) are deep learning algorithms based on graph-structured data. GCNs can process data with generalized topological graph structures and deeply explore their features and patterns. The core idea of GCNs is to use an adjacency matrix to describe the graph's topology, gradually propagating and aggregating node information through multiple layers of graph convolution operations. In each layer of graph convolution operations, the weights of a node's neighbors and the edges between them are considered, and then the node's features are updated and aggregated. Thus, through multiple layers of graph convolution operations, each node can utilize the information from its neighbors for feature extraction and classification.
[0074] This disclosure provides four image processing methods, and also relates to four image processing apparatuses, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0075] Referring to Figure 1, Figure 1 illustrates an application scenario of an image processing method according to an embodiment of the present disclosure. The image processing method specifically includes the following:
[0076] Determine a target image that includes a target spatial object and target objects within that target spatial object;
[0077] The target image is input into the image processing model to obtain a target spatial object model that includes the target spatial object and the target object.
[0078] The image processing model is a machine learning model, and the target spatial object model is determined by the image processing model based on the spatial structure features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0079] Specifically, by identifying a target image containing the target spatial object and target objects within that object, and processing the target image using an image processing model, a target spatial object model containing the target spatial object and target objects is obtained, thus achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model is determined based on the spatial structural features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the relationships between these features. When performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target object, and the relationships between the target objects are all considered. This ensures the accuracy and efficiency of the 3D modeling of the target spatial object, enabling the final target spatial object model to reproduce the details of the target spatial object. This further enhances the user's viewing experience when the target spatial object is a house that needs to be viewed.
[0080] As shown in Figure 1, Figure 1 includes end-side device 102 and cloud-side device 104.
[0081] In practice, users can send a house query request to cloud device 104 through edge device 102. In response to the house query request, cloud device 104 can determine the house information corresponding to the house query request, thereby determining the target image containing the target house and the target objects contained in the target house, and inputting the target image into the image processing model to obtain the three-dimensional house model of the target image, and sending the three-dimensional house model to edge device 102, which will then display it to the user through the display interface of edge device 102.
[0082] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0083] Cloud-side device 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Cloud-side device 104 can also be a server for a distributed system, or a server integrated with blockchain. Cloud-side device 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0084] It is worth noting that the image processing method provided in this embodiment can be executed by the cloud-side device 104. In other embodiments of this disclosure, the image processing model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the cloud-side device 104, thereby executing the image processing method provided in this embodiment. In other embodiments, the image processing method provided in this embodiment can also be jointly executed by the edge device 102 and the cloud-side device 104.
[0085] Referring to Figure 2, Figure 2 shows a flowchart of a first image processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0086] Step 202: Determine the target image containing the target spatial object and the target objects within the target spatial object.
[0087] Specifically, the image processing method provided in this disclosure can be applied to the field of 3D house modeling. In this field, the image processing method can be used to create a 3D model of a house, obtaining a house model that can then be displayed to users using virtual reality technology. The image processing method provided in this disclosure can also be applied to other application scenarios, such as the gaming field. Specifically, a real-world scene can be 3D modeled using this image processing method and then displayed to users using virtual reality technology, ensuring a realistic gaming experience for users.
[0088] For ease of understanding, this disclosure uses the application of the image processing method to the field of 3D house modeling as an example for illustration, and does not limit the application of the image processing method to other fields.
[0089] In this context, the target space object can be understood as the real-world scene space that needs to be modeled in 3D, such as a house. The target objects within the target space object can be understood as objects placed within the target space object; in the case of a house, the target objects could be furniture within the house.
[0090] In one embodiment of this disclosure, multiple target objects may be placed in the target space object.
[0091] Based on this, a target image containing a real-world scene space and the objects placed in that real-world scene space can be identified.
[0092] For example, in the field of 3D house modeling, house A may contain a table, chairs, bedside tables and a bed. Then, a target image containing house A and the table, chairs, bedside tables and bed placed in house A can be determined.
[0093] In specific implementation, determining the target image containing the target spatial object and the target objects within the target spatial object includes:
[0094] A panoramic view of the target spatial object containing the target object is taken to obtain a target image containing the target spatial object and the target object within the target spatial object; or
[0095] Acquire multiple initial images of the target spatial object, which includes the target object, from different angles;
[0096] The multiple initial images are stitched together to obtain a target image containing the target spatial object and the target objects within the target spatial object.
[0097] In this context, a target image can be understood as an image that can comprehensively display the target spatial object and the target objects within the target spatial object.
[0098] In one embodiment of this disclosure, the target image may be a panoramic image of the target spatial object; specifically, a panoramic image containing the target spatial object and the target objects within the target spatial object may be obtained by taking a panoramic photograph of the target spatial object.
[0099] In practical applications, fisheye lenses, spherical lenses, catadioptric lenses, or omnidirectional lenses can be used in conjunction with sensor units to capture panoramic images of target spatial objects, thereby obtaining panoramic images that include the target spatial object and the target object itself.
[0100] In another embodiment of this disclosure, multiple initial images of the target spatial object containing the target object can be obtained from different angles, and the multiple initial images can be stitched together to obtain a target image containing the target spatial object and the target object.
[0101] In practical applications, a single lens or a single camera can be used to stitch together a series of initial images captured by rotating the camera to obtain a panoramic image. Alternatively, a single camera equipped with a linear array sensor can scan the target object space and directly combine the linear array pixels to obtain a panoramic image. Alternatively, a camera with multiple lenses can be used in conjunction with mirrors to synthesize a panoramic image of the target space object through reflected light paths. This disclosure does not limit these methods.
[0102] Continuing with the previous example, a panoramic photo of house A can be taken to obtain a panoramic image of house A. Alternatively, house A can be photographed from different angles to obtain multiple initial images of house A from different angles, and then these multiple initial images can be stitched together to obtain a panoramic image of house A. Understandably, this panoramic image can fully display house A and the tables, chairs, bedside tables, and bed placed within it.
[0103] In summary, by determining the panoramic image of the target spatial object, it is easier to perform 3D modeling of the target spatial object, restore all objects and details placed in the target spatial object, and further ensure the accuracy of 3D modeling of the target spatial object.
[0104] Step 204: Input the target image into the image processing model to obtain a target spatial object model containing the target spatial object and the target object;
[0105] The image processing model is a machine learning model, and the target spatial object model is determined by the image processing model based on the spatial structure features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0106] A target spatial object model can be understood as a 3D model of the target spatial object and the target objects. In the case of a house, the target spatial object model can be understood as a 3D house model of the target house and the objects placed within it. This 3D house model can display the house layout and the objects placed within it. The spatial structural features of the target spatial object can be understood as the house layout features; the house layout can be understood as the structure and shape of the target house. The point cloud features of the target image can be understood as the point cloud features of the target spatial object and the target objects. The relationships between features can be understood as the relationships between spatial structural features, image features, point cloud features, and object features.
[0107] In practical applications, before inputting the target image into the image processing model to obtain a target spatial object model containing the target spatial object and the target object, the process further includes:
[0108] The image processing model is invoked through the model call interface;
[0109] The step of inputting the target image into an image processing model to obtain a target spatial object model containing the target spatial object and the target object includes:
[0110] The target image is input into the image processing model to obtain a target spatial object model containing the target spatial object and the target object.
[0111] Specifically, the image processing model can be invoked through a model call interface. In another embodiment of this disclosure, the image processing model can also be deployed on a server, and the target image can be processed directly according to the image processing model to obtain a target spatial object model containing the target spatial object and the target object.
[0112] In specific implementation, the step of inputting the target image into an image processing model to obtain a target spatial object model containing the target spatial object and the target object includes:
[0113] The target image is input into the image processing model, where feature extraction is performed on the target image to obtain the spatial structure features of the target spatial object and the image features of the target image.
[0114] Feature extraction is performed on the depth image corresponding to the target image to obtain the point cloud features of the target image and the object features of the target object;
[0115] The spatial structure features, the image features, the point cloud features, and the object features are fused to obtain fused spatial structure features, fused point cloud features, and fused object features;
[0116] Feature analysis is performed on the fused spatial structure features, the fused point cloud features, and the fused object features to obtain a target spatial object model containing the target spatial object and the target object.
[0117] Specifically, after the target image is input into the image processing model, feature extraction can be performed on the target image to obtain the spatial structure features of the target spatial object and the image features of the target image. Feature extraction can also be performed on the depth image corresponding to the target image to obtain the point cloud features of the target image and the object features of the target object. The spatial structure features, image features, point cloud features and object features are then fused to obtain fused spatial structure features, fused point cloud features and fused object features. Finally, feature parsing is performed on the fused spatial structure features, fused point cloud features and fused object features to obtain a target spatial object model containing the target spatial object and the target object.
[0118] Understandably, when the target object space includes multiple target objects, the object features of each target object can be extracted.
[0119] Continuing with the previous example, a panoramic image of house A can be input into the image processing model. The model extracts features from the panoramic image of house A to obtain its spatial structure features and image features. It also extracts features from the corresponding depth image to obtain the point cloud features of the panoramic image, as well as the object features of the table, chair, bedside table, and bed within house A. These features are then fused to obtain fused spatial structure features, fused point cloud features, and fused object features. Finally, feature analysis is performed on these fused spatial structure features, fused point cloud features, and fused object features to obtain a 3D house model containing the house layout of house A, as well as the table, chair, bedside table, and bed within house A.
[0120] In summary, by extracting features from the target image and fusing the extracted features, and simultaneously performing spatial structure and object detection tasks, and by fusing features from multiple modalities, the accuracy of the obtained target spatial object model is improved.
[0121] Further, the step of extracting features from the target image to obtain the spatial structure features of the target spatial object and the image features of the target image includes:
[0122] The target image is encoded to obtain its image features;
[0123] The image features are decoded to obtain the spatial structure features of the target spatial object.
[0124] Specifically, the target image can be encoded to obtain its image features, and the spatial structure features of the target spatial object can be determined based on these image features.
[0125] In summary, by acquiring the image features of the target image and the spatial structural features of the target spatial object, subsequent feature fusion is used to further ensure the accuracy of the 3D modeling of the target spatial object.
[0126] In specific implementation, the step of decoding the image features to obtain the spatial structure features of the target spatial object includes:
[0127] Based on the image features, an initial three-dimensional mesh structure corresponding to the target spatial object is created;
[0128] Calculate the vertex information of the initial three-dimensional mesh structure, and determine the spatial structure features of the target spatial object based on the vertex information.
[0129] The initial 3D mesh structure can be understood as the initial grid structure of the target space object (i.e., the grid layout). In practical applications, the initial 3D mesh structure can be spherical. The vertex information of the initial 3D mesh structure can be understood as the vertex offset information and vertex features of the initial 3D mesh structure.
[0130] Based on this, an initial three-dimensional mesh structure of the target spatial object can be created according to the image features of the target image, the vertex offset information and vertex features of the initial three-dimensional mesh structure can be calculated, and the spatial structure features of the target spatial object can be determined according to the vertex offset information and vertex features.
[0131] In practical applications, the initial 3D mesh structure can be represented by a triangular mesh (V, E, F), where V represents the vertices of the target spatial object's grid, E represents the edges between the grid vertices, and F represents the features associated with the grid vertices. The initial 3D mesh structure can be spherical. By calculating the vertex offsets of the spherical initial 3D mesh structure, the grid result of the target spatial object is obtained, and this grid result is the spatial structural feature of the target spatial object.
[0132] In summary, by creating an initial 3D mesh structure and determining its spatial structural features, this method can be applied to various different apartment layouts, facilitating subsequent 3D modeling of the target spatial object.
[0133] Further, the step of extracting features from the depth image corresponding to the target image to obtain the point cloud features of the target image and the object features of the target object includes:
[0134] Convert the depth image corresponding to the target image into three-dimensional point cloud data;
[0135] The three-dimensional point cloud data is downsampled to obtain sampled three-dimensional point cloud data;
[0136] Based on the sampled 3D point cloud data, the point cloud features of the target image and the contour features of the target object are determined, and the contour features are used as the object features.
[0137] In this context, a point cloud can be understood as a collection of point data. Therefore, 3D point cloud data can be understood as the collection of all point data in a target image, that is, the collection of point data in the target object space and the target object itself. The contour features of the target object can be understood as the bounding box features of the target object, which can be used to describe the shape of the target object.
[0138] Based on this, the depth image corresponding to the target image can be converted into three-dimensional point cloud data, and the three-dimensional point cloud data can be downsampled to obtain sampled three-dimensional point cloud data. Based on the sampled three-dimensional point cloud data, the point cloud features of the target image and the contour features of the target object can be determined, and the contour features can be used as the object features.
[0139] In practical applications, the Fibonacci sampling method can be used to downsample 3D point cloud data.
[0140] In summary, downsampling 3D point cloud data facilitates the subsequent acquisition of point cloud features and object features, thereby enabling feature fusion.
[0141] Specifically, fusing the spatial structure features, the image features, the point cloud features, and the object features to obtain fused spatial structure features, fused point cloud features, and fused object features includes:
[0142] Attention mechanisms are applied to the spatial structure features, image features, point cloud features, and object features to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0143] Specifically, attention mechanisms can be applied to multimodal features such as spatial structure features, image features, point cloud features, and object features to obtain fused spatial structure features, fused point cloud features, and fused object features.
[0144] In summary, by applying an attention mechanism to the multimodal features obtained after feature extraction, we can further explore the intrinsic relationships between different components (such as the walls, floors, and furniture) within the target space object.
[0145] In specific implementation, the attention mechanism processing of the spatial structure features, image features, point cloud features, and object features to obtain the fused spatial structure features, fused point cloud features, and fused object features includes:
[0146] Based on the correlation between the spatial structure features, the image features, the point cloud features, and the object features, the spatial structure features, the image features, the point cloud features, and the object features are fused to obtain the fused spatial structure features, the fused point cloud features, and the fused object features.
[0147] Specifically, taking the acquisition of fused spatial structure features as an example, the spatial structure features, image features, point cloud features, and object features can be weighted and summed based on the relationships between spatial structure features and image features, between spatial structure features and point cloud features, and between spatial structure features and object features to obtain fused spatial structure features. The processes for obtaining fused point cloud features and fused object features are similar and will not be repeated here.
[0148] Understandably, when the target spatial object includes multiple target objects, the spatial structural features, image features, point cloud features, and each object feature can be weighted and summed based on the relationships between object features and spatial structural features, between object features and image features, between object features and point cloud features, and between object features and each other, to obtain the fused object features.
[0149] In summary, by fusing multimodal features based on the relationships between features, we learn the features of multiple modalities and the relationships between features, resulting in fused spatial structure features, fused point cloud features, and fused object features with rich knowledge. This further makes the target spatial object model obtained subsequently based on the fused spatial structure features, fused point cloud features, and fused object features more accurate.
[0150] Furthermore, before fusing the spatial structure features, the image features, the point cloud features, and the object features, the process further includes:
[0151] Determine the location encoding information of the spatial structure features, the image features, the point cloud features, and the object features;
[0152] Based on the location encoding information, the spatial structure features, the image features, the point cloud features, and the object features are stitched together to obtain the stitched spatial structure features, the image features, the point cloud features, and the object features;
[0153] Accordingly, the fusion of the spatial structure features, the image features, the point cloud features, and the object features includes:
[0154] The spliced spatial structure features, image features, point cloud features, and object features are fused together.
[0155] Specifically, the location coding information of spatial structure features, image features, point cloud features, and object features can be determined. Based on this location coding information, the spatial structure features, image features, point cloud features, and object features are stitched together to obtain the stitched spatial structure features, image features, point cloud features, and object features. The stitched spatial structure features, image features, point cloud features, and object features are then fused together.
[0156] In practical applications, spatial structure features, image features, point cloud features, and object features, along with their location codes, can be summed and connected point by point based on the location coding information of spatial structure features, image features, point cloud features, and object features, to obtain the stitched spatial structure features, image features, point cloud features, and object features.
[0157] In summary, by concatenating features based on positional encoding information, it becomes easier to perform subsequent feature parsing on the features of each modality.
[0158] In specific implementation, the step of performing feature analysis on the fused spatial structure features, the fused point cloud features, and the fused object features to obtain a target spatial object model containing the target spatial object and the target object includes:
[0159] The fused spatial structure features are analyzed to obtain the target spatial structure information of the target spatial object;
[0160] The fused point cloud features are analyzed to obtain the object mesh information of the target object;
[0161] The features of the fused object are analyzed to obtain the object contour information of the target object;
[0162] Based on the target space structure information, the object mesh information, and the object contour information, a target space object model containing the target space object and the target object is determined.
[0163] The object mesh information of the target object can include the vertices and triangular faces of the target object. The object contour information of the target object can be understood as the bounding box information of the target object, which can include the center point information, size information, orientation information, and position information of the target object. The target spatial structure information of the target spatial object can be understood as the vertex information of the three-dimensional mesh structure of the target spatial object.
[0164] Based on this, feature analysis can be performed on the fused spatial structure features, fused point cloud features, and fused object features respectively to obtain the target spatial structure information of the target spatial object, the object mesh information and the object contour information of the target object, and then the target spatial structure information, object mesh information and object contour information can be fused to obtain a target spatial object model containing the target spatial object and the target object.
[0165] In summary, feature analysis enables the 3D reconstruction of the target spatial object.
[0166] In addition, before performing feature extraction on the depth image corresponding to the target image, the following steps are also included:
[0167] The target image is transformed to obtain the depth image corresponding to the target image.
[0168] In practical applications, depth estimation can be performed on the target image to obtain the corresponding depth image.
[0169] In summary, by performing depth estimation on the target image, we can provide three-dimensional prior information about the target spatial object for subsequent feature extraction.
[0170] In one embodiment of this disclosure, the image processing model includes a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature parsing unit;
[0171] Accordingly, the step of inputting the target image into the image processing model to obtain a target spatial object model containing the target spatial object and the target object includes:
[0172] The target image is input into the first feature extraction unit to obtain the spatial structure features of the target spatial object and the image features of the target image;
[0173] The depth image corresponding to the target image is input into the second feature extraction unit to obtain the point cloud features of the target image and the object features of the target object;
[0174] The spatial structure features, the image features, the point cloud features, and the object features are input into the feature fusion unit to obtain fused spatial structure features, fused point cloud features, and fused object features;
[0175] The fused spatial structure features, the fused point cloud features, and the fused object features are input into the feature parsing unit to obtain a target spatial object model containing the target spatial object and the target object.
[0176] Specifically, the first feature extraction unit of the image processing model can extract the spatial structure features of the target spatial object and the image features of the target image. The second feature extraction unit of the image processing model can extract the point cloud features of the target image and the object features of the target object. The feature fusion unit of the image processing model can fuse the features between the spatial structure features, image features, point cloud features and object features. The feature parsing unit can perform feature parsing of the fused spatial structure features, fused point cloud features and fused object features.
[0177] Furthermore, the first feature extraction unit includes a first encoder and a first decoder;
[0178] Accordingly, the step of inputting the target image into the first feature extraction unit to obtain the spatial structure features of the target spatial object and the image features of the target image includes:
[0179] The target image is input into the first encoder to obtain the image features of the target image;
[0180] The image features are input into the first decoder to obtain the spatial structure features of the target spatial object.
[0181] Specifically, the process of obtaining the spatial structure features of the target spatial object and the image features of the target image using the first encoder and the first decoder is similar to that described above, and will not be repeated here.
[0182] In specific implementation, the first decoder is a stacked graph convolutional network layer:
[0183] Accordingly, inputting the image features into the first decoder to obtain the spatial structure features of the target spatial object includes:
[0184] Based on the image features, an initial three-dimensional mesh structure corresponding to the target spatial object is created;
[0185] The vertex information of the initial 3D mesh structure is calculated through the stacked graph convolutional network layers, and the spatial structural features of the target spatial object are determined based on the vertex information.
[0186] In practical applications, the first decoder consists of two stacked Graph Convolutional Network (GCN) layers. By using the two stacked GCN layers, the vertex offsets of the initial 3D network structure can be calculated, thereby obtaining the spatial structural features of the target spatial object.
[0187] Specifically, the process of determining the spatial structure features of the target spatial object using the first decoder is similar to that described above, and will not be repeated here.
[0188] In one embodiment of this disclosure, inputting the depth image corresponding to the target image into the second feature extraction unit to obtain the point cloud features of the target image and the object features of the target object includes:
[0189] The depth image corresponding to the target image is input into the second feature extraction unit, where the depth image is converted into three-dimensional point cloud data.
[0190] The three-dimensional point cloud data is downsampled to obtain sampled three-dimensional point cloud data;
[0191] Based on the sampled 3D point cloud data, the point cloud features of the target image and the contour features of the target object are determined, and the contour features are used as the object features.
[0192] The second feature extraction unit is used to detect the object contour information (i.e., bounding box) of the target object and regress the shape code.
[0193] Specifically, the process of obtaining the point cloud features of the target image and the object features of the target object through the second feature extraction unit is similar to that described above, and will not be repeated here.
[0194] In one embodiment of this disclosure, the step of inputting the spatial structure features, the image features, the point cloud features, and the object features into a feature fusion unit to obtain fused spatial structure features, fused point cloud features, and fused object features includes:
[0195] The spatial structure features, image features, point cloud features, and object features are input into the feature fusion unit. In the feature fusion unit, attention mechanism processing is performed on the spatial structure features, image features, point cloud features, and object features to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0196] In practical applications, refer to Figure 3, which illustrates a schematic diagram of a feature fusion unit in an image processing model according to an embodiment of the present disclosure. As shown in Figure 3, the feature fusion unit includes multiple stacked Transformer encoders. Each encoder includes a multi-head self-attention (MHSA) layer and a feed-forward network (FFN) layer. The MHSA layer is the foundation of the encoder, enabling the feature fusion unit to fuse information from different modalities.
[0197] Specifically, the process of feature fusion through the feature fusion unit is similar to that described above, and will not be repeated here.
[0198] Furthermore, the feature fusion unit includes a second encoder, which includes a self-attention mechanism layer;
[0199] Accordingly, in the feature fusion unit, attention mechanism processing is performed on the spatial structure features, the image features, the point cloud features, and the object features to obtain the fused spatial structure features, the fused point cloud features, and the fused object features, including:
[0200] In the feature fusion unit, the self-attention mechanism layer performs fusion processing on the spatial structure features, image features, point cloud features, and object features based on the correlation between them, to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0201] The second encoder can be understood as a Transformer encoder, and the self-attention mechanism layer can be a multi-head self-attention mechanism layer.
[0202] Specifically, the process of feature fusion through the feature fusion unit is similar to that described above, and will not be repeated here.
[0203] Furthermore, before inputting the spatial structure features, the image features, the point cloud features, and the object features into the feature fusion unit, the method further includes:
[0204] Determine the location encoding information of the spatial structure features, the image features, the point cloud features, and the object features;
[0205] Based on the location encoding information, the spatial structure features, the image features, the point cloud features, and the object features are stitched together to obtain the stitched spatial structure features, the image features, the point cloud features, and the object features;
[0206] Accordingly, the feature fusion unit that inputs the spatial structure features, the image features, the point cloud features, and the object features includes:
[0207] The stitched spatial structure features, image features, point cloud features, and object features are input into the feature fusion unit.
[0208] Specifically, the process of stitching together the spatial structure features, the image features, the point cloud features, and the object features is similar to that described above, and will not be repeated here.
[0209] In one embodiment of this description, the feature parsing unit includes a spatial structure feature parsing subunit, a point cloud feature parsing subunit, and an object feature parsing subunit;
[0210] Accordingly, the step of inputting the fused spatial structure features, the fused point cloud features, and the fused object features into the feature parsing unit to obtain a target spatial object model containing the target spatial object and the target object includes:
[0211] The fused spatial structure features are input into the spatial structure feature parsing subunit to obtain the target spatial structure information of the target spatial object;
[0212] The fused point cloud features are input into the point cloud feature parsing subunit to obtain the object mesh information of the target object;
[0213] The fused object features are input into the object feature parsing subunit to obtain the object contour information of the target object;
[0214] Based on the target space structure information, the object mesh information, and the object contour information, a target space object model containing the target space object and the target object is determined.
[0215] Specifically, the process of feature parsing through the feature parsing unit is similar to that described above, and will not be repeated here.
[0216] In practical applications, before inputting the fused point cloud features into the point cloud feature parsing subunit, the process further includes:
[0217] Determine a sample image and label space object model that includes a sample space object and sample objects within the sample space object;
[0218] The sample image and the label space object model are input into the image processing model to obtain the sample object mesh information of the sample object output by the point cloud feature parsing subunit;
[0219] Based on the sample object mesh information and the distance between the sample space objects, and the label space object model, calculate the collision loss function between the sample objects and the sample space objects;
[0220] The point cloud feature parsing subunit is trained according to the collision loss function.
[0221] The collision loss function can be understood as the physical collision loss function between sample objects and objects in the sample space. The point cloud feature parsing subunit can include an encoder and a decoder.
[0222] In practical applications, Figure 4 shows a schematic diagram of the relationship between a sample object and a sample space object in an image processing method according to an embodiment of the present disclosure. Referring to Figure 4, as shown in (a), the sample object is completely inside the sample space object (i.e., the house); as shown in (b), the sample object intersects with the wall of the sample space object; and as shown in (c), the sample object is completely outside the sample space object. When the sample object is completely inside or outside the sample space object, there is no physical constraint. However, when the sample object intersects with the wall of the sample space object, a collision loss function between the sample object and the sample space object can be calculated.
[0223] Specifically, a sample image and a label space object model containing sample space objects and sample objects within those sample space objects can be determined. The sample image and label space object model are input into the image processing model to obtain the sample object mesh information output by the point cloud feature parsing subunit. If the sample object mesh information and the sample space object intersect, the predicted distance between the sample object and the sample space object can be calculated based on the sample object mesh information. The label distance between the sample object and the sample space object can be determined based on the label space object model. Based on the predicted distance and the label distance, a collision loss function between the sample object and the sample space object is calculated. The point cloud feature parsing subunit is then trained based on the collision loss function.
[0224] The physical collision loss function is shown in the following formula.
[0225] in, For the first The vertices of the bounding box of a sample object, X L For the vertices of the apartment grid, the ReLU function ensures that the calculation process only considers the external vertices. This applies when the bounding box of the sample object is completely outside the apartment grid. The value is 0, otherwise it is 1.
[0226] In summary, by training point cloud feature parsing sub-units based on the physical collision loss function, the final 3D modeling results can better conform to the physical world and standardize the relative relationships between objects and house types.
[0227] In one embodiment of this disclosure, the image processing model further includes an image conversion unit;
[0228] Accordingly, before inputting the depth image corresponding to the target image into the second feature extraction unit, the method further includes:
[0229] The target image is input into the image conversion unit to obtain the depth image corresponding to the target image.
[0230] Specifically, the process by which he obtains a depth image through the image conversion unit is similar to that described above, and will not be repeated here.
[0231] In practical applications, before inputting the spatial structure features, image features, point cloud features, and object features into the feature fusion unit, the method further includes:
[0232] Determine a sample image and label space object model that includes a sample space object and sample objects within the sample space object;
[0233] The sample image and the label space object model are input into the image processing model to obtain the sample space structure features of the sample space object, the sample image features of the sample image, the sample point cloud features of the sample image, and the sample object features of the sample object.
[0234] Random masking is performed on the sample spatial structure features, the sample image features, the sample point cloud features, and the sample object features to obtain the masked sample spatial structure features, the sample image features, the sample point cloud features, and the sample object features;
[0235] The feature fusion unit is trained based on the masked sample spatial structure features, sample image features, sample point cloud features, sample object features, and the label spatial object model.
[0236] Specifically, a sample image and a label space object model containing sample space objects and sample objects within those sample space objects can be determined. The sample image and label space object model are then input into an image processing model to obtain the sample space structure features of the sample space objects, the sample image features of the sample images, the sample point cloud features of the sample images, and the sample object features of the sample objects. Random masking is performed on the sample space structure features, sample image features, sample point cloud features, and sample object features to obtain the masked sample space structure features, sample image features, sample point cloud features, and sample object features. Based on the masked sample space structure features, sample image features, sample point cloud features, and sample object features, a predicted spatial object model output by the image processing model is obtained. The feature fusion unit is then trained based on this predicted spatial object model and the label space object model.
[0237] In practical applications, the formula for randomly masking sample spatial structure features, sample image features, sample point cloud features, and sample object features is as follows:
[0238] Where d is the dimension of the input feature and M is the mask matrix.
[0239] The image processing method provided in this disclosure will be further described below with reference to Figure 5, taking the application of the image processing method in house layout estimation as an example. Figure 5 shows a flowchart of the processing procedure of an image processing method provided in an embodiment of this disclosure, specifically including the following steps.
[0240] Step 502: Input the target image into the first feature extraction unit to obtain the spatial structure features (i.e., house type features) of the target spatial object (i.e., the house) and the image features of the target image.
[0241] Specifically, the target image is a panoramic image of the target house and the target objects within the target house. The image processing model includes a first feature extraction unit (i.e., house type estimation module), an image transformation unit (i.e., depth estimation module), a second feature extraction unit (i.e., object detection module), a feature fusion unit (i.e., context fusion module), a spatial structure feature parsing subunit (i.e., house type branch), a point cloud feature parsing subunit (i.e., object branch), and an object feature parsing subunit (i.e., bounding box branch).
[0242] The first feature extraction unit includes a first encoder and a first decoder. The first decoder is a stacked graph convolutional network layer. The target image is input into the first encoder to obtain the image features of the target image. The image features are input into the first decoder. The first decoder creates an initial three-dimensional mesh structure (i.e., sphere) corresponding to the target spatial object based on the image features. The vertex offsets of the initial three-dimensional mesh structure are calculated through the stacked graph convolutional network layer. Based on the vertex offsets, the spatial structure features of the target spatial object are determined.
[0243] Step 504: Input the target image into the image conversion unit to obtain the depth image corresponding to the target image.
[0244] Step 506: Input the depth image corresponding to the target image into the second feature extraction unit to obtain the point cloud features of the target image and the object features of the target object.
[0245] Specifically, in the second feature extraction unit, the depth image corresponding to the target image can be converted into three-dimensional point cloud data, and the three-dimensional point cloud data can be downsampled to obtain sampled three-dimensional point cloud data. Based on the sampled three-dimensional point cloud data, the point cloud features of the target image and the contour features of the target object are determined, and the contour features are used as object features.
[0246] Step 508: Input the spatial structure features, image features, point cloud features and object features into the feature fusion unit to obtain fused spatial structure features, fused point cloud features and fused object features.
[0247] Specifically, taking the acquisition of fused spatial structure features as an example, the spatial structure features, image features, point cloud features, and object features can be weighted and summed based on the correlation between spatial structure features and image features, the correlation between spatial structure features and point cloud features, and the correlation between spatial structure features and object features to obtain fused spatial structure features.
[0248] Step 510: Input the fused spatial structure features into the spatial structure feature parsing sub-unit to obtain the target spatial structure information (i.e., apartment layout) of the target spatial object.
[0249] Step 512: Input the fused point cloud features into the point cloud feature parsing sub-unit to obtain the object mesh information (i.e., object mesh) of the target object.
[0250] Step 514: Input the fused object features into the object feature parsing subunit to obtain the object contour information (i.e., bounding box) of the target object.
[0251] Step 516: Based on the target space structure information, object mesh information, and object contour information, determine the target space object model (i.e., the final reconstruction result) that includes the target space object and the target object.
[0252] Specifically, target space structure information, object mesh information, and object contour information can be fused to obtain a target space object model that includes the target space object and the target object.
[0253] One embodiment of this disclosure determines a target image containing a target spatial object and target objects within that object, and processes the target image using an image processing model to obtain a target spatial object model containing the target spatial object and target objects, thereby achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model is determined based on the spatial structural features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the relationships between these features. When performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target object, and the relationships between the spatial structure and the target objects are considered, thus ensuring the accuracy and efficiency of the 3D modeling of the target spatial object. This allows the final target spatial object model to reproduce the details of the target spatial object, further ensuring a better viewing experience for users when the target spatial object is a house that needs to be viewed.
[0254] Corresponding to the above method embodiments, this disclosure also provides an image processing apparatus embodiment. FIG6 shows a schematic diagram of the structure of a first image processing apparatus provided in an embodiment of this disclosure. As shown in FIG6, the apparatus includes:
[0255] The determining module 602 is configured to determine a target image containing a target space object and target objects in the target space object;
[0256] Input module 604 is configured to input the target image into an image processing model to obtain a target spatial object model containing the target spatial object and the target object;
[0257] The image processing model is a machine learning model, and the target spatial object model is determined by the image processing model based on the spatial structure features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0258] In an optional embodiment, the input module 604 is further configured to:
[0259] The target image is input into the image processing model, where feature extraction is performed on the target image to obtain the spatial structure features of the target spatial object and the image features of the target image.
[0260] Feature extraction is performed on the depth image corresponding to the target image to obtain the point cloud features of the target image and the object features of the target object;
[0261] The spatial structure features, the image features, the point cloud features, and the object features are fused to obtain fused spatial structure features, fused point cloud features, and fused object features;
[0262] Feature analysis is performed on the fused spatial structure features, the fused point cloud features, and the fused object features to obtain a target spatial object model containing the target spatial object and the target object.
[0263] In an optional embodiment, the input module 604 is further configured to:
[0264] The target image is encoded to obtain its image features;
[0265] The image features are decoded to obtain the spatial structure features of the target spatial object.
[0266] In an optional embodiment, the input module 604 is further configured to:
[0267] Based on the image features, an initial three-dimensional mesh structure corresponding to the target spatial object is created;
[0268] Calculate the vertex information of the initial three-dimensional mesh structure, and determine the spatial structure features of the target spatial object based on the vertex information.
[0269] In an optional embodiment, the input module 604 is further configured to:
[0270] Convert the depth image corresponding to the target image into three-dimensional point cloud data;
[0271] The three-dimensional point cloud data is downsampled to obtain sampled three-dimensional point cloud data;
[0272] Based on the sampled 3D point cloud data, the point cloud features of the target image and the contour features of the target object are determined, and the contour features are used as the object features.
[0273] In an optional embodiment, the input module 604 is further configured to:
[0274] Attention mechanisms are applied to the spatial structure features, image features, point cloud features, and object features to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0275] In an optional embodiment, the input module 604 is further configured to:
[0276] Based on the correlation between the spatial structure features, the image features, the point cloud features, and the object features, the spatial structure features, the image features, the point cloud features, and the object features are fused to obtain the fused spatial structure features, the fused point cloud features, and the fused object features.
[0277] In an optional embodiment, the input module 604 is further configured to:
[0278] Determine the location encoding information of the spatial structure features, the image features, the point cloud features, and the object features;
[0279] Based on the location encoding information, the spatial structure features, the image features, the point cloud features, and the object features are stitched together to obtain the stitched spatial structure features, the image features, the point cloud features, and the object features;
[0280] Accordingly, the fusion of the spatial structure features, the image features, the point cloud features, and the object features includes:
[0281] The spliced spatial structure features, image features, point cloud features, and object features are fused together.
[0282] In an optional embodiment, the input module 604 is further configured to:
[0283] The fused spatial structure features are analyzed to obtain the target spatial structure information of the target spatial object;
[0284] The fused point cloud features are analyzed to obtain the object mesh information of the target object;
[0285] The features of the fused object are analyzed to obtain the object contour information of the target object;
[0286] Based on the target space structure information, the object mesh information, and the object contour information, a target space object model containing the target space object and the target object is determined.
[0287] In an optional embodiment, the input module 604 is further configured to:
[0288] The target image is transformed to obtain the depth image corresponding to the target image.
[0289] In an optional embodiment, the determining module 602 is further configured to:
[0290] A panoramic view of the target spatial object containing the target object is taken to obtain a target image containing the target spatial object and the target object within the target spatial object; or
[0291] Acquire multiple initial images of the target spatial object, which includes the target object, from different angles;
[0292] The multiple initial images are stitched together to obtain a target image containing the target spatial object and the target objects within the target spatial object.
[0293] In one optional embodiment, the image processing model includes a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature parsing unit;
[0294] The input module 604 is further configured as follows:
[0295] The target image is input into the first feature extraction unit to obtain the spatial structure features of the target spatial object and the image features of the target image;
[0296] The depth image corresponding to the target image is input into the second feature extraction unit to obtain the point cloud features of the target image and the object features of the target object;
[0297] The spatial structure features, the image features, the point cloud features, and the object features are input into the feature fusion unit to obtain fused spatial structure features, fused point cloud features, and fused object features;
[0298] The fused spatial structure features, the fused point cloud features, and the fused object features are input into the feature parsing unit to obtain a target spatial object model containing the target spatial object and the target object.
[0299] In an optional embodiment, the first feature extraction unit includes a first encoder and a first decoder;
[0300] Accordingly, the input module 604 is further configured as follows:
[0301] The target image is input into the first encoder to obtain the image features of the target image;
[0302] The image features are input into the first decoder to obtain the spatial structure features of the target spatial object.
[0303] In an optional embodiment, the first decoder is a stacked graph convolutional network layer:
[0304] Accordingly, the input module 604 is further configured as follows:
[0305] Based on the image features, an initial three-dimensional mesh structure corresponding to the target spatial object is created;
[0306] The vertex information of the initial 3D mesh structure is calculated through the stacked graph convolutional network layers, and the spatial structural features of the target spatial object are determined based on the vertex information.
[0307] In an optional embodiment, the input module 604 is further configured to:
[0308] The depth image corresponding to the target image is input into the second feature extraction unit, where the depth image is converted into three-dimensional point cloud data.
[0309] The three-dimensional point cloud data is downsampled to obtain sampled three-dimensional point cloud data;
[0310] Based on the sampled 3D point cloud data, the point cloud features of the target image and the contour features of the target object are determined, and the contour features are used as the object features.
[0311] In an optional embodiment, the input module 604 is further configured to:
[0312] The spatial structure features, image features, point cloud features, and object features are input into the feature fusion unit. In the feature fusion unit, attention mechanism processing is performed on the spatial structure features, image features, point cloud features, and object features to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0313] In an optional embodiment, the feature fusion unit includes a second encoder, which includes a self-attention mechanism layer;
[0314] Accordingly, the input module 604 is further configured as follows:
[0315] In the feature fusion unit, the self-attention mechanism layer performs fusion processing on the spatial structure features, image features, point cloud features, and object features based on the correlation between them, to obtain the fused spatial structure features, fused point cloud features, and fused object features.
[0316] In an optional embodiment, the input module 604 is further configured to:
[0317] Determine the location encoding information of the spatial structure features, the image features, the point cloud features, and the object features;
[0318] Based on the location encoding information, the spatial structure features, the image features, the point cloud features, and the object features are stitched together to obtain the stitched spatial structure features, the image features, the point cloud features, and the object features;
[0319] Accordingly, the feature fusion unit that inputs the spatial structure features, the image features, the point cloud features, and the object features includes:
[0320] The stitched spatial structure features, image features, point cloud features, and object features are input into the feature fusion unit.
[0321] In an optional embodiment, the feature parsing unit includes a spatial structure feature parsing subunit, a point cloud feature parsing subunit, and an object feature parsing subunit;
[0322] Accordingly, the input module 604 is further configured as follows:
[0323] The fused spatial structure features are input into the spatial structure feature parsing subunit to obtain the target spatial structure information of the target spatial object;
[0324] The fused point cloud features are input into the point cloud feature parsing subunit to obtain the object mesh information of the target object;
[0325] The fused object features are input into the object feature parsing subunit to obtain the object contour information of the target object;
[0326] Based on the target space structure information, the object mesh information, and the object contour information, a target space object model containing the target space object and the target object is determined.
[0327] In an optional embodiment, the device further includes a training module configured to:
[0328] Determine a sample image and label space object model that includes a sample space object and sample objects within the sample space object;
[0329] The sample image and the label space object model are input into the image processing model to obtain the sample object mesh information of the sample object output by the point cloud feature parsing subunit;
[0330] Based on the sample object mesh information and the distance between the sample space objects, and the label space object model, calculate the collision loss function between the sample objects and the sample space objects;
[0331] The point cloud feature parsing subunit is trained according to the collision loss function.
[0332] In an optional embodiment, the image processing model further includes an image conversion unit;
[0333] Accordingly, the input module 604 is further configured as follows:
[0334] The target image is input into the image conversion unit to obtain the depth image corresponding to the target image.
[0335] In an optional embodiment, the training module is further configured to:
[0336] Determine a sample image and label space object model that includes a sample space object and sample objects within the sample space object;
[0337] The sample image and the label space object model are input into the image processing model to obtain the sample space structure features of the sample space object, the sample image features of the sample image, the sample point cloud features of the sample image, and the sample object features of the sample object.
[0338] Random masking is performed on the sample spatial structure features, the sample image features, the sample point cloud features, and the sample object features to obtain the masked sample spatial structure features, the sample image features, the sample point cloud features, and the sample object features;
[0339] The feature fusion unit is trained based on the masked sample spatial structure features, sample image features, sample point cloud features, sample object features, and the label spatial object model.
[0340] One embodiment of this disclosure determines a target image containing a target spatial object and target objects within that object, and processes the target image using an image processing model to obtain a target spatial object model containing the target spatial object and target objects, thereby achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model is determined based on the spatial structural features of the target spatial object, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the relationships between these features. When performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target object, and the relationships between the spatial structure and the target objects are considered, thus ensuring the accuracy and efficiency of the 3D modeling of the target spatial object. This allows the final target spatial object model to reproduce the details of the target spatial object, further ensuring a better viewing experience for users when the target spatial object is a house that needs to be viewed.
[0341] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0342] Referring to Figure 7, Figure 7 shows a flowchart of a second image processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0343] Step 702: Determine the target image containing the target spatial object and the target objects within the target spatial object;
[0344] Step 704: Input the target image into the image processing model to obtain a target spatial object model containing the target spatial object and the target object;
[0345] The image processing model is a machine learning model, comprising a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature parsing unit. The first feature extraction unit extracts image features of the target image and spatial structure features of the target spatial object. The second feature extraction unit extracts point cloud features of the target image and object features of the target object. The feature fusion unit fuses the spatial structure features, image features, point cloud features, and object features. The feature parsing unit obtains a target spatial object model containing the target spatial object and the target object.
[0346] Specifically, one embodiment of this disclosure determines a target image containing a target spatial object and target objects within that target spatial object, and processes the target image using an image processing model to obtain a target spatial object model containing the target spatial object and target objects, thereby achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model extracts image features from the target image and spatial structure features of the target spatial object through a first feature extraction unit, extracts point cloud features from the target image and object features of the target object through a second feature extraction unit, fuses spatial structure features, image features, point cloud features, and object features through a feature fusion unit, and obtains a target spatial object model containing the target spatial object and target objects through a feature parsing unit. In performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target objects, and the relationships between the spatial structure of the target spatial object and the target objects are considered. This multi-dimensional feature approach ensures the accuracy and efficiency of 3D modeling of the target spatial object, enabling the final target spatial object model to reproduce the details of the target spatial object, further ensuring a better viewing experience for users when the target spatial object is a house to be viewed.
[0347] The above is an illustrative scheme of an image processing method according to this embodiment. It should be noted that the technical solution of this image processing method belongs to the same concept as the technical solution of the image processing method described above. For details not described in detail in the technical solution of the image processing method, please refer to the description of the technical solution of the image processing method described above.
[0348] Corresponding to the above method embodiments, this disclosure also provides an image processing apparatus embodiment. FIG8 shows a schematic diagram of the structure of a second image processing apparatus provided in one embodiment of this disclosure. As shown in FIG8, the apparatus includes:
[0349] The determination module 802 is configured to determine a target image containing a target spatial object and target objects in the target spatial object;
[0350] The input module 804 is configured to input the target image into an image processing model to obtain a target spatial object model that includes the target spatial object and the target object.
[0351] The image processing model is a machine learning model, comprising a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature parsing unit. The first feature extraction unit extracts image features of the target image and spatial structure features of the target spatial object. The second feature extraction unit extracts point cloud features of the target image and object features of the target object. The feature fusion unit fuses the spatial structure features, image features, point cloud features, and object features. The feature parsing unit obtains a target spatial object model containing the target spatial object and the target object.
[0352] Specifically, one embodiment of this disclosure determines a target image containing a target spatial object and target objects within that target spatial object, and processes the target image using an image processing model to obtain a target spatial object model containing the target spatial object and target objects, thereby achieving 3D modeling of the target spatial object. Furthermore, this target spatial object model extracts image features from the target image and spatial structure features of the target spatial object through a first feature extraction unit, extracts point cloud features from the target image and object features of the target object through a second feature extraction unit, fuses spatial structure features, image features, point cloud features, and object features through a feature fusion unit, and obtains a target spatial object model containing the target spatial object and target objects through a feature parsing unit. In performing 3D modeling of the target spatial object, the spatial structure of the target spatial object, the object features of the target objects, and the relationships between the spatial structure of the target spatial object and the target objects are considered. This multi-dimensional feature approach ensures the accuracy and efficiency of 3D modeling of the target spatial object, enabling the final target spatial object model to reproduce the details of the target spatial object, further ensuring a better viewing experience for users when the target spatial object is a house to be viewed.
[0353] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0354] Referring to Figure 9, Figure 9 shows a flowchart of a third image processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0355] Step 902: Determine a target image containing the target house and the target objects within the target house;
[0356] Step 904: Input the target image into the image processing model to obtain a three-dimensional house model containing the target house and the target house;
[0357] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0358] Specifically, the target house can be understood as the target spatial object mentioned above, the three-dimensional house model can be understood as the model of the target spatial object mentioned above, and the house type features can be understood as spatial structural features.
[0359] One embodiment of this disclosure determines a target image containing a target house and target objects within the target house, and processes the target image using an image processing model to obtain a 3D house model containing the target house and target objects, thus achieving 3D modeling of the target house. Furthermore, this 3D house model is determined based on the house layout features, image features of the target image, point cloud features of the target image, object features of the target objects, and the relationships between these features. When performing 3D modeling of the target house, the layout of the target house, the object features of the target objects, and the relationships between the layout of the target house and the target objects are considered, thereby ensuring the accuracy and efficiency of the 3D modeling of the target house. This allows the final 3D house model to reproduce the details of the target house, further enhancing the user's house viewing experience.
[0360] The above is an illustrative scheme of an image processing method according to this embodiment. It should be noted that the technical solution of this image processing method belongs to the same concept as the technical solution of the image processing method described above. For details not described in detail in the technical solution of the image processing method, please refer to the description of the technical solution of the image processing method described above.
[0361] Corresponding to the above method embodiments, this disclosure also provides an image processing apparatus embodiment. FIG10 shows a schematic diagram of the structure of a third image processing apparatus provided in an embodiment of this disclosure. As shown in FIG10, the apparatus includes:
[0362] The determining module 1002 is configured to determine a target image containing a target house and target objects in the target house;
[0363] The input module 1004 is configured to input the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0364] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0365] One embodiment of this disclosure determines a target image containing a target house and target objects within the target house, and processes the target image using an image processing model to obtain a 3D house model containing the target house and target objects, thus achieving 3D modeling of the target house. Furthermore, this 3D house model is determined based on the house layout features, image features of the target image, point cloud features of the target image, object features of the target objects, and the relationships between these features. When performing 3D modeling of the target house, the layout of the target house, the object features of the target objects, and the relationships between the layout of the target house and the target objects are considered, thereby ensuring the accuracy and efficiency of the 3D modeling of the target house. This allows the final 3D house model to reproduce the details of the target house, further enhancing the user's house viewing experience.
[0366] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0367] Referring to Figure 11, Figure 11 shows a flowchart of a fourth image processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
[0368] Step 1102: Receive a house query request sent by the client, wherein the house query request includes house information of the target house;
[0369] Step 1104: Based on the house information, determine a target image containing the target house and the target objects in the target house;
[0370] Step 1106: Input the target image into the image processing model to obtain a three-dimensional house model containing the target house and the target house;
[0371] Step 1108: Send the 3D house model to the client and display it through the client's display interface;
[0372] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0373] Specifically, in the field of user house viewing, users can send a house query request to the server through the client. The server can determine the target image containing the target house and the target objects in the target house based on the house information carried in the house query request, and input the target image into the image processing model. After the image processing model processes the image, a 3D house model containing the target house and the target objects is obtained. The 3D house model is sent to the client and displayed to the user through the client's display interface.
[0374] One embodiment of this disclosure receives a house query request sent by a client, thereby determining a target image containing the target house and target objects within the target house. The target image is then processed using an image processing model to obtain a 3D house model containing the target house and target objects, thus achieving 3D modeling of the target house. Furthermore, this 3D house model is determined based on the house's floor plan features, the target image's image features, the target image's point cloud features, the target object's object features, and the relationships between these features. When performing 3D modeling of the target house, the floor plan of the target house, the object features of the target objects, and the relationships between the floor plan and the target objects are considered, thereby ensuring the accuracy and efficiency of the 3D modeling of the target house. This allows the final 3D house model to accurately reproduce the details of the target house, further enhancing the user's house viewing experience.
[0375] The above is an illustrative scheme of an image processing method according to this embodiment. It should be noted that the technical solution of this image processing method belongs to the same concept as the technical solution of the image processing method described above. For details not described in detail in the technical solution of the image processing method, please refer to the description of the technical solution of the image processing method described above.
[0376] Corresponding to the above method embodiments, this disclosure also provides an image processing apparatus embodiment. FIG12 shows a schematic diagram of the structure of a fourth image processing apparatus provided in one embodiment of this disclosure. As shown in FIG12, the apparatus includes:
[0377] The receiving module 1202 is configured to receive a house query request sent by a client, wherein the house query request includes house information of the target house;
[0378] The determining module 1204 is configured to determine a target image containing the target house and the target object in the target house based on the house information;
[0379] The input module 1206 is configured to input the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house.
[0380] The sending module 1208 is configured to send the three-dimensional house model to the client and display it through the client's display interface;
[0381] The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model based on the house type features of the target house, the image features of the target image, the point cloud features of the target image, the object features of the target object, and the correlation between the features.
[0382] One embodiment of this disclosure receives a house query request sent by a client, thereby determining a target image containing the target house and target objects within the target house. The target image is then processed using an image processing model to obtain a 3D house model containing the target house and target objects, thus achieving 3D modeling of the target house. Furthermore, this 3D house model is determined based on the house's floor plan features, the target image's image features, the target image's point cloud features, the target object's object features, and the relationships between these features. When performing 3D modeling of the target house, the floor plan of the target house, the object features of the target objects, and the relationships between the floor plan and the target objects are considered, thereby ensuring the accuracy and efficiency of the 3D modeling of the target house. This allows the final 3D house model to accurately reproduce the details of the target house, further enhancing the user's house viewing experience.
[0383] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0384] Figure 13 shows a structural block diagram of a computing device 1300 according to an embodiment of the present disclosure. The components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0385] The computing device 1300 also includes an access device 1340, which enables the computing device 1300 to communicate via one or more networks 1360. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0386] In one embodiment of this disclosure, the aforementioned components of the computing device 1300, as well as other components not shown in FIG. 13, may also be connected to each other, for example, via a bus. It should be understood that the computing device block diagram shown in FIG. 13 is merely for illustrative purposes and is not intended to limit the scope of this disclosure. Those skilled in the art can add or replace other components as needed.
[0387] The computing device 1300 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1300 can also be a mobile or stationary server.
[0388] The processor 1320 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above method.
[0389] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0390] An embodiment of this disclosure also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0391] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are relatively simple in description because they are substantially similar to the method embodiments; relevant parts can be referred to the descriptions of the method embodiments.
[0392] An embodiment of this disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0393] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.
[0394] The foregoing has described specific embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0395] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0396] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this disclosure are not limited to the described order of actions, because according to the embodiments of this disclosure, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this disclosure.
[0397] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0398] The preferred embodiments disclosed above are merely illustrative of this disclosure. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments of this disclosure. These embodiments are selected and specifically described in this disclosure to better explain the principles and practical applications of the embodiments of this disclosure, thereby enabling those skilled in the art to better understand and utilize this disclosure. This disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. An image processing method, comprising: determining a target image containing a target space object and a target object in the target space object; inputting the target image into an image processing model to obtain a target space object model containing the target space object and the target object; wherein the image processing model is a machine learning model, and the target space object model is determined by the image processing model according to spatial structure features of the target space object, image features of the target image, point cloud features of the target image, object features of the target object, and a correlation between the features.
2. The image processing method of claim 1, wherein the inputting the target image into the image processing model to obtain the target space object model containing the target space object and the target object comprises: inputting the target image into the image processing model, and performing feature extraction on the target image in the image processing model to obtain spatial structure features of the target space object and image features of the target image; performing feature extraction on a depth image corresponding to the target image to obtain point cloud features of the target image and object features of the target object; fusing the spatial structure features, the image features, the point cloud features, and the object features to obtain fused spatial structure features, fused point cloud features, and fused object features; performing feature analysis on the fused spatial structure features, the fused point cloud features, and the fused object features to obtain the target space object model containing the target space object and the target object.
3. The image processing method of claim 2, wherein the performing feature extraction on the target image to obtain the spatial structure features of the target space object and the image features of the target image comprises: performing encoding processing on the target image to obtain the image features of the target image; performing decoding processing on the image features to obtain the spatial structure features of the target space object.
4. The image processing method of claim 3, wherein the performing decoding processing on the image features to obtain the spatial structure features of the target space object comprises: creating an initial three-dimensional grid structure corresponding to the target space object according to the image features; calculating vertex information of the initial three-dimensional grid structure, and determining the spatial structure features of the target space object according to the vertex information.
5. The image processing method of any one of claims 2-4, wherein the performing feature extraction on a depth image corresponding to the target image to obtain the point cloud features of the target image and the object features of the target object comprises: converting the depth image corresponding to the target image into three-dimensional point cloud data; performing down-sampling processing on the three-dimensional point cloud data to obtain sampled three-dimensional point cloud data; determining the point cloud features of the target image and contour features of the target object according to the sampled three-dimensional point cloud data, and taking the contour features as the object features.
6. The image processing method of any one of claims 2-5, wherein the fusing of the spatial structure feature, the image feature, the point cloud feature, and the object feature to obtain fused spatial structure feature, fused point cloud feature, and fused object feature comprises: performing attention mechanism processing on the spatial structure feature, the image feature, the point cloud feature, and the object feature to obtain the fused spatial structure feature, the fused point cloud feature, and the fused object feature.
7. The image processing method of claim 6, wherein the performing of the attention mechanism processing on the spatial structure feature, the image feature, the point cloud feature, and the object feature to obtain the fused spatial structure feature, the fused point cloud feature, and the fused object feature comprises: fusing the spatial structure feature, the image feature, the point cloud feature, and the object feature according to an association relationship among the spatial structure feature, the image feature, the point cloud feature, and the object feature to obtain the fused spatial structure feature, the fused point cloud feature, and the fused object feature.
8. The image processing method of any one of claims 2-7, further comprising, before the fusing of the spatial structure feature, the image feature, the point cloud feature, and the object feature: determining position encoding information of the spatial structure feature, the image feature, the point cloud feature, and the object feature; splicing the spatial structure feature, the image feature, the point cloud feature, and the object feature according to the position encoding information to obtain spliced spatial structure feature, image feature, point cloud feature, and object feature; and correspondingly, the fusing of the spatial structure feature, the image feature, the point cloud feature, and the object feature comprises: fusing the spliced spatial structure feature, image feature, point cloud feature, and object feature.
9. The image processing method of any one of claims 2-8, wherein the feature analysis of the fused spatial structure feature, the fused point cloud feature, and the fused object feature to obtain a target spatial object model containing the target spatial object and the target object comprises: performing feature analysis on the fused spatial structure feature to obtain target spatial structure information of the target spatial object; performing feature analysis on the fused point cloud feature to obtain object mesh information of the target object; performing feature analysis on the fused object feature to obtain object contour information of the target object; and determining a target spatial object model containing the target spatial object and the target object according to the target spatial structure information, the object mesh information, and the object contour information.
10. The image processing method of any one of claims 2-9, further comprising, before the feature extraction of the depth image corresponding to the target image: performing image conversion on the target image to obtain a depth image corresponding to the target image.
11. The image processing method of any one of claims 1-10, wherein the determining the target image containing a target spatial object and a target object in the target spatial object comprises: taking a panoramic photo of the target spatial object containing the target object to obtain the target image containing the target spatial object and the target object in the target spatial object; or obtaining a plurality of initial images of the target spatial object containing the target object at different angles; stitching the plurality of initial images to obtain the target image containing the target spatial object and the target object in the target spatial object.
12. The image processing method of any one of claims 1-11, wherein the inputting the target image into the image processing model to obtain a target spatial object model containing the target spatial object and the target object further comprises: calling the image processing model through a model calling interface; and the inputting the target image into the image processing model to obtain the target spatial object model containing the target spatial object and the target object comprises: inputting the target image into the called image processing model to obtain the target spatial object model containing the target spatial object and the target object.
13. The image processing method of claim 1, wherein the image processing model comprises a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature analysis unit; and correspondingly, the inputting the target image into the image processing model to obtain the target spatial object model containing the target spatial object and the target object comprises: inputting the target image into the first feature extraction unit to obtain a spatial structure feature of the target spatial object and an image feature of the target image; inputting a depth image corresponding to the target image into the second feature extraction unit to obtain a point cloud feature of the target image and an object feature of the target object; inputting the spatial structure feature, the image feature, the point cloud feature, and the object feature into the feature fusion unit to obtain a fused spatial structure feature, a fused point cloud feature, and a fused object feature; inputting the fused spatial structure feature, the fused point cloud feature, and the fused object feature into the feature analysis unit to obtain the target spatial object model containing the target spatial object and the target object.
14. The image processing method of claim 13, wherein the first feature extraction unit comprises a first encoder and a first decoder; and correspondingly, the inputting the target image into the first feature extraction unit to obtain the spatial structure feature of the target spatial object and the image feature of the target image comprises: inputting the target image into the first encoder to obtain the image feature of the target image; inputting the image feature into the first decoder to obtain the spatial structure feature of the target spatial object.
15. The image processing method of claim 14, wherein the first decoder is a stacked graph convolutional network layer. Correspondingly, the inputting the image feature into the first decoder to obtain the spatial structure feature of the target spatial object comprises: creating an initial three-dimensional mesh structure corresponding to the target spatial object according to the image feature; calculating vertex information of the initial three-dimensional mesh structure through the stacked graph convolution network layer, and determining the spatial structure feature of the target spatial object according to the vertex information.
16. The image processing method according to any one of claims 13-15, wherein the inputting the depth image corresponding to the target image into the second feature extraction unit to obtain the point cloud feature of the target image and the object feature of the target object comprises: inputting the depth image corresponding to the target image into the second feature extraction unit, and converting the depth image into three-dimensional point cloud data in the second feature extraction unit; performing down-sampling processing on the three-dimensional point cloud data to obtain sampled three-dimensional point cloud data; determining the point cloud feature of the target image and the contour feature of the target object according to the sampled three-dimensional point cloud data, and taking the contour feature as the object feature.
17. The image processing method according to any one of claims 13-16, wherein the inputting the spatial structure feature, the image feature, the point cloud feature and the object feature into the feature fusion unit to obtain the fused spatial structure feature, the fused point cloud feature and the fused object feature comprises: inputting the spatial structure feature, the image feature, the point cloud feature and the object feature into the feature fusion unit, and performing attention mechanism processing on the spatial structure feature, the image feature, the point cloud feature and the object feature in the feature fusion unit to obtain the fused spatial structure feature, the fused point cloud feature and the fused object feature.
18. The image processing method according to claim 17, wherein the feature fusion unit comprises a second encoder, and the second encoder comprises a self-attention mechanism layer. Correspondingly, the performing attention mechanism processing on the spatial structure feature, the image feature, the point cloud feature and the object feature in the feature fusion unit to obtain the fused spatial structure feature, the fused point cloud feature and the fused object feature comprises: in the feature fusion unit, performing fusion processing on the spatial structure feature, the image feature, the point cloud feature and the object feature according to the correlation relationship among the spatial structure feature, the image feature, the point cloud feature and the object feature through the self-attention mechanism layer to obtain the fused spatial structure feature, the fused point cloud feature and the fused object feature.
19. The image processing method according to any one of claims 13-18, further comprising, before the inputting the spatial structure feature, the image feature, the point cloud feature and the object feature into the feature fusion unit: determining position encoding information of the spatial structure feature, the image feature, the point cloud feature and the object feature. According to the position coding information, the spatial structure feature, the image feature, the point cloud feature and the object feature are spliced to obtain the spliced spatial structure feature, the spliced image feature, the spliced point cloud feature and the spliced object feature; Correspondingly, the inputting the spatial structure feature, the image feature, the point cloud feature and the object feature into the feature fusion unit comprises: The spliced spatial structure feature, the spliced image feature, the spliced point cloud feature and the spliced object feature are inputted into the feature fusion unit.
20. The image processing method according to any one of claims 13-19, the feature analysis unit comprises a spatial structure feature analysis subunit, a point cloud feature analysis subunit and an object feature analysis subunit; Correspondingly, the inputting the fused spatial structure feature, the fused point cloud feature and the fused object feature into the feature analysis unit to obtain a target spatial object model comprising the target spatial object and the target object comprises: The fused spatial structure feature is inputted into the spatial structure feature analysis subunit to obtain target spatial structure information of the target spatial object; The fused point cloud feature is inputted into the point cloud feature analysis subunit to obtain object grid information of the target object; The fused object feature is inputted into the object feature analysis subunit to obtain object contour information of the target object; According to the target spatial structure information, the object grid information and the object contour information, a target spatial object model comprising the target spatial object and the target object is determined.
21. The image processing method according to claim 20, before the inputting the fused point cloud feature into the point cloud feature analysis subunit, further comprising: determining a sample image and a label spatial object model comprising a sample spatial object and a sample object in the sample spatial object; inputting the sample image and the label spatial object model into the image processing model to obtain sample object grid information of the sample object outputted by the point cloud feature analysis subunit; calculating a collision loss function between the sample object and the sample spatial object according to a distance between the sample object grid information and the sample spatial object and the label spatial object model; and training the point cloud feature analysis subunit according to the collision loss function.
22. The image processing method according to any one of claims 13-21, the image processing model further comprises an image conversion unit; Correspondingly, before the inputting the depth image corresponding to the target image into the second feature extraction unit, further comprising: inputting the target image into the image conversion unit to obtain a depth image corresponding to the target image.
23. The image processing method according to any one of claims 13-18, before the inputting the spatial structure feature, the image feature, the point cloud feature and the object feature into the feature fusion unit, further comprising: determining a sample image and a label spatial object model comprising a sample spatial object and a sample object in the sample spatial object; inputting the sample image and the label space object model into the image processing model to obtain a sample space structure feature of the sample space object, a sample image feature of the sample image, a sample point cloud feature of the sample image, and a sample object feature of the sample object; performing random mask processing on the sample space structure feature, the sample image feature, the sample point cloud feature, and the sample object feature to obtain the sample space structure feature, the sample image feature, the sample point cloud feature, and the sample object feature after being masked; training the feature fusion unit according to the sample space structure feature, the sample image feature, the sample point cloud feature, and the sample object feature after being masked, and the label space object model.
24. An image processing method, comprising: determining a target image containing a target space object and a target object in the target space object; inputting the target image into an image processing model to obtain a target space object model containing the target space object and the target object; wherein the image processing model is a machine learning model, the image processing model comprises a first feature extraction unit, a second feature extraction unit, a feature fusion unit, and a feature analysis unit, the first feature extraction unit is configured to extract an image feature of the target image and a space structure feature of the target space object, the second feature extraction unit is configured to extract a point cloud feature of the target image and an object feature of the target object, the feature fusion unit is configured to fuse the space structure feature, the image feature, the point cloud feature, and the object feature, and the feature analysis unit is configured to obtain the target space object model containing the target space object and the target object.
25. An image processing method, comprising: determining a target image containing a target house and a target object in the target house; inputting the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house; wherein the image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model according to a house type feature of the target house, an image feature of the target image, a point cloud feature of the target image, an object feature of the target object, and a correlation relationship between the features.
26. An image processing method, comprising: receiving a house query request sent by a client, wherein the house query request comprises house information of a target house; determining a target image containing the target house and a target object in the target house according to the house information; inputting the target image into an image processing model to obtain a three-dimensional house model containing the target house and the target house; sending the three-dimensional house model to the client to be displayed through a display interface of the client. The image processing model is a machine learning model, and the three-dimensional house model is determined by the image processing model according to a house type feature of the target house, an image feature of the target image, a point cloud feature of the target image, an object feature of the target object, and a correlation between the features. 27.A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, so as to implement the steps of the method in any one of claims 1 to 26. 28.A computer readable storage medium storing computer programs / instructions, which, when executed by a processor, implement the steps of the method in any one of claims 1 to 26. 29.A computer program product comprising computer programs / instructions, which, when executed by a processor, implement the steps of the method in any one of claims 1 to 26.
Citation Information
Patent Citations
AR technology-based house inspection experience system
CN108022305A
House realistic picture display method and device, terminal equipment and storage medium
CN111145352A
Three-dimensional house model determination method and device, electronic equipment and storage medium
CN115761138A
Method for data collection and model generation of house
US20190378330A1