Three-dimensional model delivery method and device, equipment and storage medium
By receiving user voice data and scene image data to generate a 3D model that matches the environment, and accurately deploying it on the deployment reference plane, the problem of complex operation and insufficient adaptive capability in existing technologies is solved, achieving efficient and accurate 3D model deployment and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing methods for deploying 3D models are cumbersome, technically demanding, and lack adaptability to the deployment environment, resulting in inconsistencies between virtual models and real-world scenes, which negatively impacts user experience.
By receiving real-time user voice data and scene image data uploaded by mobile devices with augmented reality capabilities, an initial 3D model is generated. Through plane detection and spatial feature correction, a target 3D model matching the current environment is obtained, and finally, precise delivery is carried out on the delivery reference plane.
It improves the efficiency and accuracy of 3D modeling, enhances the realism and immersion of user interaction, reduces technical costs, avoids interaction failures caused by model incompatibility with the environment, and improves the effectiveness and experience of user operation.
Smart Images

Figure CN121767896A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of augmented reality technology, and in particular to a method, apparatus, device and storage medium for projecting a three-dimensional model. Background Technology
[0002] With the rapid development of Augmented Reality (AR) technology, its applications in fields such as gaming, education, industrial design, and marketing are becoming increasingly widespread. Against this backdrop, a key technical challenge in this field is how to accurately and naturally project 3D models into the real environment, achieving seamless visual integration with the physical space to create an immersive interactive experience for users.
[0003] However, existing methods for deploying 3D models primarily rely on manual operations performed by users on mobile devices, such as selecting templates from a pre-set model library or triggering model generation by entering text commands. These methods are not only cumbersome and technically demanding for non-professional users, but the generated 3D models often lack adaptability to the deployment environment, resulting in visual inconsistencies between the actual appearance and the real-world scene, thus negatively impacting the user experience.
[0004] Therefore, there is an urgent need to propose a new method to solve the above problems. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and storage medium for deploying 3D models, which improves the accuracy and efficiency of 3D model modeling and deployment, thereby optimizing the user experience.
[0006] In a first aspect, embodiments of the present invention provide a method for projecting a three-dimensional model, the method comprising:
[0007] Receive real-time user voice data and scene image data of the current environment uploaded by mobile devices with augmented reality capabilities;
[0008] An initial three-dimensional model is generated based on the real-time user voice data;
[0009] Plane detection is performed on the scene image data to obtain the projection reference plane;
[0010] Based on the spatial features of the scene image data, the parameters of the initial 3D model are corrected to obtain a target 3D model that matches the current environment.
[0011] The target 3D model and the projection reference plane are sent to the mobile device so that the mobile device can project the target 3D model onto the projection reference plane.
[0012] The technical solution of this invention first receives real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality capabilities, providing a data foundation for subsequent 3D modeling and model deployment. Next, an initial 3D model is generated based on the real-time user voice data, lowering the technical threshold for 3D modeling and making it more suitable for non-professional users, thus making the interaction process more intuitive and efficient. Simultaneously, generating the model directly through voice description can more accurately capture the user's specific expectations regarding shape and attributes, effectively reducing the problem of model discrepancies with expectations caused by manual operation deviations, thereby achieving high-precision matching of user semantic intent and improving applicability. Then, plane detection is performed on the scene image data to obtain the deployment reference plane, ensuring the spatial accuracy of the 3D model deployment and enhancing the realism and credibility of user interaction. Afterwards, the initial 3D model is parameter-corrected based on the spatial characteristics of the scene image data to obtain a target 3D model matching the current environment, effectively enhancing the realism and immersion of subsequent AR interactions, improving the environmental robustness of model deployment, and avoiding interaction failures caused by model-environment incompatibility (such as the model exceeding the screen range or overlapping with real objects), thereby significantly improving the effectiveness of user operations and reducing the technical costs of subsequent model deployment and interaction. Finally, the target 3D model and the projection reference plane are sent to the mobile device, enabling the mobile device to project the target 3D model onto the projection reference plane. This achieves accurate projection of the model in space, significantly enhancing the user's immersion and interactive experience. Therefore, the technical solution of this invention solves the problem in existing technologies where the lack of adaptability to the projection environment leads to inconsistencies between the visual representation of the virtual model and the real scene, thus affecting the user experience.
[0013] Secondly, embodiments of the present invention also provide a device for projecting a three-dimensional model, the device comprising:
[0014] The receiving module is used to receive real-time user voice data and scene image data of the current environment uploaded by mobile devices with augmented reality capabilities;
[0015] The generation module is used to generate an initial three-dimensional model based on the real-time user voice data;
[0016] The detection module is used to perform planar detection on the scene image data to obtain the projection reference plane;
[0017] The correction module is used to correct the parameters of the initial 3D model based on the spatial features of the scene image data to obtain a target 3D model that matches the current environment.
[0018] The delivery module is used to send the target 3D model and the delivery reference plane to the mobile device, so that the mobile device can deliver the target 3D model onto the delivery reference plane.
[0019] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0020] At least one processor; and a memory communicatively connected to said at least one processor;
[0021] The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the method for projecting a three-dimensional model according to any embodiment of the present invention.
[0022] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, implement the method for projecting a three-dimensional model as described in any embodiment of the present invention.
[0023] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the 3D model projection device, or it may be packaged separately from the processor of the 3D model projection device; this application does not impose any limitations on this.
[0024] The descriptions of the second, third, and fourth aspects in this application can be referenced to the detailed description of the first aspect; and the beneficial effects described in the second, third, and fourth aspects can be referenced to the analysis of the beneficial effects of the first aspect, which will not be repeated here.
[0025] In this application, the name of the device for projecting the 3D model does not limit the device or functional module itself. In actual implementation, these devices or functional modules may appear under other names. As long as the function of each device or functional module is similar to that of this application, it falls within the scope of the claims of this application and its equivalents.
[0026] These or other aspects of this application will become more readily apparent in the following description. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a method for projecting a three-dimensional model according to an embodiment of the present invention;
[0029] Figure 2 A flowchart illustrating another method for projecting a three-dimensional model provided in an embodiment of the present invention;
[0030] Figure 3 This is a schematic diagram of the structure of a three-dimensional model projection device provided in an embodiment of the present invention;
[0031] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0032] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0033] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0034] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0035] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0036] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subroutine, etc. Moreover, without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0037] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0038] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0039] Figure 1 This is a flowchart illustrating a method for projecting a 3D model according to an embodiment of the present invention. This embodiment is applicable to situations where users need to perform 3D modeling and projecting. The method can be executed by a 3D model projecting device, which can be implemented in software and / or hardware. For example, the device can be an electronic device. (See reference...) Figure 1 The method for projecting the 3D model in this embodiment specifically includes the following steps:
[0040] Step 110: Receive real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality capabilities.
[0041] Specifically, mobile devices with augmented reality capabilities are smart terminals equipped with augmented reality technology (such as smartphones, AR glasses, tablets, etc.). Real-time user voice data refers to voice information (such as commands, descriptive language, etc.) input by the user in real time through the mobile device's microphone. Scene image data of the current environment refers to a real-time image sequence (or video frames) of the real physical environment captured by the mobile device's camera, including information such as objects, textures, lighting, and spatial structure in the environment.
[0042] In practice, once the user completes voice input and scene image confirmation on the mobile device and clicks "Upload Data," the server receives the data packet from the device, thereby acquiring real-time user voice data and scene image data of the current environment. The specific process is as follows: After receiving the user's modeling instruction (e.g., clicking the "Start Modeling" button within the application), the mobile device activates its microphone to enter real-time voice acquisition mode, continuously recording the user's voice description (e.g., generating a 1-meter-high blue round table), while simultaneously displaying the voice input status in real-time on the user-interactive visual interface (e.g., waveform animation). After the user completes the voice description, they can click the "Shoot Environment" button to trigger the camera to acquire scene image data of the current environment (e.g., a single frame image or a short video frame sequence). After acquisition, the interface will display a preview image for user confirmation. If the user has no objection to the voice content and scene images, clicking the "Upload Data" button will cause the mobile device to package the voice data, which has undergone preliminary local noise reduction, and the scene image data, including image frames and synchronously recorded device sensor posture information, and send them to a server with data processing capabilities (such as a cloud platform) via the network. After receiving the data packet, the server will first perform integrity verification (such as verifying data format, metadata integrity, etc.). After the verification is successful, the data packet will be parsed to obtain the real-time user voice data and the scene image data of the current environment.
[0043] In this embodiment, the above steps provide a data foundation for subsequent 3D modeling and model deployment.
[0044] Step 120: Generate an initial 3D model based on real-time user voice data.
[0045] Specifically, the initial 3D model is a 3D virtual model generated based on user voice data.
[0046] In practice, speech recognition technology can be used to convert real-time user speech data into text descriptions. Then, the text is input into a pre-trained 3D generative model to obtain an initial 3D model. The 3D generative model is a model obtained by training a deep learning model (such as generative adversarial networks, diffusion models, etc.) on the text corresponding to the historical user speech data of different users and the accurate 3D virtual model that has been manually annotated or verified.
[0047] Optionally, to improve the quality of model generation, the following modular loss function configuration can be adopted during the training process of the 3D generation model: (1) Shape generation part: Chamfer distance and intersection-union ratio loss are used together. Chamfer distance is used to measure the difference between the predicted shape and the real shape in the point cloud distribution, while intersection-union ratio loss optimizes the overlap between the two in 3D space, jointly ensuring the structural accuracy of the generated shape. (2) Texture generation part: Perceptual loss and color consistency loss are used together. Perceptual loss is based on the pre-trained VGG network to extract high-level features of the image to compare the difference between the predicted texture and the real texture at the semantic level; color consistency loss constrains the stability of texture color under multiple views, avoids rendering inconsistency, and thus enhances the realism of the texture.
[0048] In addition, to improve training stability and model performance, a two-stage training strategy can be adopted: in the first stage, the texture generation sub-network is frozen and only the shape generation part is trained; in the second stage, the texture network is unfrozen and joint fine-tuning of shape and texture is performed to achieve overall collaborative optimization of generation.
[0049] In this embodiment, through the above steps, users do not need to master professional modeling tools or input complex parameters; they can quickly generate 3D models simply by describing them in natural language. This is especially suitable for non-professional users, making the interaction process more intuitive and efficient. Furthermore, generating models directly through voice description can more accurately capture users' specific expectations regarding shape and attributes, effectively reducing discrepancies between the model and expectations caused by manual operation errors. This achieves high-precision matching of the user's semantic intent and improves applicability.
[0050] Step 130: Perform planar detection on the scene image data to obtain the projection reference plane.
[0051] Specifically, the projection reference plane is a virtual reference plane (corresponding to the physical plane in the real environment) obtained by performing planar detection on the scene image data and used to place the 3D model. It includes spatial parameters such as position, angle, and size range.
[0052] In practice, the scene image data can be preprocessed first (e.g., distortion correction, brightness equalization, color normalization, noise reduction) to improve image quality and reduce the impact of ambient lighting changes and sensor noise on detection accuracy. Then, feature detection algorithms (e.g., oriented FAST and rotated BRIEF feature detection algorithms, scale-invariant feature transform) or deep learning-based keypoint detection networks are used to extract discriminative key feature points (e.g., texture intersections, object edge corners, planar texture inflection points) from the preprocessed image. A descriptor for each feature point (e.g., BRIEF descriptor, SIFT descriptor) is calculated to accurately characterize the pixel distribution patterns around the feature point. Based on the descriptor of each feature point, and combined with the device attitude data (such as angle and displacement) recorded in real time by the mobile device's sensors (such as gyroscopes and accelerometers), the simultaneous localization and mapping (SLAM) technology is used to calculate the three-dimensional coordinates of each feature point in the world coordinate system (such as the origin being the initial physical position when the mobile device first starts AR function and begins to collect scene data, and the X-axis, Y-axis, and Z-axis corresponding to the horizontal left and right, vertical up and down, and horizontal front and back coordinate systems in real space, respectively), thus obtaining a three-dimensional point cloud covering the current environment.
[0053] Next, a planar model is fitted from the 3D point cloud using a random sampling consensus algorithm: the candidate plane equation is calculated by randomly selecting three non-collinear feature points multiple times, and the distances from the remaining feature points to the plane are counted (if the distance is less than a preset error threshold, it is determined to be an interior point). Through iterative optimization, the plane with the most interior points and the smallest fitting error (average distance between interior points) is finally selected as the target candidate plane.
[0054] Then, the validity of the target candidate plane is judged (such as area verification, flatness verification, viewing angle verification, etc.). If it passes the verification, it is determined as the projection reference plane. If it fails, the plane with the second best number of interior points and fitting error is selected from the candidate planes recorded in the previous iteration as the target candidate plane, and the validity is verified again until a plane that passes the verification is found.
[0055] It should be noted that if the scene image data consists of continuous image frames (rather than a single frame), after obtaining the descriptor of each feature point, cross-frame matching is required using the feature point descriptor: brute-force matching or fast approximate nearest neighbor matching algorithms are used to associate feature points corresponding to the same physical point in different frames. Combined with the tracking and map update module of synchronous positioning and map building technology, the 3D coordinates of the feature points are dynamically optimized (such as removing abnormal feature points in motion-blurred frames and supplementing newly added feature points in new frames). Then, the 3D point cloud construction and plane fitting steps are performed.
[0056] In this embodiment, the above steps can ensure the spatial accuracy of the 3D model projection and enhance the realism and credibility of user interaction.
[0057] Step 140: Based on the spatial features of the scene image data, the parameters of the initial 3D model are corrected to obtain a target 3D model that matches the current environment.
[0058] Specifically, the spatial features of scene image data are information extracted from scene images to characterize the spatial attributes of the real environment. Examples of spatial features include scale features, texture features, depth features, obstacle features, lighting features, and pose features. The target 3D model is the final 3D model that, after parameter correction, matches the current environment in terms of scale, pose, and visual effects.
[0059] In practice, convolutional neural networks or depth estimation models can be used to extract features from scene image data to obtain spatial features. These spatial features, along with the structural parameters of the initial 3D model (such as vertex coordinates, size scale, and texture mapping information), are then input into a pre-trained correction model to obtain correction parameters for the initial 3D model (such as scaling factors, pose rotation matrices, and texture adjustment weights). The original parameters of the initial 3D model are then adjusted based on these correction parameters (e.g., modifying vertex coordinates by scaling factors and adjusting model pose by rotation matrices) to generate a target 3D model that matches the current environment. The correction model refers to the model obtained by training a deep learning model (such as a Transformer-based feature fusion model or an end-to-end parameter prediction network) using spatial features from a large number of historical scenes, the corresponding initial 3D model (such as the base model for user speech generation), and manually labeled or verified real-world adapted 3D models (i.e., the target model that perfectly matches the historical scene).
[0060] In this embodiment, the above steps effectively enhance the realism and immersion of subsequent AR interactions. They improve the environmental robustness of model deployment, avoiding interaction failures caused by model-environment incompatibility (such as models exceeding screen limits or overlapping with real objects), thereby significantly improving the effectiveness of user operations and reducing the technical costs of subsequent model deployment and interaction.
[0061] Step 150: Send the target 3D model and the projection reference plane to the mobile device so that the mobile device can project the target 3D model onto the projection reference plane.
[0062] In practice, after obtaining the target 3D model, the data of the target 3D model (such as vertex coordinates, texture maps, material parameters, etc.) can be integrated with the parameters of the projection reference plane (such as plane equations, center coordinates, normal vectors, and dimensions in the world coordinate system) in GLB or GLTF format. Then, the integrated data is compressed and packaged using a model compression algorithm (such as Draco) and sent to the mobile device. After receiving the data packet, the mobile device first decompresses and parses it to obtain the relevant parameters of the target 3D model and the projection reference plane. Then, it starts the built-in AR rendering engine (such as ARKit or ARCore), and combines the parsed model and reference plane parameters to project the target 3D model onto the projection reference plane.
[0063] In addition, users can make real-time adjustments to the deployed target 3D model through touch interaction on mobile devices to improve the flexibility of interaction and scene adaptability, such as: using two fingers to zoom to adjust the model size, using one or two fingers to rotate to adjust the model posture, and using one finger to drag to change the model position.
[0064] In this embodiment, the above steps achieve accurate model placement in space, thereby significantly enhancing the user's immersion and interactive experience.
[0065] The 3D model deployment method provided in this invention first receives real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality capabilities, providing a data foundation for subsequent 3D modeling and deployment. Next, an initial 3D model is generated based on the real-time user voice data, lowering the technical threshold for 3D modeling and making it more suitable for non-professional users, thus making the interaction process more intuitive and efficient. Simultaneously, generating the model directly through voice description can more accurately capture the user's specific expectations regarding shape and attributes, effectively reducing the problem of model discrepancies caused by manual operation deviations, thereby achieving high-precision matching of user semantic intent and improving applicability. Then, plane detection is performed on the scene image data to obtain the deployment reference plane, ensuring the spatial accuracy of the 3D model deployment and enhancing the realism and credibility of user interaction. Afterwards, the initial 3D model is parameter-corrected based on the spatial features of the scene image data to obtain a target 3D model matching the current environment, effectively enhancing the realism and immersion of subsequent AR interactions. This invention enhances the environmental robustness of model deployment, avoiding interaction failures caused by model-environment incompatibility (such as models exceeding screen limits or overlapping with real objects), thereby significantly improving the effectiveness of user operations and reducing the technical costs of subsequent model deployment and interaction. Finally, the target 3D model and deployment reference plane are sent to the mobile device, enabling the mobile device to deploy the target 3D model on the deployment reference plane, achieving precise model deployment in space and significantly improving user immersion and interactive experience. Therefore, the technical solution of this invention solves the problem of existing technologies lacking adaptability to the deployment environment, leading to inconsistencies between the visual representation of virtual models and real-world scenes, thus affecting user experience.
[0066] Figure 2 This is a flowchart illustrating another method for projecting a three-dimensional model according to an embodiment of the present invention. This embodiment is a specific implementation based on the above embodiment. In this embodiment, the method may further include:
[0067] Step 210: Receive real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality capabilities.
[0068] Step 211: Generate an initial 3D model based on real-time user voice data.
[0069] Further, step 211 may specifically include: parsing real-time user voice data to obtain user command text; performing semantic understanding processing on the user command text to obtain object attribute data; performing voxel generation processing on the object attribute data to obtain a three-dimensional voxel representation; converting the three-dimensional voxel representation into a triangular mesh; generating a material map corresponding to the triangular mesh based on the appearance attribute data in the object attribute data; and fusing and rendering the triangular mesh and the corresponding material map to obtain an initial three-dimensional model.
[0070] Specifically, user command text is the text content obtained by parsing and converting real-time user voice data. Object attribute data is the output of semantic understanding processing; it is a structured data set describing the features of the target object. For example, object attribute data includes geometric attribute data and appearance attribute data. Geometric attribute data describes the object's shape, size, structure, and other spatial features (e.g., 1.5 meters high, round tabletop, three drawers, etc.); appearance attribute data describes the object's surface visual features (e.g., color "red," material "wood," texture "wood grain," etc.). Three-dimensional voxel representation is a form of three-dimensional spatial data representation that discretizes the three-dimensional structure of the target object into a large number of regularly arranged cubic voxels, each carrying corresponding attribute information (e.g., whether it belongs to an object, color value, etc.). Triangular mesh is a commonly used form of three-dimensional model representation, a mesh structure formed by connecting a large number of triangular faces through vertices, capable of accurately describing the surface contour and geometric shape of the object. Material maps are two-dimensional image files generated based on the object's appearance attribute data, containing visual information such as the object's surface color, texture, reflectivity, and transparency.
[0071] In practice, speech recognition models (such as Transformer-based end-to-end ASR models, basic CTC-Conformer models, CTC-Attention hybrid models, Efficient Conformer-CTC, etc.) can be used to parse real-time user speech data to obtain user command text. Next, the user command text is normalized (e.g., removing interjections and correcting typos), and then the sentence is broken down into word units (e.g., red / round desktop / wooden / coffee table) using word segmentation tools (e.g., Jieba for Chinese, NLTK for English). Then, semantic understanding models (e.g., BERT pre-trained models) and named entity recognition are used to extract the core attributes of the object from the word segmentation results [e.g., geometric attributes (shape, structure, size), appearance attributes (color, material, texture), category attributes (object type, basic structure used to constrain the model)]. The extracted attributes are then integrated into structured data such as JSON, for example: {"Category": "Coffee Table", "Geometric": {"Desktop Shape": "Round", "Height": "0.45 meters"}, "Appearance": {"Color": "Red", "Material": "Wood"}}.
[0072] Next, a three-dimensional voxel space (e.g., a 100×100×50 voxel grid, where each voxel represents 1 cubic centimeter) is defined based on the object's geometric properties (such as size); at the same time, the origin and coordinate axes are set (e.g., the origin is the center of the bottom surface, and the Z-axis is the height direction). Then, the voxel filling logic is formulated based on the geometric properties [e.g., for a coffee table surface, a circular voxel area with a diameter of 0.6 meters is generated at a height of Z=0.4 meters (the voxel state is set to "exist"); the support is below the table surface (Z=0 to 0.4 meters), generating 4 columnar voxel structures to support the table surface].
[0073] After filling the voxel space according to the rules, each "existing" voxel carries basic attributes (such as belonging to a "tablet" or "support"), forming a 3D structural data of stacked voxels. Then, the moving cube algorithm is used to traverse the voxel space, identify the boundaries between "existing" and "non-existent" voxels, calculate isosurfaces, and generate triangular patches to extract the outer surface contour of the object. The initially generated triangular mesh can then be simplified and smoothed to obtain a triangular mesh model.
[0074] Next, based on the material information (e.g., "wood") in the appearance attributes, the corresponding base texture template (e.g., wood grain texture) is retrieved from the preset material library. If the user specifies a color (e.g., "red"), the base texture's tone is adjusted; otherwise, the default color is used. Simultaneously, the texture's scaling ratio is adjusted according to the object's structure (e.g., tabletop, support), generating a 2D texture file (e.g., PNG format) containing information such as color, diffuse, and specular highlights. UV mapping is then used to bind the texture coordinates to the vertices of the triangular mesh. The preset material library is a standardized database containing various mainstream material digital resources, pre-established according to actual conditions or needs.
[0075] Finally, the triangular mesh is associated with the material maps using a rendering engine (such as Blender or Unity's real-time rendering module), allowing the textures to be applied to the mesh surface according to UV mapping rules (e.g., wood grain texture accurately covering the coffee table surface). Default lighting effects such as ambient light and diffuse light can be added to simulate the visual appearance of objects under natural light (e.g., the natural reflection of a wooden surface) to enhance the model's three-dimensionality. After rendering, a data file (such as GLB format) of the initial 3D model with complete geometric structure and appearance features is output.
[0076] In this embodiment, the above steps can accurately match user needs, ensure model fidelity, and thus improve the accuracy of the generated initial 3D model.
[0077] Furthermore, after step 211, the method further includes: sending the initial 3D model to the mobile device and receiving user feedback modification information uploaded by the mobile device; correcting the initial 3D model based on the user feedback modification information to obtain an updated initial 3D model.
[0078] Specifically, user feedback modification information refers to the modification instructions input by the user on the device after viewing the initial 3D model on the mobile device, targeting parts of the model that do not meet expectations (such as being too large, having color deviations, or missing structures). User feedback modification information can be text commands, touch operation records, voice feedback, etc.
[0079] In practice, after obtaining the initial 3D model, it can be sent to a mobile device for user viewing, and user feedback and modification information uploaded by the mobile device can be received. Then, based on this information, the initial 3D model is corrected to obtain an updated initial 3D model. Specifically, the user feedback modification information can first be parsed using a semantic analysis model (such as a BERT-based text understanding model) to identify the modification requirements. Then, based on the parsing results, the modification object, modification dimension, and specific parameters are determined, and associated with the corresponding initial 3D model file to perform the modification operation. After the modification is completed, an updated initial 3D model is generated, and a modification log (such as a comparison of parameters before and after modification) can be recorded for subsequent traceability or secondary adjustments. For example, if the modification requirement involves size or posture adjustment, the vertex coordinates of the model's triangular mesh can be adjusted using 3D modeling tools to ensure that the geometry meets the requirements; if the modification requirement involves color or material, the corresponding material map is called from the preset material library, UV mapping and texture binding are re-performed, and the original material parameters are replaced.
[0080] Optionally, in order to quickly deploy the updated initial 3D model to the real scene and achieve a seamless connection between user needs and scene adaptation, after obtaining the updated initial 3D model, planar detection can be performed on the scene image data to obtain the projection reference plane; then the updated initial 3D model and the projection reference plane are sent to the mobile device so that the mobile device can project the updated initial 3D model on the projection reference plane.
[0081] In this embodiment, the above steps give users more control over model adjustments, optimize the interactive experience, and improve the matching degree between the model and user needs.
[0082] Step 212: Perform semantic segmentation and planar region recognition on the scene image to obtain each candidate planar region.
[0083] Specifically, the candidate planar region is a flat area selected from the scene image that has the potential to deploy a virtual model.
[0084] In the specific implementation, semantic segmentation models (such as Mask R-CNN, U-Net, etc.) are first used to perform semantic segmentation on the scene image to obtain semantic segmentation masks corresponding to different object categories. Based on preset matching rules and the category information of the initial 3D model, target semantic segmentation masks matching the model category are selected from all masks. Next, connected component analysis (such as using a flood fill algorithm) is performed on the target semantic segmentation masks to divide regions belonging to the same semantic category and with continuous pixels into multiple independent candidate region blocks. Then, edge detection algorithms are used, combined with geometric features such as the overlap between the contour and the fitted image, to verify the planar attributes of each candidate region block, excluding non-flat regions such as curved surfaces and concave / convex areas. Finally, the candidate region blocks that pass all planar feature verifications are determined as candidate planar regions. The preset matching rules refer to the mapping rules pre-set according to actual conditions or needs, used to associate the initial 3D model category with the semantic segmentation mask category. For example, for furniture models such as coffee tables and lamps, only semantic segmentation masks labeled with the tabletop and floor are matched.
[0085] In this embodiment, the above steps can filter effective planes, eliminate interference from invalid areas, reduce the computational cost of determining the subsequent deployment reference plane, and provide a reliable spatial basis for the accurate deployment of the virtual model.
[0086] Step 213: Calculate the three-dimensional spatial plane equation of each candidate plane region based on the image features of each candidate plane region, and obtain the three-dimensional spatial plane equation of each candidate plane region.
[0087] Specifically, image features are data used to describe the visual or spatial attributes of candidate planar regions. For example, image features can be appearance features, geometric features, depth features, etc. The three-dimensional spatial plane equation is a mathematical equation that accurately describes the spatial position and orientation of the candidate planar region in a three-dimensional world coordinate system.
[0088] In practice, for the current candidate plane region, the image features of the current candidate plane region can be extracted using image segmentation results and depth perception technology. For example, the pixel coordinate range of the region can be determined from the scene image, the depth values of uniformly sampled points (e.g., 30-50 points) in the region can be obtained through a depth camera or depth estimation model (e.g., MonoDepth), the camera intrinsic parameters of the shooting device can be retrieved at the same time, and finally the data can be summarized to obtain the image features.
[0089] Then, based on the camera pinhole imaging model, the image features are converted into three-dimensional spatial coordinates. For example, a camera coordinate system is established with the camera's optical center as the origin. Substituting the pixel coordinates, depth values, and camera intrinsic parameters of each sampling point, the three-dimensional coordinates of each sampling point in the camera coordinate system are calculated using the camera perspective inverse projection formula, forming a discrete set of three-dimensional coordinate points. The converted three-dimensional coordinate points are then preprocessed (e.g., outlier removal). Finally, the least squares method is used to perform plane fitting on the preprocessed three-dimensional coordinate points to obtain the three-dimensional spatial plane equation of the current candidate planar region.
[0090] In this embodiment, the above steps provide a data foundation for the subsequent calculation of the recommended values for each candidate planar region.
[0091] Step 214: Calculate the recommended value for each candidate plane region based on the preset geometric constraints and the three-dimensional spatial plane equation of each candidate plane region.
[0092] Specifically, the preset geometric constraints are geometric rules pre-defined based on actual conditions or needs, used to determine whether candidate planar regions are suitable for deploying virtual models. The recommended value is a numerical score obtained by quantifying the suitability of each candidate planar region based on the preset geometric constraints.
[0093] In specific implementation, the geometric constraints defined for virtual model deployment can be retrieved from the system's preset configuration, and corresponding scoring weights and calculation rules can be set for each constraint. For example: (1) Flatness constraint: The tilt angle of the candidate plane is required to be ≤5 degrees. The scoring rule is that if the tilt angle is ≤5 degrees, 25 points are awarded; 5 points are deducted for each degree exceeding the standard. (2) Area constraint: The actual area of the candidate plane is required to be ≥1.2 times the bottom area of the initial three-dimensional model. The scoring rule is that if the area meets the standard, 25 points are awarded; 5 points are deducted for each 10% lower than the standard value. (3) Distance constraint: The distance from the candidate plane to the mobile device is required to be within the range of 1-3 meters. The scoring rule is that if the distance is within the range, 25 points are awarded; if the distance is <1 meter or >3 meters, 5 points are deducted for each 0.5 meter deviation. (4) No occlusion constraint: There are no obvious occluders in the candidate plane. The scoring rule is that, combined with the depth data of the scene image, if more than 90% of the area in the plane is occluded, 25 points are awarded; 5 points are deducted for each 10% increase in the proportion of occluded area.
[0094] Then, for each candidate plane region, extract the key parameters of its three-dimensional space plane equation one by one, and calculate the score of a single constraint condition in combination with the above rules. For example: (1) Extract the normal vector of the plane equation, calculate the angle between the normal vector and the reference vector [such as the reference normal vector (0,0,1) of the horizontal plane] by vector dot product, that is, the plane tilt angle, and then give the flatness score according to the rules. (2) Determine the boundary range of the plane in three-dimensional space by plane equation, and then calculate the actual area of the plane by polygon area formula. After comparing with the bottom area of the model, give the area score according to the rules. (3) Select any point on the plane, combine the camera optical center coordinates, use the distance formula from point to plane to calculate the average distance, and obtain the distance score according to the rules. (4) Overlay the two-dimensional image region corresponding to the plane equation with the depth map, analyze the continuity of pixel depth values in the region, count the proportion of unoccluded areas, and give the unoccluded score according to the rules.
[0095] Finally, the scores of a single candidate planar region under all constraints are summed to obtain the final recommended value for that region. For example, the recommended value of a candidate planar region = flatness score + area score + distance score + unobstructed score.
[0096] In this embodiment, the above steps can assign quantitative evaluation indicators to the candidate planar region, thereby providing a data basis for the subsequent determination of the deployment reference surface.
[0097] Furthermore, after step 214, the method further includes: for the current candidate planar region, determining the scene priority coefficient of the current candidate planar region based on the attribute information of the initial 3D model; updating the recommended value of the current candidate planar region according to the scene priority coefficient of the current candidate planar region, and obtaining the updated recommended value of the current candidate planar region.
[0098] Specifically, attribute information refers to metadata inherent in the initial 3D model that describes its own inherent characteristics. For example, attribute information can be a category (such as decorative painting, coffee table, table, etc.). The scene priority coefficient is a quantified value of the scene adaptation level determined for the current candidate planar region based on the attribute information of the initial 3D model.
[0099] In the specific implementation, after calculating the recommended value of each candidate plane region based on preset geometric constraints and the 3D spatial plane equations of each candidate plane region, for the current candidate plane region, the scene priority coefficient of the current candidate plane region can be obtained by querying the model attribute and plane type priority correspondence table based on the attribute information of the initial 3D model and the type of the current candidate plane region (such as wall, tabletop, ground). Then, the recommended value of the current candidate plane region is updated according to the scene priority coefficient, resulting in the updated recommended value. The specific calculation formula is: Updated recommended value = Original recommended value × Scene priority coefficient. The model attribute and plane type priority correspondence table is a rule base pre-established according to the actual scene or requirements, used to store the mapping relationship from the "model attribute - plane type" combination to the "priority coefficient".
[0100] In this embodiment, the accuracy of the recommended values for the obtained candidate planar regions is improved through the above steps.
[0101] Step 215: Determine the candidate plane area with the highest recommendation value as the deployment reference plane.
[0102] In practice, after obtaining the recommended values of each candidate plane region, the candidate plane region with the highest recommended value can be determined as the projection reference plane to ensure that the subsequent 3D model can accurately and reasonably fit the real scene plane and avoid problems such as misaligned projection position and poor scene adaptability.
[0103] Step 216: Based on the spatial features of the scene image data, the parameters of the initial 3D model are corrected to obtain a target 3D model that matches the current environment.
[0104] Optionally, to further improve the adaptation efficiency and fusion accuracy of the target 3D model and the scene, the spatial features of the scene image data and the projection reference plane can be used to correct the parameters of the initial 3D model to obtain the target 3D model. For example, the spatial features of the scene image data, the spatial parameters of the projection reference plane and the initial 3D model are input into a pre-trained comprehensive correction model to obtain the target 3D model. The comprehensive correction model is a model obtained by training a deep learning model with historical spatial features under different scenes, corresponding historical projection reference plane parameters and historical initial 3D model data as input samples, and historical real 3D models that completely match the real scene (i.e. ideal correction results) as output labels.
[0105] Optional spatial features include illumination features, scale features, and texture features.
[0106] Further, step 216 may specifically include: correcting the geometric scale parameters of the initial 3D model based on scale features to obtain a reference model that is coordinated with the scene proportions; and correcting the visual appearance parameters of the reference model based on lighting features and texture features to obtain the target 3D model.
[0107] Specifically, scale features are features extracted from scene image data that reflect the actual size and spatial proportions of objects within the scene. Geometric scale parameters are core parameters inherent to the initial 3D model, describing its spatial dimensions, such as the model's length, width, height, and the relative dimensions of its components (e.g., the diameter of a coffee table top and the height of its legs). The baseline model is an intermediate model obtained after scale feature correction of the initial 3D model. Lighting features are features extracted from scene image data that describe ambient lighting conditions, such as the direction, intensity, color temperature, and shadow distribution within the scene. Texture features are features extracted from scene image data that describe the detailed texture style of object surfaces, such as the material texture, color tone, and texture clarity of objects within the scene. The visual appearance parameters of the baseline model are parameters that determine its visual presentation effect, such as the model's color, material, gloss, and transparency.
[0108] In practice, the scale features of the scene image data and the geometric scale parameters of the initial 3D model are first input into a pre-trained size correction model to obtain a baseline model that is proportionate to the scene. Then, the lighting features and texture features of the scene image data, along with the visual appearance parameters of the baseline model, are input into a pre-trained appearance correction model to obtain the target 3D model. The size correction model is obtained by supervised training of a deep learning model (such as a convolutional neural network or Transformer model) using a labeled dataset containing "scene scale features - model geometric parameters - standard adaptation results". The appearance correction model is obtained by training a deep learning model (such as a generative adversarial network or image inpainting model) using a labeled dataset containing "lighting features and texture features - model appearance parameters - visual fusion target".
[0109] In this embodiment, the above steps improve the accuracy of the obtained target 3D model, reduce the cost of manual adjustment, and improve modeling efficiency.
[0110] Furthermore, before step 216, the process includes: sending an automatic correction confirmation message to the mobile device and receiving user feedback information uploaded by the mobile device; if the user feedback information indicates agreement, triggering the execution of parameter correction of the initial 3D model based on the spatial features of the scene image data to obtain a target 3D model that matches the current environment.
[0111] Specifically, the automatic correction confirmation message is a request sent to the mobile device asking the user whether to allow the system to automatically correct the initial 3D model parameters. The user feedback message is the response instruction given by the mobile device to the "automatic correction confirmation message," which is divided into two states: agree and disagree.
[0112] In practice, before correcting the parameters of the initial 3D model based on the spatial features of the scene image data to obtain a target 3D model matching the current environment, an automatic correction confirmation message can be pushed to the user's mobile device so that the mobile device can display this information to the user. Then, user feedback information uploaded by the mobile device is received. If the user's feedback is "agree," the initial 3D model's parameters are corrected based on the spatial features of the scene image data to obtain a target 3D model matching the current environment. If the user's feedback is "disagree," the user can be redirected to the manual correction interface or the initial 3D model and the projection reference plane data can be sent to the mobile device so that the mobile device can directly project the initial 3D model onto the projection reference plane.
[0113] In this embodiment, the above steps not only improve interaction efficiency, but also enhance user experience by respecting the differences in users' subjective needs.
[0114] Step 217: Send the target 3D model and the projection reference plane to the mobile device so that the mobile device can project the target 3D model onto the projection reference plane.
[0115] The 3D model deployment method provided in this invention first receives real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality capabilities, providing a data foundation for subsequent 3D modeling and deployment. Next, an initial 3D model is generated based on the real-time user voice data, lowering the technical threshold for 3D modeling and making it more suitable for non-professional users, thus making the interaction process more intuitive and efficient. Simultaneously, generating the model directly through voice description can more accurately capture the user's specific expectations regarding shape and attributes, effectively reducing the problem of model discrepancies with expectations caused by manual operation deviations, thereby achieving high-precision matching of user semantic intent and improving applicability. Then, semantic segmentation and planar region recognition processing are performed on the scene image to obtain candidate planar regions. This allows for the filtering of effective planes and the elimination of interference from invalid regions, reducing the computational cost of determining the subsequent deployment reference plane and providing a reliable spatial foundation for the accurate deployment of the virtual model. Finally, the 3D spatial plane equations of the corresponding candidate planar regions are calculated based on the image features of each candidate planar region, providing a data foundation for the subsequent calculation of recommended values for each candidate planar region. Based on preset geometric constraints and the 3D spatial plane equations of each candidate plane region, recommended values are calculated for each candidate plane region, assigning quantitative evaluation indicators to them and providing a data foundation for determining the subsequent deployment reference plane. Then, the candidate plane region with the highest recommended value is determined as the deployment reference plane to ensure that the subsequent 3D model accurately and reasonably fits the real scene plane, avoiding problems such as misaligned deployment positions and poor scene adaptability. Next, the initial 3D model is parameter-corrected based on the spatial characteristics of the scene image data to obtain a target 3D model that matches the current environment, effectively enhancing the realism and immersion of subsequent AR interactions. This improves the environmental robustness of model deployment, avoiding interaction failures caused by model incompatibility with the environment (such as the model exceeding the screen range or overlapping with real objects), thus significantly improving the effectiveness of user operations and reducing the technical costs of subsequent model deployment and interaction. Finally, the target 3D model and the deployment reference plane are sent to the mobile device, enabling the mobile device to deploy the target 3D model on the deployment reference plane, achieving accurate model deployment in space and significantly improving the user's immersion and interactive experience. Therefore, the technical solution of the present invention solves the problem that the existing technology lacks the ability to adapt to the deployment environment, resulting in the visual performance of the virtual model being inconsistent with the real scene, which in turn affects the user experience.
[0116] Figure 3 This is a schematic diagram of a three-dimensional model projection device provided in an embodiment of the present invention. This device belongs to the same inventive concept as the three-dimensional model projection methods in the above embodiments. Details not described in detail in the embodiments of the three-dimensional model projection device can be found in the embodiments of the three-dimensional model projection methods described above. Figure 3As shown, the device includes:
[0117] like Figure 3 As shown, the device includes:
[0118] The receiving module 310 is used to receive real-time user voice data and scene image data of the current environment uploaded by a mobile device with augmented reality function;
[0119] Generation module 320 is used to generate an initial three-dimensional model based on the real-time user voice data;
[0120] Detection module 330 is used to perform planar detection on the scene image data to obtain the projection reference plane;
[0121] The correction module 340 is used to correct the parameters of the initial three-dimensional model based on the spatial features of the scene image data to obtain a target three-dimensional model that matches the current environment.
[0122] The delivery module 350 is used to send the target 3D model and the delivery reference plane to the mobile device, so that the mobile device can deliver the target 3D model onto the delivery reference plane.
[0123] Based on the above embodiments, the generation module 320 is specifically used for:
[0124] The real-time user voice data is parsed to obtain user command text; the user command text is semantically understood to obtain object attribute data; the object attribute data is voxel-generated to obtain a three-dimensional voxel representation; the three-dimensional voxel representation is converted into a triangular mesh; a material map corresponding to the triangular mesh is generated based on the appearance attribute data in the object attribute data; the triangular mesh and the corresponding material map are merged and rendered to obtain an initial three-dimensional model.
[0125] Based on the above embodiments, the detection module 330 is specifically used for:
[0126] Semantic segmentation and planar region recognition are performed on the scene image to obtain candidate planar regions; the three-dimensional spatial plane equation of the corresponding candidate planar region is calculated based on the image features of the candidate planar regions to obtain the three-dimensional spatial plane equation of each candidate planar region; the recommended value of each candidate planar region is calculated based on the preset geometric constraints and the three-dimensional spatial plane equation of each candidate planar region; the candidate planar region with the highest recommended value is determined as the projection reference plane.
[0127] Based on the above embodiments, the device further includes:
[0128] The update module is used to calculate the recommended value of each candidate plane region based on preset geometric constraints and the three-dimensional spatial plane equation of each candidate plane region, and then, for the current candidate plane region, determine the scene priority coefficient of the current candidate plane region based on the attribute information of the initial three-dimensional model; update the recommended value of the current candidate plane region according to the scene priority coefficient of the current candidate plane region, and obtain the updated recommended value of the current candidate plane region.
[0129] Based on the above embodiments, the spatial features include illumination features, scale features, and texture features; the correction module 340 is specifically used for:
[0130] The geometric scale parameters of the initial 3D model are corrected based on the scale features to obtain a reference model that is in harmony with the scene proportions; the visual appearance parameters of the reference model are corrected based on the lighting features and the texture features to obtain the target 3D model.
[0131] Based on the above embodiments, the device further includes:
[0132] The user correction module is used to send the initial 3D model to the mobile device after generating the initial 3D model based on the real-time user voice data, and to receive user feedback modification information uploaded by the mobile device; and to correct the initial 3D model based on the user feedback modification information to obtain an updated initial 3D model.
[0133] Based on the above embodiments, the device further includes:
[0134] The user confirmation module is used to send automatic correction confirmation information to the mobile device before correcting the parameters of the initial 3D model based on the spatial features of the scene image data to obtain a target 3D model that matches the current environment, and to receive user feedback information uploaded by the mobile device; if the user feedback information indicates agreement, it triggers the execution of parameter correction of the initial 3D model based on the spatial features of the scene image data to obtain a target 3D model that matches the current environment.
[0135] The three-dimensional model casting device provided in the embodiments of the present invention can execute the three-dimensional model casting method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0136] It is worth noting that in the embodiments of the above-mentioned three-dimensional model projection device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0137] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Figure 4 A block diagram of an exemplary electronic device 4 suitable for implementing embodiments of the present invention is shown. Figure 4 The electronic device 4 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0138] like Figure 4 As shown, electronic device 4 is represented in the form of a general-purpose computing electronic device. The components of electronic device 4 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0139] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0140] Electronic device 4 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 4, including volatile and non-volatile media, removable and non-removable media.
[0141] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 4 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (… Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0142] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0143] Electronic device 4 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with electronic device 4, and / or with any device that enables electronic device 4 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through input / output (I / O) interface 22. Furthermore, electronic device 4 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. Figure 4 As shown, network adapter 20 communicates with other modules of electronic device 4 via bus 18. It should be understood that, although... Figure 4 Not shown, it can be combined with electronic device 4 to use other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0144] The processing unit 16 executes various functional applications and page displays by running programs stored in the system memory 28, such as implementing the three-dimensional model projection method provided in the embodiments of the present invention. Of course, those skilled in the art will understand that the processor can also implement the technical solutions of the three-dimensional model projection method provided in any embodiment of the present invention.
[0145] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements, for example, the method for projecting a three-dimensional model provided in this invention.
[0146] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0147] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0148] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0149] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0150] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0151] Furthermore, the acquisition, storage, use, and processing of data in the technical solution of this invention all comply with relevant laws and regulations.
[0152] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A method of dropping a three-dimensional model, characterized by, The method comprises: receiving real-time user voice data uploaded by a mobile device with augmented reality function and scene image data of a current environment; generating an initial three-dimensional model based on the real-time user voice data; performing plane detection on the scene image data to obtain a placement reference surface; performing parameter correction on the initial three-dimensional model based on spatial features of the scene image data to obtain a target three-dimensional model matched with the current environment; sending the target three-dimensional model and the placement reference surface to the mobile device to enable the mobile device to place the target three-dimensional model on the placement reference surface.
2. The method of claim 1, wherein, The method for generating an initial three-dimensional model based on real-time user voice data comprises: parsing the real-time user voice data to obtain user instruction text; performing semantic understanding processing on the user instruction text to obtain object attribute data; performing voxel generation processing on the object attribute data to obtain three-dimensional voxel representation; converting the three-dimensional voxel representation into a triangular mesh; generating a material map corresponding to the triangular mesh based on appearance attribute data in the object attribute data; performing fusion rendering on the triangular mesh and the corresponding material map to obtain an initial three-dimensional model.
3. The method of claim 1, wherein, The method for performing plane detection on the scene image data to obtain a placement reference surface comprises: performing semantic segmentation and plane area identification processing on the scene image to obtain each candidate plane area; calculating a three-dimensional spatial plane equation of each candidate plane area based on image features of the candidate plane areas to obtain the three-dimensional spatial plane equation of each candidate plane area; calculating a recommendation value of each candidate plane area based on a preset geometric constraint condition and the three-dimensional spatial plane equation of each candidate plane area; determining the candidate plane area with the highest recommendation value as the placement reference surface.
4. The method of claim 3, wherein, After calculating the recommendation value of each candidate plane area based on a preset geometric constraint condition and the three-dimensional spatial plane equation of each candidate plane area, the method further comprises: for a current candidate plane area, determining a scene priority coefficient of the current candidate plane area based on attribute information of the initial three-dimensional model; updating the recommendation value of the current candidate plane area according to the scene priority coefficient of the current candidate plane area to obtain an updated recommendation value of the current candidate plane area.
5. The method of claim 1, wherein, The spatial features comprise illumination features, scale features, and texture features. The method for performing parameter correction on the initial three-dimensional model based on spatial features of the scene image data to obtain a target three-dimensional model matched with the current environment comprises: correcting geometric scale parameters of the initial three-dimensional model based on the scale features to obtain a reference model coordinated with the scene scale; correcting visual appearance parameters of the reference model based on the illumination features and the texture features to obtain the target three-dimensional model.
6. The method of claim 1, wherein, After generating an initial three-dimensional model based on real-time user voice data, the method further comprises: sending the initial three-dimensional model to the mobile device and receiving user feedback modification information uploaded by the mobile device; correcting the initial three-dimensional model based on the user feedback modification information to obtain an updated initial three-dimensional model.
7. The method of claim 1, wherein, Before the initial three-dimensional model is parameter-corrected based on the spatial features of the scene image data to obtain a target three-dimensional model matching the current environment, the method further comprises: sending automatic correction confirmation information to the mobile device and receiving user feedback information uploaded by the mobile device; if the user feedback information is an approval, triggering execution of parameter correction of the initial three-dimensional model based on the spatial features of the scene image data to obtain a target three-dimensional model matching the current environment.
8. A device for dispensing a three-dimensional model, characterized in that The device comprises: a receiving module configured to receive real-time user voice data uploaded by a mobile device with an augmented reality function and scene image data of a current environment; a generating module configured to generate an initial three-dimensional model based on the real-time user voice data; a detecting module configured to perform plane detection on the scene image data to obtain a placement reference surface; a correcting module configured to perform parameter correction of the initial three-dimensional model based on the spatial features of the scene image data to obtain a target three-dimensional model matching the current environment; a placing module configured to send the target three-dimensional model and the placement reference surface to the mobile device to enable the mobile device to place the target three-dimensional model on the placement reference surface.
9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the three-dimensional model placing method of any one of claims 1-7.
10. A storage medium containing computer-executable instructions, wherein: The computer executable instructions, when executed by a computer processor, are used to execute the three-dimensional model placing method of any one of claims 1-7.