A method, apparatus, storage medium, and electronic device for generating a three-dimensional model view
By using pre-trained orientation discriminant model to adjust the orientation and render the three-dimensional model in multiple views, the problems of slow loading speed of 3D models and low efficiency in generating multi-angle views are solved, and multiple views are generated quickly and efficiently, supporting the processing of large-scale 3D models.
Patent Information
- Application Number
- CN202510345872.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing technology has challenges in the slow loading speed of three-dimensional models, low efficiency in generating multi-angle views, and insufficient demand for diversification, which limits the further application and development of cultural relics digitalization.
By obtaining the target three-dimensional model and adjusting its orientation using a pre-trained orientation discriminant model, multi-view rendering is performed based on the adjusted three-dimensional model to generate rendering results containing multiple views.
It realizes the rapid and efficient generation of multiple three-dimensional model views, significantly improves the generation speed and shortens the rendering time, supports the processing of large-scale three-dimensional models, and meets the needs of digitizing a large number of cultural relics.
Smart Images

Figure CN119863600B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a method, apparatus, storage medium, and electronic device for generating three-dimensional model views. Background Art
[0002] With the rapid development of three-dimensional digitization of cultural relics and landscapes, globally, the trend of using three-dimensional scanning and modeling technology for cultural relic digitization has become increasingly obvious. Technologies such as laser scanning and photogrammetry are widely used in high-precision modeling of cultural relics. Many international museums have established virtual museums, using three-dimensional models to display their collections, providing detailed background information and interactive functions.
[0003] However, the existing technology still faces challenges such as slow loading of three-dimensional models, low efficiency in generating multi-angle views, and insufficient response to diverse requirements, which limit the further application and development of cultural relic digitization.
[0004] Therefore, how to efficiently generate views of three-dimensional models has become a technical problem that needs to be urgently solved by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method, apparatus, storage medium, and electronic device for generating three-dimensional model views that overcome or at least partially solve the above problems. The technical solutions are as follows:
[0006] A method for generating three-dimensional model views includes:
[0007] Obtain a target three-dimensional model;
[0008] Use a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model to obtain the target three-dimensional model with the adjusted orientation;
[0009] Based on the target three-dimensional model with the adjusted orientation, perform multi-view rendering on the target three-dimensional model according to preset parameters to obtain a view rendering result, where the view rendering result includes multiple views of the target three-dimensional model.
[0010] Optionally, the step of using a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model to obtain the target three-dimensional model with the adjusted orientation includes:
[0011] Input the target 3D model into the orientation discrimination model, so that the orientation discrimination model extracts the first geometric features and the first functional semantic features of the target 3D model, and performs orientation prediction based on the first geometric features and the first functional semantic features to obtain the output first predicted rotation parameters, and use the first predicted rotation parameters to rotate the target 3D model to obtain the target 3D model with the adjusted orientation.
[0012] Optionally, the training process of the orientation discrimination model includes:
[0013] Obtain a 3D model training set, where the 3D model training set includes a plurality of standardized and already orientation-adjusted 3D training models;
[0014] Input each of the 3D training models in the 3D model training set into the orientation discrimination model for training, extract the second geometric features and the second functional semantic features of the 3D training model, and then rotate the 3D training model according to the random rotation parameters, and extract the third geometric features and the third functional semantic features of the rotated 3D training model;
[0015] Use the second geometric features and the third geometric features to obtain a geometric consistency loss;
[0016] Use the second functional semantic features and the third functional semantic features to obtain a functional semantic constraint loss;
[0017] Based on the third geometric features and the third functional semantic features, obtain the second predicted rotation parameters, and use the second predicted rotation parameters and the random rotation parameters to obtain an orientation prediction loss;
[0018] Use the geometric consistency loss, the functional semantic constraint loss and the orientation prediction loss to obtain a comprehensive loss;
[0019] Based on the comprehensive loss, adjust the network parameters of the orientation discrimination model and retrain until the preset training end condition is reached to obtain the trained orientation discrimination model.
[0020] Optionally, the standardization process includes:
[0021] Obtain a 3D training model;
[0022] Move the geometric center of the 3D training model to the target origin, and adjust the 3D training model to a preset specification size to obtain the standardized 3D training model.
[0023] Optionally, the geometric features include: local differential geometric features, mesoscale structural features, and global structural features.
[0024] Optionally, after performing multi-view rendering on the target 3D model according to preset rendering parameters based on the target 3D model with the adjusted orientation to obtain a view rendering result, the method further includes:
[0025] Using a cover image selection model to select a view from the view rendering result as the cover image of the target 3D model.
[0026] Optionally, the preset parameters include: light source parameters, camera parameters, and rendered image parameters.
[0027] A 3D model view generation device includes: a 3D model acquisition unit, an orientation adjustment unit, and a view rendering result acquisition unit.
[0028] The 3D model acquisition unit is configured to acquire a target 3D model;
[0029] The orientation adjustment unit is configured to use a pre-trained orientation discrimination model to adjust the orientation of the target 3D model to obtain the target 3D model with the adjusted orientation;
[0030] The view rendering result acquisition unit is configured to perform multi-view rendering on the target 3D model according to preset parameters based on the target 3D model with the adjusted orientation to obtain a view rendering result, where the view rendering result includes multiple views of the target 3D model.
[0031] A computer-readable storage medium stores a program, and when the program is executed by a processor, the 3D model view generation method is implemented.
[0032] An electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is configured to call program instructions in the memory to execute the 3D model view generation method.
[0033] With the above technical solution, a method, device, storage medium, and electronic device for generating a three-dimensional model view provided by the present invention obtain a target three-dimensional model; use a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model to obtain a target three-dimensional model with an adjusted orientation; based on the target three-dimensional model with an adjusted orientation, perform multi-view rendering on the target three-dimensional model according to preset parameters to obtain a view rendering result, where the view rendering result includes multiple views of the target three-dimensional model. The present invention adjusts the orientation of the target three-dimensional model through an automated technique and performs rendering with preset parameters based on this, thereby quickly and efficiently generating multiple views, thus supporting the processing of a large number of three-dimensional models, significantly improving the generation speed and shortening the rendering time to meet the needs of a large number of cultural relic digitizations.
[0034] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0036] Figure 1 It shows a schematic flowchart of an implementation manner of the method for generating a three-dimensional model view provided by an embodiment of the present invention;
[0037] Figure 2 It shows a schematic diagram of the model structure of the orientation discrimination model provided by an embodiment of the present invention;
[0038] Figure 3 It shows a schematic flowchart of the training process of the orientation discrimination model provided by an embodiment of the present invention;
[0039] Figure 4 It shows a logical block diagram of the self-supervised training process provided by an embodiment of the present invention;
[0040] Figure 5 It shows a logical block diagram of a method for generating a three-dimensional model view provided by an embodiment of the present invention;
[0041] Figure 6 It shows a schematic diagram of the structure of the device for generating a three-dimensional model view provided by an embodiment of the present invention;
[0042] Figure 7 It shows a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. Detailed implementation mode
[0043] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0044] With the continuous development of the trend of three-dimensional digitization of cultural relic landscapes, the application of three-dimensional models is increasing. However, three-dimensional models are usually larger in volume than static pictures, resulting in significantly slower loading and rendering speeds than pictures. This problem is particularly prominent when users search or browse three-dimensional models in the information stream. In order to allow users to preview three-dimensional models more intuitively, it is usually necessary to generate cover images and multiple perspective views based on three-dimensional data, generally including six different perspectives: up, down, left, right, front, and back. Currently, these covers and six views mostly rely on manual generation, and the operation process is cumbersome. Usually, it is necessary to open the three-dimensional model, adjust the cultural relic to the corresponding direction, and perform rendering output, with relatively low overall efficiency. The number of manual operations for each cultural relic reaches up to seven times, and when the background color needs to be changed, a large amount of manpower is also required for repeated operations.
[0045] Globally, the trend of cultural relic digitization is becoming more and more obvious, and the application of three-dimensional model technology in cultural relic protection, display, and research is becoming more and more extensive. Foreign countries are in the leading position in the field of three-dimensional scanning and modeling technology, and use technologies such as laser scanning, structured light scanning, and photogrammetry to perform high-precision digital modeling of cultural relics. For example, the British Museum and the Pompeii archaeological site in Italy have both used advanced three-dimensional scanning technology to record and study cultural relics and sites in detail. At the same time, many internationally renowned museums, such as the Louvre Museum and the Metropolitan Museum of Art in New York, have created virtual museums to display their collections of cultural relics through three-dimensional models, enabling global audiences to conduct virtual visits through the Internet. These platforms not only display the appearance of cultural relics but also provide detailed background information and interactive functions.
[0046] In China, the Palace Museum is at the forefront of cultural relic digitization. Through the "Digital Multibay" project, it has created high-precision three-dimensional cultural relic models using digitalization, cloud computing, and artificial intelligence technologies and displayed hundreds of digitized cultural relics. In addition, the construction of the Digital Twin Platform of the Longmen Grottoes has also provided tourists with a brand-new digital tourism experience, and domestic cultural and museum institutions have also made remarkable progress in digital acquisition and online exhibition.
[0047] Although the existing technology has achieved remarkable achievements in 3D scanning and modeling, there are still some challenges in the application and display of 3D models. First, the loading and rendering speed of 3D models is relatively slow, directly affecting the user experience. Second, the generation of covers and multi-angle views still relies on manual operations, which is inefficient and costly. In addition, the existing technology lacks flexibility in meeting diverse requirements and is difficult to quickly adjust and batch process, which to a certain extent restricts the further development and application of cultural relic digitalization.
[0048] Based on this, in the embodiments of the present invention, a method for generating 3D model views is provided. First, a target 3D model is obtained, and the orientation of the target 3D model is adjusted using a pre-trained orientation discrimination model to obtain an adjusted target 3D model. Second, based on the adjusted target 3D model, multi-view rendering is performed according to preset parameters to generate a rendering result including multiple views. The present invention realizes the orientation adjustment of the target 3D model and the rendering based on preset parameters through automation technology, and quickly and efficiently generates multiple views. This process supports the processing of large-scale 3D models, significantly improves the view generation speed and shortens the rendering time, and effectively meets the needs of a large number of cultural relic digitalizations.
[0049] As Figure 1 shown, a schematic flowchart of an implementation manner of the method for generating 3D model views provided by the embodiments of the present invention, the method may include:
[0050] S100. Obtain a target 3D model.
[0051] Among them, a three-dimensional model (Three-Dimensional Model, 3D Model) refers to a three-dimensional computer model created for a specific object (such as cultural relics, buildings, natural landscapes, etc.) during the three-dimensional digitization process. The three-dimensional model captures the shape, structure, and details of the object through digital technology for display, analysis, and interaction in a computer environment.
[0052] Specifically, in the process of obtaining the target 3D model in the embodiments of the present invention, a 3D model processing tool (such as Blender) can be used to normalize the model. This process includes moving the geometric center of the model to the origin and standardizing the length, width, and height of the model. It should be noted that this step is not limited to a specific 3D model editing tool, and users can also choose other open-source libraries and tools. Through this standardization process, consistent and standardized input data can be provided for the subsequent orientation judgment model, thereby improving the accuracy and efficiency of processing.
[0053] S110. Use a pre-trained orientation discrimination model to adjust the orientation of the target 3D model to obtain an adjusted target 3D model with a good orientation.
[0054] Among them, the orientation discrimination model is a pre-trained machine learning model used to determine the direction or orientation of a 3D model. By analyzing the geometric features of the 3D model, the orientation discrimination model can identify the correct orientation of the 3D model, thereby providing guidance for subsequent processing.
[0055] Specifically, in the embodiments of the present invention, the orientation of the input target 3D model can be judged according to the orientation discrimination model, and then the target 3D model can be rotated or flipped accordingly according to the judgment result to achieve the correct orientation. Through orientation adjustment, the user can obtain a target 3D model with consistent orientation and in line with the usage scenario, thereby improving the effect and accuracy of the 3D model in display, analysis or other applications.
[0056] Furthermore, in the embodiments of the present invention, the orientation discrimination model can make full use of the geometric structure (mesh) and functional semantic information (including texture information and shape information) of the 3D model, and combine the two to realize the judgment of the orientation of the 3D model.
[0057] S120. Based on the target 3D model with adjusted orientation, perform multi-view rendering on the target 3D model according to preset parameters to obtain a view rendering result, where the view rendering result includes multiple views of the target 3D model.
[0058] Among them, the preset parameters refer to a set of parameters set in advance when performing 3D rendering, graphics processing or other computer vision-related tasks on a 3D model. These preset parameters are used to define the rendering environment and conditions to ensure the consistency and controllability of the results. By providing the preset parameters, the user can achieve the expected visual effect during the rendering process and ensure the unity and repeatability of the rendering results.
[0059] Optionally, the preset parameters may include: light source parameters, camera parameters, and rendering picture parameters.
[0060] Among them, the light source parameters refer to various parameters regarding the light source settings during 3D rendering. The light source parameters affect the lighting effect and visual performance of the 3D model in the scene. The light source parameters may include light source type, light source position, light source orientation, light source color, light intensity, and attenuation rate.
[0061] Among them, the camera parameters refer to various parameters regarding the virtual camera settings during rendering. The camera parameters determine the viewing angle of the 3D scene and its visual effect. The camera parameters may include camera position, camera orientation, field of view (FOV), focal length, and depth of field.
[0062] Among them, the rendering image parameters refer to the settings when generating the final view of the three-dimensional model. The rendering image parameters directly affect the output quality and format of the view. The rendering image parameters may include image resolution, image format, background color, size, compression ratio, and post-processing effects.
[0063] Among them, multi-view rendering refers to the process of rendering the same three-dimensional model from multiple different angles and positions. In the embodiments of the present invention, by adjusting the position and orientation of the virtual camera, images of the three-dimensional model from multiple perspectives can be generated. Different views respectively show different sides of the three-dimensional model, helping users to more comprehensively understand the shape, details, and structure of the three-dimensional model.
[0064] Among them, the view rendering result is a set of images generated after multi-view rendering. Each view in the view rendering result is the result of rendering the target three-dimensional model from a specific angle and position. These views together show the appearance of the three-dimensional model from different perspectives and can be used for analysis, display, or further processing. The view rendering result may include multiple views among the front view, rear view, left view, right view, top view, bottom view, and the view at a 45-degree angle in the direct front of the target three-dimensional model.
[0065] Specifically, in the embodiments of the present invention, the target three-dimensional model with the adjusted orientation can be imported into a specified rendering software, a corresponding scene can be created in the rendering software, the light source and camera perspective can be set according to the preset parameters, and then the rendering process can be started to render each configured camera perspective one by one. After completion, the corresponding image file is saved to obtain the view rendering result.
[0066] A method for generating a three-dimensional model view provided by the present invention includes: obtaining a target three-dimensional model; using a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model to obtain a target three-dimensional model with the adjusted orientation; based on the target three-dimensional model with the adjusted orientation, performing multi-view rendering on the target three-dimensional model according to preset parameters to obtain a view rendering result, where the view rendering result includes multiple views of the target three-dimensional model. The present invention adjusts the orientation of the target three-dimensional model through an automated technology and performs preset parameter rendering based on this, thereby quickly and efficiently generating multiple views, thus supporting the processing of a large number of three-dimensional models, significantly improving the generation speed and shortening the rendering time to meet the needs of a large number of cultural relic digitizations.
[0067] Optionally, based on one or more corresponding embodiments described above, in another optional embodiment provided by the embodiments of the present invention, using a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model to obtain a target three-dimensional model with the adjusted orientation may specifically include: Figure 1
[0068] Input the target 3D model into the orientation discrimination model, so that the orientation discrimination model extracts the first geometric features and the first functional semantic features of the target 3D model, and performs orientation prediction based on the first geometric features and the first functional semantic features to obtain the output first predicted rotation parameters. Use the first predicted rotation parameters to rotate the target 3D model to obtain the target 3D model with the adjusted orientation.
[0069] Among them, geometric features refer to the attributes that describe the shape, structure, and spatial relationship of a 3D model. Geometric features usually include information such as the boundary, angle, surface curvature, volume, and area of the 3D model.
[0070] Among them, the functional semantic feature function refers to the attributes related to the function, use, and meaning of the 3D model. The functional semantic feature function usually involves the role and usage mode of the object corresponding to the 3D model in a specific context.
[0071] Among them, the rotation parameter is a numerical value used to describe the rotation of a 3D model in 3D space. When performing orientation adjustment, the rotation parameter usually represents information such as the axis and angle of rotation, so as to correctly orient the 3D model in the desired direction.
[0072] The model structure of the orientation discrimination model provided by the embodiments of the present invention can be as Figure 2 shown. After the 3D model is input into the orientation discrimination model, it undergoes feature extraction through the functional semantic feature module and the multi-scale geometric feature module respectively, and then after being processed by the encoder and the decoder, the geometric features and functional semantic features of the 3D model are output, and then orientation prediction is performed based on the geometric features and functional semantic features.
[0073] Optionally, the geometric features provided by the embodiments of the present invention may include local differential geometric features, mesoscale structural features, and global structural features.
[0074] First, the embodiments of the present invention can first extract the local differential geometric features of the 3D model. For example: the principal curvature, mean curvature, Gaussian curvature, and normal vector of each point on the 3D model, that is:
[0075]
[0076] Among them, is a point on the 3D model; is the local differential geometric feature of point ; , are the principal curvatures of point ; is the mean curvature of point , that is ; is the point The Gaussian curvature of, i.e., ; is the normal vector of point .
[0077] Then, perform a connection operation on the local differential geometric features to obtain the mesoscale structure features, i.e.:
[0078]
[0079] Wherein, is an object composed of multiple points on the 3D model; is the mesoscale structure feature of object ; represents the graph connection operation; is the point within the range of object ; represents the component connection relationship, which is used to depict the connection relationship between each local element within object , and is usually reflected by a topological graph and an adjacency matrix.
[0080] Next, perform a symmetry and pooling operation on the mesoscale structure features to obtain the global structure features, i.e.:
[0081]
[0082] Wherein, represents the global structure feature of the 3D model; represents the pooling operation; represents the symmetry operation.
[0083] Finally, use the local differential geometric features, mesoscale structure features, and global structure features as the geometric features of the 3D model, and combine them with the functional semantic features of the 3D model for orientation prediction.
[0084] Optionally, in the process of extracting the functional semantic features of the 3D model in the embodiments of the present invention, the 3D model can be first divided into components, and on this basis, each functional component can be functionally classified, i.e.:
[0085]
[0086] Wherein, represents the probability that point belongs to the component category under the condition of a given point on the surface of the 3D model ; is a function about point , which is used to extract the feature information related to the functional component in the area where this point is located; is the activation function.
[0087] Then, based on each functional component, a functional relationship diagram is constructed, that is:
[0088]
[0089] Among them, represents the functional relationship diagram; represents the vertex set in the functional relationship diagram, where each vertex corresponds to a functional component; represents the edge set in the functional relationship diagram. The edges are used to connect different vertices, reflecting the relationships existing between the functional components.
[0090] Among them, the edge weight representing the functional dependence strength in the functional relationship diagram is:
[0091]
[0092] Among them, represents the weight of the edge connecting vertex and vertex in the functional relationship diagram, that is, the weight of the edge connecting the th functional component and the th functional component; is the feature information related to the th functional component; is the feature information related to the th functional component; represents the relationship between the functional components.
[0093] Finally, the functional semantic features of the 3D model are combined with the geometric features of the 3D model for orientation prediction.
[0094] Optionally, the embodiments of the present invention can splice the geometric features and functional semantic features of the 3D model into a model comprehensive feature, and then perform orientation prediction based on the model comprehensive feature, that is:
[0095]
[0096] Among them, represents the model comprehensive feature; represents the geometric feature; represents the functional semantic feature.
[0097] By inputting the target 3D model into the orientation discrimination model, the embodiments of the present invention can extract the first geometric feature and the first functional semantic feature of the target 3D model, and perform orientation prediction based on these features to obtain the first predicted rotation parameter. Rotating the target 3D model using the first predicted rotation parameter can effectively adjust the orientation of the model, making it more consistent and accurate in visual display, thereby helping to improve the generation efficiency of the views of the 3D model.
[0098] Optionally, as Figure 3 shown, the flowchart of the training process of the orientation discrimination model provided by the embodiment of the present invention. The training process of the orientation discrimination model may include:
[0099] S300. Obtain a 3D model training set, where the 3D model training set includes a plurality of standardized 3D training models with adjusted orientations.
[0100] Specifically, in the embodiment of the present invention, based on a certain number of existing 3D models with adjusted orientations, these 3D models are standardized to obtain 3D training models. The 3D training models are used to train the orientation discrimination model. After the training is completed, the orientation discrimination model can receive any 3D model data and calculate the required rotation angle to adjust the 3D model to the correct orientation. After adjustment, manual orientation confirmation can be performed, and the confirmed sample data is incorporated into the 3D model training set, thereby gradually enriching and enhancing the training data of the orientation discrimination model. After multiple iterations, when the accuracy of the orientation discrimination model reaches a predetermined standard, the manual confirmation step can be cancelled to achieve a fully automated orientation adjustment process.
[0101] S310. Input each 3D training model in the 3D model training set into the orientation discrimination model for training, extract the second geometric feature and the second functional semantic feature of the 3D training model, and then rotate the 3D training model according to the random rotation parameter, and extract the third geometric feature and the third functional semantic feature of the rotated 3D training model.
[0102] S320. Use the second geometric feature and the third geometric feature to obtain a geometric consistency loss.
[0103] Among them, the geometric consistency loss is used to evaluate the consistency of the geometric features of the 3D model in different orientations, that is, the difference between the geometric features of the 3D model before and after rotation. In the embodiment of the present invention, the difference between the second geometric feature and the third geometric feature of the 3D model before and after rotation can be compared to calculate the geometric consistency loss. For example, the geometric consistency loss can be the mean square error or cosine similarity between the second geometric feature and the third geometric feature.
[0104] S330. Use the second functional semantic feature and the third functional semantic feature to obtain a functional semantic constraint loss.
[0105] Among them, the functional semantic constraint loss is used to evaluate the consistency of the functional semantic features of the 3D model in different orientations, that is, whether the functions and meanings of the 3D model in different perspectives remain stable. In the embodiments of the present invention, the functional semantic constraint loss can be calculated by comparing the differences between the second functional semantic features and the third functional semantic features before and after rotation. For example, the functional semantic constraint loss can be the cross-entropy loss between the second functional semantic features and the third functional semantic features.
[0106] S340. Based on the third geometric feature and the third functional semantic feature, obtain the second predicted rotation parameter, and use the second predicted rotation parameter and the random rotation parameter to obtain the orientation prediction loss.
[0107] Among them, the orientation prediction loss is used to evaluate the difference between the rotation parameter predicted by the model and the actual random rotation parameter, so as to reflect the accuracy of the orientation discrimination model in judging the object orientation. In the embodiments of the present invention, the orientation prediction rotation parameter can be calculated by comparing the differences between the second predicted rotation parameter and the random rotation parameter.
[0108] S350. Use the geometric consistency loss, the functional semantic constraint loss, and the orientation prediction loss to obtain the comprehensive loss.
[0109] Among them, the comprehensive loss is a total loss value formed by combining the geometric consistency loss, the functional semantic constraint loss, and the orientation prediction loss. The present invention can provide a comprehensive optimization objective by comprehensively considering the consistency of geometric and semantic features and the accuracy of orientation prediction, so as to better guide the parameter update of the model during the training process. For example, in the embodiments of the present invention, the weighted average method can be adopted, that is, each loss is multiplied by the corresponding weight and then added to form the comprehensive loss.
[0110] S360. Based on the comprehensive loss, adjust the network parameters of the orientation discrimination model, and retrain until the preset training end condition is reached to obtain the trained orientation discrimination model.
[0111] In the embodiments of the present invention, based on the calculated comprehensive loss, the network parameters of the orientation discrimination model can be adjusted. Through the backpropagation algorithm, the orientation discrimination model will continuously update the parameters according to the feedback information of the loss. The training process will continue until the preset training end condition is met (for example, reaching a certain loss value or the number of training rounds), so as to obtain a trained orientation discrimination model.
[0112] Optionally, the preset training end condition can be that the comprehensive loss reaches within the preset loss threshold or the number of training rounds reaches the preset round threshold.
[0113] The training process of the orientation discrimination model provided by the embodiments of the present invention is a self-supervised training process. Figure 4The following is a logic block diagram of the self-supervised training process provided by an embodiment of the present invention. In the design of the loss function, the orientation prediction loss, the geometric consistency loss, and the functional semantic constraint loss are calculated respectively and combined. The form of the final loss function can be:
[0114]
[0115] Wherein, is the comprehensive loss; is the orientation prediction loss; is the geometric consistency loss; is the functional semantic constraint loss; and are hyperparameters that can be set according to actual needs.
[0116] During the self-supervised training process, for each 3D model, a random rotation transformation is applied. The rotated model will generate new comprehensive multi-scale geometric features and functional semantic features, and the corresponding losses are calculated based on these features. At the same time, a fully connected multi-layer perceptron (MLP) module is used to input the features of the rotated model and output the predicted rotation parameters .
[0117] The orientation prediction loss is also called the rotation recovery loss. In the self-supervised training process, the rotation recovery loss is designed as a self-supervised loss, aiming to make the rotation parameters predicted by the orientation discrimination model be restored to a state matching the original unrotated 3D model. The objective is:
[0118]
[0119] The design of the orientation prediction loss function is:
[0120]
[0121] Wherein, is the norm calculation; is the identity matrix.
[0122] To ensure that the 3D model features after predicted rotation correction are consistent with the original model features (i.e., calculate the multi-scale geometric features and functional semantic features of the 3D model before and after rotation correction respectively), the designed feature consistency loss function is:
[0123]
[0124] Wherein, is the feature consistency loss; is the feature extraction network; is the feature of the rotated 3D model; is the feature of the original model.
[0125] In the embodiment of the present invention, by training a standardized 3D model training set with adjusted orientations, the orientation discrimination model is used to extract the second geometric feature and the second functional semantic feature, which can effectively obtain the geometric and semantic information of the 3D model in different orientations. Further, through random rotation processing, the third geometric feature and the third functional semantic feature after rotation are extracted, so as to calculate the geometric consistency loss and the functional semantic constraint loss, providing a solid foundation for generating efficient 3D model views. In addition, by adjusting the model parameters through the comprehensive loss, the trained orientation discrimination model has stronger robustness and generalization ability, so that the orientation can be adjusted more accurately in practical applications.
[0126] Optionally, based on one or more corresponding embodiments above, Figure 3 In another optional embodiment provided by the embodiment of the present invention, the standardization process includes:
[0127] Obtain a 3D training model; move the geometric center of the 3D training model to the target origin, and adjust the 3D training model to a preset specification size to obtain a standardized 3D training model.
[0128] In the embodiment of the present invention, by calculating the geometric center of the 3D model and moving it to the origin of the coordinate system, the position of the model in space is ensured to be unified. This symmetry helps the stability of the 3D model in subsequent processing, especially when analyzing different perspectives or poses, and helps to eliminate the influence of position deviation.
[0129] In the embodiment of the present invention, adjusting the 3D training model to a preset specification size enables different models to be compared and analyzed at the same scale. This unified scale standardization makes the subsequent training and inference processes more efficient and reduces the errors caused by the different sizes of the models.
[0130] In the embodiment of the present invention, through the standardization of position and size, a consistent reference framework is ensured among multiple 3D models in the processing flow. The standardized 3D model simplifies the model data input and reduces the complexity of model preprocessing, making the training process more efficient. In addition, the unified positioning and size adjustment of the 3D model help the machine learning algorithm to converge faster, enhancing the learning ability and prediction accuracy of the orientation discrimination model.
[0131] Optionally, based on the above Figure 1Based on one or more corresponding embodiments, in another alternative embodiment provided by the embodiments of the present invention, after performing multi-view rendering on the target 3D model according to the preset rendering parameters based on the target 3D model with the adjusted orientation and obtaining the view rendering result, the method may further include:
[0132] Use the cover image selection model to select a view from the view rendering result as the cover image of the target 3D model.
[0133] Among them, the cover image selection model can be a pre-trained visual aesthetics discrimination model, which judges its aesthetic quality by analyzing the visual features of the view.
[0134] Specifically, the embodiments of the present invention can use the cover image selection model to extract the visual features of each view in the view rendering result, such as: color, composition, texture, contrast, light and shadow, etc. Then, based on these visual features, an aesthetic score is calculated, and the view with the highest aesthetic score in the view rendering result is selected as the cover image of the target 3D model.
[0135] The training of the cover image selection model usually depends on an image dataset with annotations, which contains images scored by humans to identify their aesthetic quality. The cover image selection model uses machine learning or deep learning algorithms (such as support vector machines, decision trees, and deep neural networks, etc.) to train the extracted visual features and establish the relationship between the input visual features and the aesthetic score. After training, the cover image selection model can evaluate the views of the 3D model, output an aesthetic score, and judge the aesthetic quality of the views.
[0136] The embodiments of the present invention significantly reduce the time for manual screening and evaluation through the automated process of selecting the cover image, thereby supporting the processing of a large number of 3D models, reducing the time from view rendering to cover image selection of the 3D model, and improving the fast display ability of the 3D model. In addition, the embodiments of the present invention can select a cover image with aesthetic value to attract the user's attention to the 3D model, enhance the visual appeal of the 3D model during display, and thus obtain a better display effect.
[0137] To facilitate understanding of the overall process of the 3D model view generation method provided by the embodiments of the present invention, here in combination with Figure 5 It is described as follows: Figure 5The following is a logic block diagram of a three-dimensional model view generation method provided by an embodiment of the present invention. Through designing an automated process and leveraging intelligent algorithms, the embodiment of the present invention realizes the automatic generation of the cover image and six views of the three-dimensional model, thereby significantly reducing the need for manual operations. The operation process includes: first, performing standardization processing on the geometric center and size of the loaded three-dimensional model; then, using a trained orientation discrimination model to adjust the orientation; next, applying preset lighting and camera parameters to complete view rendering; finally, using a cover image selection model to determine the final cover image.
[0138] The orientation discrimination model is a deep learning model constructed based on a front discrimination algorithm. Initially, it can be trained by manually confirming sample data, and subsequently, the three-dimensional training model with a positive orientation can be fully automated to supplement it to further optimize the performance of the model. Based on this, the embodiment of the present invention can automatically adjust the angle and direction of the model, quickly generate the required views, support the processing of a large number of three-dimensional models, thereby improving the generation speed, shortening the rendering time, and meeting the needs of a large number of cultural relic digitizations.
[0139] In addition, the embodiment of the present invention supports flexible configuration options, can meet diverse product requirements, including different background colors, resolutions, and formats, etc. Users can easily adjust the parameters through a friendly system interface and apply them to multiple models in batches to avoid repetitive labor. The introduction of the automated process significantly reduces human errors and ensures the standardization of the output results. At the same time, the orientation discrimination model provided by the embodiment of the present invention also has good scalability, can adapt to the development of future technologies and new requirements, and supports continuous optimization and upgrading.
[0140] Although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous.
[0141] It should be understood that the various steps recorded in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0142] Corresponding to the above method embodiment, the embodiment of the present invention also provides a three-dimensional model view generation device, the structure of which is as Figure 6 shown, and may include: a three-dimensional model acquisition unit 10, an orientation adjustment unit 20, and a view rendering result acquisition unit 30.
[0143] The three-dimensional model acquisition unit 10 is used to acquire a target three-dimensional model;
[0144] An orientation adjustment unit 20, configured to perform orientation adjustment on a target 3D model by using a pre-trained orientation discrimination model, so as to obtain a target 3D model with adjusted orientation;
[0145] A view rendering result obtaining unit 30, configured to perform multi-view rendering on the target 3D model according to preset parameters based on the target 3D model with adjusted orientation, so as to obtain a view rendering result, where the view rendering result includes multiple views of the target 3D model.
[0146] Optionally, the orientation adjustment unit 20 may specifically be configured to input the target 3D model into the orientation discrimination model, so that the orientation discrimination model extracts a first geometric feature and a first functional semantic feature of the target 3D model, and performs orientation prediction based on the first geometric feature and the first functional semantic feature to obtain an output first predicted rotation parameter, and rotates the target 3D model by using the first predicted rotation parameter to obtain a target 3D model with adjusted orientation.
[0147] Optionally, the 3D model view generation device may further include: a model training unit.
[0148] The model training unit is configured to obtain a 3D model training set, where the 3D model training set includes multiple standardized 3D training models with adjusted orientation; sequentially input each 3D training model in the 3D model training set into the orientation discrimination model for training, extract a second geometric feature and a second functional semantic feature of the 3D training model, and then rotate the 3D training model according to a random rotation parameter, and extract a third geometric feature and a third functional semantic feature of the rotated 3D training model; obtain a geometric consistency loss by using the second geometric feature and the third geometric feature; obtain a functional semantic constraint loss by using the second functional semantic feature and the third functional semantic feature; obtain a second predicted rotation parameter based on the third geometric feature and the third functional semantic feature, and obtain an orientation prediction loss by using the second predicted rotation parameter and the random rotation parameter; obtain a comprehensive loss by using the geometric consistency loss, the functional semantic constraint loss, and the orientation prediction loss; adjust network parameters of the orientation discrimination model based on the comprehensive loss, and re-train until a preset training end condition is reached, so as to obtain a trained orientation discrimination model.
[0149] Optionally, the 3D model view generation device may further include: a model standardization unit.
[0150] The model standardization unit is configured to obtain a 3D training model; move the geometric center of the 3D training model to a target origin, and adjust the 3D training model to a preset specification size to obtain a standardized 3D training model.
[0151] Optionally, the geometric features include: local differential geometric features, mesoscale structure features, and global structure features.
[0152] Optionally, the three-dimensional model view generation device may further include: a cover image selection unit.
[0153] The cover image selection unit is configured to select a view from the view rendering result as the cover image of the target three-dimensional model by using a cover image selection model.
[0154] Optionally, the preset parameters include: light source parameters, camera parameters, and rendering image parameters.
[0155] A three-dimensional model view generation device provided by the present invention is configured to: obtain a target three-dimensional model; perform orientation adjustment on the target three-dimensional model by using a pre-trained orientation discrimination model to obtain a target three-dimensional model with adjusted orientation; based on the target three-dimensional model with adjusted orientation, perform multi-view rendering on the target three-dimensional model according to preset parameters to obtain a view rendering result, where the view rendering result includes multiple views of the target three-dimensional model. The present invention performs orientation adjustment on the target three-dimensional model through an automated technique and performs preset parameter rendering based on this, thereby quickly and efficiently generating multiple views, thus supporting the processing of large-scale three-dimensional models, significantly improving the generation speed and shortening the rendering time to meet the needs of a large number of cultural relic digitizations.
[0156] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0157] The three-dimensional model view generation device includes a processor and a memory. The above three-dimensional model obtaining unit 10, orientation adjustment unit 20, view rendering result obtaining unit 30, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.
[0158] The processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the orientation of the target three-dimensional model is adjusted through an automated technique and preset parameter rendering is performed based on this, thereby quickly and efficiently generating multiple views, thus supporting the processing of large-scale three-dimensional models, significantly improving the generation speed and shortening the rendering time to meet the needs of a large number of cultural relic digitizations.
[0159] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the three-dimensional model view generation method is implemented.
[0160] An embodiment of the present invention provides a processor, where the processor is configured to run a program, and when the program runs, the three-dimensional model view generation method is executed.
[0161] As shown Figure 7 in the figure, an embodiment of the present invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003. Among them, the processor 1001 and the memory 1002 communicate with each other through the bus 1003. The processor 1001 is used to call program instructions in the memory 1002 to execute the above-mentioned three-dimensional model view generation method. The electronic device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0162] The present invention also provides a computer program product, which is suitable for executing a program initialized with the steps of the three-dimensional model view generation method when executed on an electronic device.
[0163] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0164] In a typical configuration, an electronic device includes one or more processors (CPUs), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, etc.
[0165] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory includes at least one memory chip. The memory is an example of a computer-readable medium.
[0166] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0167] In the description of the present invention, it should be understood that if terms such as "upper", "lower", "front", "rear", "left", and "right" are used to indicate the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of the present invention.
[0168] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0169] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0170] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the present invention.
Claims
1. A method for generating a three-dimensional model view, characterized in that: include: Obtaining a target three-dimensional model; Using a pre-trained orientation discrimination model to adjust the orientation of the target three-dimensional model, to obtain the target three-dimensional model with the adjusted orientation; Based on the target three-dimensional model with the adjusted orientation, performing multi-view rendering on the target three-dimensional model according to preset parameters to obtain a view rendering result, wherein the view rendering result includes multiple views of the target three-dimensional model; The step of adjusting the orientation of the target three-dimensional model by using a pre-trained orientation discrimination model to obtain the target three-dimensional model with the adjusted orientation includes: The target three-dimensional model is input into the orientation discrimination model so that the orientation discrimination model extracts the first geometric feature and the first functional semantic feature of the target three-dimensional model, and performs orientation prediction based on the first geometric feature and the first functional semantic feature to obtain the output first predicted rotation parameter, and the target three-dimensional model is rotated using the first predicted rotation parameter to obtain the target three-dimensional model with adjusted orientation.
2. The method according to claim 1, characterized in that The training process of the orientation discrimination model includes: Obtaining a three-dimensional model training set, wherein the three-dimensional model training set includes a plurality of standardized three-dimensional training models whose orientations have been adjusted; Inputting each of the three-dimensional training models in the three-dimensional model training set into the orientation discrimination model in sequence for training, extracting a second geometric feature and a second functional semantic feature of the three-dimensional training model, and then rotating the three-dimensional training model according to a random rotation parameter to extract a third geometric feature and a third functional semantic feature of the rotated three-dimensional training model; Using the second geometric feature and the third geometric feature, obtaining a geometric consistency loss; Using the second functional semantic feature and the third functional semantic feature, obtaining a functional semantic constraint loss; Based on the third geometric feature and the third functional semantic feature, obtain a second predicted rotation parameter, and use the second predicted rotation parameter and the random rotation parameter to obtain an orientation prediction loss; Using the geometric consistency loss, the functional semantic constraint loss and the orientation prediction loss, a comprehensive loss is obtained; Based on the comprehensive loss, the network parameters of the orientation discrimination model are adjusted, and retraining is performed until a preset training end condition is reached to obtain the trained orientation discrimination model.
3. The method according to claim 2, characterized in that The standardized process includes: Obtaining a three-dimensional training model; The geometric center of the three-dimensional training model is moved to the target origin, and the three-dimensional training model is adjusted to a preset specification size to obtain the standardized three-dimensional training model.
4. The method according to any one of claims 2 to 3, characterized in that The geometric features include: local differential geometric features, mesoscale structural features and global structural features.
5. The method according to claim 1, characterized in that After performing multi-view rendering on the target three-dimensional model based on the adjusted orientation according to preset rendering parameters to obtain a view rendering result, the method further includes: A cover image selection model is used to select a view from the view rendering results as a cover image of the target three-dimensional model.
6. The method according to claim 1, characterized in that The preset parameters include: light source parameters, camera parameters and rendering image parameters.
7. A three-dimensional model view generation device, characterized in that: include: 3D model acquisition unit, orientation adjustment unit and view rendering result acquisition unit, The three-dimensional model obtaining unit is used to obtain a target three-dimensional model; The orientation adjustment unit is used to adjust the orientation of the target three-dimensional model using a pre-trained orientation discrimination model to obtain the target three-dimensional model with the adjusted orientation; The view rendering result obtaining unit is used to perform multi-view rendering on the target three-dimensional model according to preset parameters based on the target three-dimensional model with the adjusted orientation, to obtain a view rendering result, wherein the view rendering result includes multiple views of the target three-dimensional model; The step of adjusting the orientation of the target three-dimensional model by using a pre-trained orientation discrimination model to obtain the target three-dimensional model with the adjusted orientation includes: The target three-dimensional model is input into the orientation discrimination model so that the orientation discrimination model extracts the first geometric feature and the first functional semantic feature of the target three-dimensional model, and performs orientation prediction based on the first geometric feature and the first functional semantic feature to obtain the output first predicted rotation parameter, and the target three-dimensional model is rotated using the first predicted rotation parameter to obtain the target three-dimensional model with adjusted orientation.
8. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the three-dimensional model view generating method according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: The electronic device includes at least one processor, and at least one memory and a bus connected to the processor; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the three-dimensional model view generation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN115100339A
Model rendering method and device and electronic equipment
CN117671097A