Texture mapping matching detection method and device, equipment and storage medium

By mapping the three-dimensional model to the texture space and using a neural network model to detect the compatibility of the texture map and the texture coordinate area, the problems of low efficiency and low accuracy in the existing technology are solved, and efficient and accurate texture map matching detection is achieved.

CN120823302APending Publication Date: 2025-10-21TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410444542.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies are inefficient and consume a lot of hardware resources when detecting the matching between three-dimensional models and texture maps. In addition, manual judgment is easily affected by subjective factors, resulting in low accuracy.

Method used

By mapping the 3D model to the texture space, obtaining the texture coordinate map, and using a neural network model to predict the regional fit between the texture map and the texture coordinate region, the matching is automatically detected, eliminating the need for view rendering and manual judgment processes.

Benefits of technology

It improves the efficiency and accuracy of texture map matching detection, saves hardware resources and manpower costs, and achieves objective matching detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823302A_ABST
    Figure CN120823302A_ABST
Patent Text Reader

Abstract

The invention discloses a texture mapping matching detection method and device, equipment and a storage medium. The method comprises the steps that a three-dimensional model and a texture map of an object are acquired, and a texture image area in the texture map is used for determining the display effect of a model area; the three-dimensional model is mapped to a texture space, a texture coordinate graph is obtained, and a texture coordinate area in the texture coordinate graph is obtained by mapping a model area contained in the three-dimensional model; obtaining a region adaptation degree between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate graph; and detecting the matching between the texture map and the three-dimensional model based on the obtained region adaptation degree to obtain a matching detection result. According to the invention, processing resources can be saved, and the matching detection efficiency and accuracy of the texture mapping can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, specifically to the field of image processing technology, and in particular to a texture map matching detection method, device, equipment and storage medium. Background Art

[0002] With the development of image technology, texture mapping has been widely used. By adding a matching texture map to the surface of an object's three-dimensional model, the object can be made to look more realistic. Currently, the method for detecting whether an object's three-dimensional model matches its texture map is usually as follows: first, a three-dimensional software tool is used to render multiple views of the three-dimensional model based on the texture map to obtain a rendering image, and then the rendering image is manually observed to subjectively determine whether the three-dimensional model and the texture map match. As can be seen, in the process of determining whether the three-dimensional model and the texture map match, the existing method requires a lot of time, energy, and hardware resources to render the three-dimensional model views. This results in low efficiency in texture map matching detection and consumes a lot of hardware resources. In addition, detecting whether the texture map matches based on the visual effect of the rendering image is easily interfered by the rendering work and the subjective intention of the person. The accuracy and rules of the judgment vary from person to person. Long-term work will significantly reduce the efficiency and accuracy of texture map matching detection. Summary of the Invention

[0003] The embodiments of the present application provide a texture map matching detection method, apparatus, device, and storage medium, which can save labor costs and improve the efficiency and accuracy of texture map matching detection.

[0004] In one aspect, an embodiment of the present application provides a method for detecting the matching of a texture map, the method comprising:

[0005] Acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and each texture image region is used to determine the display effect of each model region;

[0006] Mapping the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, and a texture coordinate region is obtained by mapping a model region included in the three-dimensional model;

[0007] Obtaining a degree of regional fit between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping a model region determined by the corresponding texture image region;

[0008] Based on the obtained regional fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

[0009] On the other hand, an embodiment of the present application provides a texture map matching detection device, the device comprising:

[0010] An acquisition unit, configured to acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and a texture image region is used to determine a display effect of a model region;

[0011] a processing unit configured to map the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, where a texture coordinate region is obtained by mapping a model region included in the three-dimensional model;

[0012] The processing unit is further configured to obtain a degree of regional fit between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping a model region determined by the corresponding texture image region;

[0013] The processing unit is further configured to detect the matching between the texture map and the three-dimensional model based on the acquired regional fitness, and obtain a matching detection result.

[0014] In one embodiment, when the processing unit is used to obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0015] Performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect image; and constructing input data of a neural network model using the image superposition effect image;

[0016] The neural network model is called to predict the regional adaptation between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the input data.

[0017] In another embodiment, the image superposition effect diagram includes: a first superposition effect diagram; accordingly, when the processing unit is used to perform image superposition on the texture map and the texture coordinate map to obtain the image superposition effect diagram, it can be specifically used to:

[0018] Obtaining a mask coordinate map corresponding to the texture coordinate map, wherein the mask coordinate map includes a plurality of mask regions, and one mask region corresponds to one texture coordinate region in the texture coordinate map;

[0019] Perform image superposition on the texture map and the mask coordinate map to obtain a first superposition effect image.

[0020] In another embodiment, when the processing unit is used to obtain the mask coordinate map corresponding to the texture coordinate map, it can be specifically used to:

[0021] dilating each texture coordinate region in the texture coordinate map along its boundary; wherein the dilating process for the i-th texture coordinate region comprises: sampling a plurality of points on the boundary of the i-th texture coordinate region, determining an extension coefficient of the corresponding point according to a curvature of each point, and extending the corresponding point along a normal of the corresponding point according to the extension coefficient to dilate the i-th texture coordinate region, where i is a positive integer and is less than or equal to the number of the texture coordinate regions;

[0022] In the texture coordinate map, mask processing is performed on each expanded texture coordinate region to obtain a coordinate mask map.

[0023] In another embodiment, the image superposition effect diagram includes: a second superposition effect diagram; accordingly, when the processing unit is used to perform image superposition on the texture map and the texture coordinate map to obtain the image superposition effect diagram, it can be specifically used to:

[0024] Performing transparent processing on the texture coordinate map using a preset transparency to obtain a transparently processed texture coordinate map;

[0025] The texture coordinate image after the transparency processing and the texture map are superimposed to obtain a second superimposed effect image.

[0026] In another embodiment, any image includes N channels of image data, where N is a positive integer; accordingly, when the processing unit is used to construct input data of the neural network model using the image superposition effect diagram, it can be specifically used to:

[0027] Acquire image data of each channel in the texture map, image data of each channel in the texture coordinate map, and image data of each channel in the image overlay effect map;

[0028] In the channel axis direction, the acquired image data of M channels are stacked to obtain the input data of the neural network model; where M is an integer greater than N.

[0029] In another embodiment, the input layer of the neural network model includes M input channels, and different input channels are used to input image data of different channels in the input data.

[0030] In another embodiment, the output layer of the neural network model includes two output dimensions, and different output dimensions are used to output different prediction probability values;

[0031] Accordingly, when the processing unit is used to detect the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtain the matching detection result, it can be specifically used to:

[0032] Calling the neural network model to perform a binary classification task based on the obtained regional fitness to obtain a target classification result; the target classification result includes: a first predicted probability value that there is a match between the texture map and the texture coordinate map, and a second predicted probability value that there is no match between the texture map and the texture coordinate map;

[0033] After the neural network model outputs the target classification result through two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

[0034] In another embodiment, when the processing unit is used to obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0035] For a p-th texture image region in the texture map, calculating an intersection-over-union ratio between the p-th texture image region and a corresponding texture coordinate region in the texture coordinate map, where the value of p is a positive integer and is less than or equal to the number of texture image regions included in the texture map;

[0036] The calculated intersection-over-union ratio is used as the region adaptation degree between the p-th texture image region and the corresponding texture coordinate region.

[0037] In another embodiment, when the processing unit is used to obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0038] For a p-th texture image region in the texture map, Q points are selected as boundary points on the region boundary of the p-th texture coordinate region corresponding to the p-th texture image region, where Q is a positive integer;

[0039] Sampling R pixels in the texture map along the normal of the qth boundary point, where q∈[1,Q], R is an integer greater than 1, the R pixels are all located on the normal of the qth boundary point, and at least two of the R pixels are located on both sides of the region boundary of the pth texture coordinate region;

[0040] Determining pixel values ​​of the R pixel points based on the texture map, and calculating pixel value gradients pixel by pixel in the R pixel points according to the pixel values ​​of the R pixel points; and integrating the calculated pixel value gradients to obtain the pixel value gradient of the qth boundary point;

[0041] The degree of regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region is determined according to the pixel value gradients of the Q boundary points.

[0042] In another embodiment, when the processing unit is used to detect the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtain the matching detection result, it can be specifically used to:

[0043] Determine the regional weight of each texture image region based on the importance of the model region corresponding to each texture image region, where the regional weight is proportional to the importance;

[0044] Using the regional weight of each texture image region, weighted summing is performed on the regional fitness corresponding to the acquired corresponding texture image region to obtain the image fitness between the texture map and the texture coordinate map;

[0045] Based on the image fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

[0046] In another aspect, an embodiment of the present application provides a computer device, the computer device including an input interface and an output interface, and the computer device further including:

[0047] processors and computer storage media;

[0048] The processor is adapted to implement one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded by the processor and executed by the above-mentioned texture map matching detection method.

[0049] On the other hand, an embodiment of the present application provides a computer storage medium, which stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the above-mentioned texture map matching detection method.

[0050] On the other hand, an embodiment of the present application provides a computer program product, which includes one or more instructions; when one or more instructions in the computer program product are executed by a processor, the above-mentioned texture map matching detection method is implemented.

[0051] The embodiment of the present application can map the three-dimensional model of the object to the texture space to obtain a texture coordinate map, and then detect the matching between the texture map and the three-dimensional model (i.e., detect whether the texture map matches the three-dimensional model) based on the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map. This not only eliminates complex processing processes such as view rendering, thereby improving the efficiency of texture map matching detection and saving hardware resources; it also eliminates the process of manual participation in judgment, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0053] Figure 1a This is a schematic diagram of mapping a three-dimensional model to a texture space provided by an embodiment of the present application;

[0054] Figure 1b This is a model rendering of a three-dimensional model provided in an embodiment of the present application;

[0055] Figure 1c This is a model rendering of another three-dimensional model provided in an embodiment of the present application;

[0056] Figure 2 1 is a flow chart of a texture map matching detection method provided in an embodiment of the present application;

[0057] Figure 3a This is a schematic diagram of the principle of intersection-over-union calculation provided in an embodiment of the present application;

[0058] Figure 3b This is a schematic diagram of sampling pixel points along a perpendicular line of a boundary point of a texture coordinate region provided by an embodiment of the present application;

[0059] Figure 3c This is a schematic diagram of modifying the output layer of an image classification network provided by an embodiment of the present application;

[0060] Figure 4 1 is a flow chart of a texture map matching detection method provided by another embodiment of the present application;

[0061] Figure 5ais a schematic diagram of the normal and tangent lines of any point on a curve provided in an embodiment of the present application;

[0062] Figure 5b This is a schematic diagram of generating a mask coordinate map provided by an embodiment of the present application;

[0063] Figure 5c This is a schematic diagram of generating a first superposition effect diagram provided in an embodiment of the present application;

[0064] Figure 5d This is a schematic diagram of generating a second superposition effect diagram provided in an embodiment of the present application;

[0065] Figure 5e This is a schematic diagram of generating input data for a neural network model provided in an embodiment of the present application;

[0066] Figure 5f This is a schematic diagram of modifying the output layer and input layer of an image classification network provided by an embodiment of the present application;

[0067] Figure 5g This is a schematic diagram of the structure of a neural network model provided in an embodiment of the present application;

[0068] Figure 5h This is a schematic diagram of the structure of another neural network model provided in an embodiment of the present application;

[0069] Figure 6a is a schematic diagram of a positive example of matching a three-dimensional model with a texture map provided in an embodiment of the present application;

[0070] Figure 6b is a schematic diagram of a negative example of a mismatch between a three-dimensional model and a texture map provided by an embodiment of the present application;

[0071] Figure 6c This is a flow chart of model reasoning of a neural network model provided in an embodiment of the present application;

[0072] Figure 7 1 is a schematic structural diagram of a texture map matching detection device provided in an embodiment of the present application;

[0073] Figure 8 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0075] The embodiment of the present application proposes a method for determining whether a texture map and a three-dimensional model match based on a texture coordinate map of a three-dimensional model (also referred to as a uv coordinate map, where u represents the horizontal direction and v represents the vertical direction) (hereinafter referred to as a texture map matching detection method). This method can directly analyze the texture coordinate map and texture map of the three-dimensional model through an algorithm to detect whether the current three-dimensional model and texture map match. This not only eliminates complex processing processes such as view rendering, thereby improving detection efficiency and saving hardware resources; it also eliminates the need for manual participation in the judgment process, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection, saving labor costs and improving detection efficiency. Among them:

[0076] ① A three-dimensional model is a polygonal representation of an object. This object can be a real-world entity (such as a real house or animal) or a fictional object (such as a game character or game resource in a game scene). Specifically, a three-dimensional model of an object can include multiple model regions, each representing a part of the object (e.g., a game character's face or hand). Each model region can include multiple meshes (a type of planar graphic) drawn based on the corresponding part's shape. The meshes can be triangles, quadrilaterals, etc. Furthermore, each vertex in the three-dimensional model is associated with a texture coordinate. Two-dimensional (2D) texture coordinates are typically represented by (u, v), with u and v both taking the value of (0, 1). The texture coordinates associated with any vertex indicate the location of the texture data to be applied to that vertex in the texture map. For example, if the texture coordinates of a vertex are (0.2, 0.3), then the texture data located at (0.2, 0.3) in the texture map should be applied (i.e., mapped) to that vertex, resulting in the display effect corresponding to that texture data.

[0077] ② The texture coordinate map of a 3D model refers to the image obtained by mapping the 3D model to the texture space. The so-called texture space can be understood as a coordinate system established based on the reference point in the texture map (such as the point in the lower left corner). Specifically, the texture coordinate map may include multiple texture coordinate areas. A texture coordinate area is obtained by mapping a model area contained in the 3D model, such as Figure 1a As shown. It can be seen that the process of mapping a 3D model to a texture space can also be called UV unfolding (i.e., the process of unfolding the 3D model into a 2D plane in the texture space based on the texture coordinates of the vertices of each mesh in the 3D model).

[0078] ③The texture map is an image used to determine the display effect of an object. By applying (mapping) the texture map to the surface of the three-dimensional model of the object in a specific way, the object can be made to look more realistic. Specifically, the texture map may include multiple texture image areas, each of which can be a two-dimensional graphic. One texture image area corresponds to a model area contained in the three-dimensional model of the object, that is, one texture image area is applied (mapped) to a model area to determine the display effect of the part of the object represented by the corresponding model area. In actual applications, there may be a situation where the three-dimensional model and the texture map match, or there may be a situation where the three-dimensional model and the texture map do not match: when the three-dimensional model matches the texture map, the model rendering obtained by applying the texture map to the surface of the three-dimensional model conforms to human visual habits and has a high aesthetic score, such as Figure 1b When the 3D model and the texture map do not match, the model rendering obtained by applying the texture map to the surface of the 3D model will have problems such as strange visual effects and inconsistent with common sense in the real world. Figure 1c As shown in the model rendering using the 12 logo, Figure 1c The texture marked with 13 is an incorrect texture.

[0079] In a specific implementation, the method for determining whether a texture map and a three-dimensional model match based on a texture coordinate map of a three-dimensional model (i.e., a texture map matching detection method) proposed in an embodiment of the present application can be executed by a computer device, where the computer device can be a terminal or a server; or, the method can be jointly executed by a terminal and a server, without limitation. Among them, the terminal can be a smart phone, a computer (such as a tablet computer, a laptop computer, a desktop computer, etc.), a smart wearable device (such as a smart watch, smart glasses), an intelligent voice interaction device, a smart home appliance (such as a smart TV), a vehicle-mounted terminal or an aircraft, etc.; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, etc. Furthermore, the terminal and server can be located inside or outside the blockchain network, without limitation; further, the terminal and server can also upload any data stored internally to the blockchain network for storage to prevent the internally stored data from being tampered with, thereby improving data security.

[0080] The following takes computer equipment as the execution subject as an example. Figure 2The method flow chart shown in FIG. 1 illustrates the specific implementation process of the texture map matching detection method proposed in the embodiment of the present application. Figure 2 As shown, the method may include the following steps S201-S204:

[0081] S201, obtaining a three-dimensional model and texture map of an object.

[0082] The three-dimensional model obtained in step S201 includes multiple model regions, each of which includes multiple meshes, and each mesh vertex is associated with a texture coordinate. Furthermore, the texture map obtained in step S201 can be any texture map to be matched; the texture map can include multiple texture image regions, with each texture image region being applied to (mapped to) each model region.

[0083] S202, mapping the three-dimensional model to a texture space to obtain a texture coordinate map.

[0084] The texture coordinate map includes multiple texture coordinate regions, each of which is obtained by mapping a model region contained in a three-dimensional model. For any model region in the three-dimensional model, the corresponding vertices can be mapped to the texture space based on the texture coordinates of the vertices of each mesh in the model region, obtaining the mapping point corresponding to each vertex. The position of the mapping point corresponding to any vertex is the position indicated by the texture coordinate of the corresponding vertex; then, the mapping points corresponding to each vertex can be connected to obtain a texture coordinate region. It can be seen that the texture coordinate region obtained in this way is a connected region.

[0085] S203: Obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map.

[0086] The texture coordinate region corresponding to a texture image region is the texture coordinate region obtained by mapping the model region to which the texture image region is applied. For example, if texture image region a is applied to model region a contained in a 3D model, then the texture coordinate region a obtained by mapping model region a contained in the 3D model to the texture space is the texture coordinate region corresponding to texture image region a.

[0087] In a specific implementation, the computer device can obtain the regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map by calculating the intersection-over-union ratio. Figure 3a The area of ​​the area marked with 31 in the figure), and the area of ​​the two areas combined (such as Figure 3aIn this specific implementation, step S203 can be implemented as follows: for the pth texture image region in the texture map, where p is a positive integer and less than or equal to the number of texture image regions contained in the texture map, the intersection-and-union ratio (IOR) between the pth texture image region and the corresponding texture coordinate region in the texture coordinate map is calculated; and the calculated IOR is used as the regional fit between the pth texture image region and the corresponding texture coordinate region. For example, if the area of ​​the region where the pth texture image region intersects with its corresponding texture coordinate region is 20, and the area of ​​the region formed by the combination of the pth texture image region and its corresponding texture coordinate region is 25, then the IOR between the pth texture image region and its corresponding texture coordinate region is 20 / 25=0.8. Therefore, the regional fit between the pth texture image region and the corresponding texture coordinate region can be determined to be 0.8. Since the calculation process of the IOR is relatively simple, the efficiency of obtaining the regional fit between the texture image region and the texture coordinate region can be improved by calculating the IOR.

[0088] In another specific implementation, considering that the various texture coordinate areas in the texture coordinate map are separated from each other, when each texture image area in the texture map is adapted to the corresponding texture coordinate area in the texture coordinate map, the various texture image areas in the texture map should also be separated from each other. Since there is a blank between any texture image area and other texture image areas in the texture map in this case, there will usually be a sudden change in pixel values ​​on both sides of the area boundary of the texture coordinate area corresponding to any texture image area. Based on this, the computer device can also obtain the area adaptation between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map by calculating the pixel value gradient of at least one point on the area boundary of the corresponding texture coordinate area based on each texture image area, thereby improving the accuracy of the area adaptation.

[0089] In this specific implementation, the specific implementation method of step S203 can be: for the p-th texture image area in the texture map, select Q points as boundary points (Q is a positive integer) on the area boundary of the p-th texture coordinate area corresponding to the p-th texture image area, and sample R pixel points in the texture map along the normal (perpendicular line) of the q-th boundary point, q∈[1,Q], R is an integer greater than 1, the R pixel points are all located on the normal of the q-th boundary point, and at least two pixel points among the R pixel points are located on both sides of the area boundary of the p-th texture coordinate area; for example, a circle is used to represent the q-th boundary point, and a triangle is used to represent the sampled R pixel points, so the schematic diagram of the R pixel points can be seen in Figure 3bAs shown. After sampling R pixels, the pixel values ​​of the R pixels can be determined based on the texture map. Specifically, based on the position of the r-th pixel (r∈[1, R]) in the texture coordinate map (assuming it is (3, 2)), the pixel value of the pixel at the same position (i.e., the position (3, 2)) can be obtained from the texture map as the pixel value of the r-th pixel. Furthermore, based on the pixel values ​​of the R pixels, the pixel value gradient can be calculated pixel by pixel in the R pixels. The pixel value gradient of the currently calculated pixel is used to reflect the pixel value change amplitude between the corresponding pixel and the previous pixel. The calculated pixel value gradients are integrated (such as the mean) to obtain the pixel value gradient of the q-th boundary point. Based on this principle, after obtaining the pixel value change gradients of the Q boundary points, the regional adaptability between the p-th texture image region and the corresponding p-th texture coordinate region can be determined based on the pixel value gradients of the Q boundary points.

[0090] The embodiment of the present application does not limit the specific implementation method for determining the regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region based on the pixel value gradients of the Q boundary points. For example, the pixel value gradients of the Q boundary points can be compared with the gradient threshold respectively. If the pixel value gradients of at least a preset number of boundary points among the Q boundary points are greater than or equal to the gradient threshold (i.e., the pixel value gradients are sufficiently large), then the regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region can be set to a first preset adaptation; otherwise, the regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region can be set to a second preset adaptation, which is less than the first preset adaptation. For another example, the pixel value gradients of the Q boundary points can be integrated (such as the mean or sum), and thus the regional fitness between the p-th texture image region and the corresponding p-th texture coordinate region can be determined based on the integrated pixel value gradient. The determined regional fitness can be equal to the integrated pixel value gradient, or can be proportional to the integrated pixel value gradient; alternatively, the regional fitness between the p-th texture image region and the corresponding p-th texture coordinate region can be determined based on the size relationship between the integrated pixel value gradient and the gradient threshold. If the integrated pixel value gradient is greater than the gradient threshold, the determined regional fitness is the first preset fitness; if the integrated pixel value gradient is less than the gradient threshold, the determined regional fitness is the second preset fitness.

[0091] In another specific implementation, a computer device may utilize AI (Artificial Intelligence) technology and a neural network model to obtain the regional fit between each texture image region in a texture map and the corresponding texture coordinate region in a texture coordinate map. Based on the efficient computation and numerical prediction capabilities of the neural network model, the regional fit can be intelligently obtained, improving the efficiency and convenience of obtaining the regional fit. AI technology refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology in computer science; it primarily aims to understand the essence of intelligence and produce new intelligent machines that can react in a manner similar to human intelligence, enabling intelligent machines to possess multiple functions such as perception, reasoning, and decision-making. Accordingly, AI technology is a comprehensive discipline that may include, but is not limited to, machine learning (ML) and deep learning. Machine learning is the core of AI and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Specifically, machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Deep learning, on the other hand, is a machine learning technique that utilizes deep neural network systems. Both machine learning and deep learning typically include artificial neural networks, large language models, and supervised learning techniques.

[0092] In this specific implementation, the specific implementation method of step S203 can be: calling the neural network model to predict the regional adaptation between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map based on the texture map and the texture coordinate map; in this case, the input of the neural network model is the texture map and the texture coordinate map. Alternatively, considering that the positions of the texture image areas and the texture coordinate areas with corresponding relationships in the texture map and the texture coordinate map are the same or similar in texture space, if the texture map and the texture coordinate map are superimposed, each texture image area in the texture map can be superimposed and displayed on the texture coordinate area corresponding to it in the texture coordinate map. This can more intuitively reflect whether each texture image area and its corresponding texture coordinate area are adapted, thereby being more conducive to the prediction of regional adaptation by the neural network model and improving the accuracy of the predicted regional adaptation. Based on this, the specific implementation of step S203 may include: performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect map; using the image superposition effect map to construct input data of the neural network model, thereby calling the neural network model to predict the regional adaptability between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map based on the input data; in this case, the input of the neural network model is the input data mentioned here.

[0093] In one embodiment, the neural network model mentioned above may be a model used only to predict regional fitness; in this case, the output of the neural network model is the predicted fitness of each region. In this case, the training method of the neural network model may be: obtaining a sample texture map, a sample texture coordinate map, and fitness label data, wherein the fitness label data includes: the regional fitness between each texture image region in the sample texture map and the corresponding texture coordinate region in the sample texture coordinate map; constructing sample data using the sample texture map and the sample texture coordinate map, wherein the sample data includes the sample texture map and the sample texture coordinate map, or includes an image obtained by superimposing the sample texture map and the sample texture coordinate map; calling the neural network model to predict the regional fitness between each texture image region in the sample texture map and the corresponding texture coordinate region in the sample texture coordinate map based on the sample data, thereby obtaining a fitness prediction result; optimizing the model parameters of the neural network model based on the difference between the fitness prediction result and the fitness label data, specifically, calculating the model loss value of the neural network model based on the difference between the fitness prediction result and the fitness label data, thereby optimizing the model parameters of the neural network model in the direction of reducing the model loss value.

[0094] In another embodiment, the neural network model mentioned above can be a model for predicting regional fitness and performing a binary classification task based on the regional fitness; the binary classification task here refers to: a task of predicting the probability value of a match between any texture map and the corresponding texture coordinate map and the probability value of a non-match. In this case, the input of the neural network model is the same as the input mentioned above, and the output of the neural network model is the classification result obtained by performing the binary classification task (i.e., the predicted probability value of a match between any texture map and the corresponding texture coordinate map, and the predicted probability value of a non-match between any texture map and the corresponding texture coordinate map). Then in this case, the training method of the neural network model can be: obtain a sample texture map, a sample texture coordinate map and an annotation label, the annotation label is used to indicate whether the sample texture map matches the sample texture coordinate map, which can specifically include an annotation probability value of matching between the sample texture map and the sample texture coordinate map. The so-called annotation probability value refers to the numerical value of the annotation, which can be 0 or 1. When the value is 0, it indicates that the sample texture map and the sample texture coordinate map do not match. When the value is 1, it indicates that the sample texture map and the sample texture coordinate map match; use the sample texture map and the sample texture coordinate map to construct sample data, and call the neural network model to perform a matching on each texture image area in the sample texture map according to the sample data. The regional fitness between the domain and the corresponding texture coordinate region in the sample texture coordinate map is predicted, and a two-class classification task is performed based on the fitness prediction result (including the predicted fitness of each region) to obtain a sample classification result, which includes: a predicted probability value of a match between the sample texture map and the sample texture coordinate map, and a predicted probability value of a no match between the sample texture map and the sample texture coordinate map; based on the difference between the annotation label and the sample classification result, the model parameters of the neural network model are optimized, specifically, based on the difference between the annotation label and the sample classification result, the model loss value of the neural network model can be calculated, thereby optimizing the model parameters of the neural network model in the direction of reducing the model loss value.

[0095] Among them, when the computer device calculates the model loss value of the neural network model based on the difference between the annotated label and the sample classification result, it can directly call the binary cross entropy loss function or other loss function, perform the loss value operation based on the difference between the annotated label and the sample classification result, and obtain the model loss value of the neural network model. Alternatively, in order to better guide the neural network model to perform the binary classification task by predicting the regional fitness, the computer device can also obtain fitness label data and control the neural network model to output the fitness prediction result, so that when calculating the model loss value based on the difference between the annotated label and the sample classification result, the model loss value can be further calculated by referring to the difference between the fitness label data and the fitness prediction result. That is to say, the computer device can calculate a classification loss value based on the difference between the labeled label and the sample classification result, and the classification loss value can be calculated by calling the binary cross entropy loss function or other loss functions; and, based on the difference between the fitness label data and the fitness prediction result, calculate the fitness prediction loss value, and the fitness prediction loss value can be calculated by calling the cross entropy loss function, or the L2 norm loss function (also known as the least square error loss function) and other loss functions; then, the classification loss value and the fitness prediction loss value can be integrated to obtain the model loss value of the neural network model, thereby optimizing the model parameters of the neural network model in the direction of reducing the model loss value.

[0096] Optionally, after obtaining the model loss value in the above manner, the computer device can directly optimize the model parameters of the neural network model in the direction of reducing the model loss value, or obtain the sample weight corresponding to the sample type of the sample data based on the correspondence between the preset sample type and the sample weight, and use the obtained sample weight to weight the model loss value to update the model loss value, and after the model loss value is updated, trigger the execution of the step of optimizing the model parameters of the neural network model in the direction of reducing the model loss value; wherein, when the sample texture map in the sample data matches the sample texture coordinate map, the sample type of the sample data is a positive sample type, otherwise it is a negative sample type, and the correspondence between the sample type and the sample weight mentioned here can be set based on actual needs. For example, if the actual needs indicate that you want to improve the ability of the neural network model to recognize positive samples, then the sample weight corresponding to the positive sample type can be set to be greater than the sample weight corresponding to the negative sample type, and so on.

[0097] Based on the above description, it should be noted that: ① The direction of reducing the model loss value mentioned above refers to the model optimization direction with the goal of minimizing the model loss value; by optimizing the model in this direction, the model loss value generated by the neural network model after each optimization is smaller than the model loss value generated by the neural network model before optimization. For example, if the model loss value calculated this time is 0.85, then after optimizing the model parameters of the neural network model in the direction of reducing the model loss value, the model loss value generated by the neural network model after this optimization should be smaller than 0.85. ② The neural network model mentioned above can be any type of model, such as a convolution-based deep learning model, a multimodal language model, etc. The so-called multimodal language model refers to a pre-trained model that establishes feature representations of two or more data modalities. It can identify the number of texture image regions contained in the input texture map according to the instructions of the task, and thus predict and output the corresponding number of regional fitness. The so-called pre-trained model, also known as the cornerstone model or large model, refers to a deep neural network (DNN) with a large number of parameters. Training it on massive amounts of unlabeled data can leverage the function approximation capabilities of the large-parameter DNN to enable the pre-trained model to extract common features from the data. Moreover, after fine-tuning (fine-tuning), supervised fine-tuning (Supervised Fine-Tuning), parameter-efficient fine-tuning (PEFT), prompt-tuning (a new fine-tuning paradigm based on pre-trained language models), the pre-trained model can be applied to downstream tasks. Therefore, the pre-trained model can achieve ideal results in few-shot learning (Few-shot) or zero-shot learning (Zero-shot) scenarios. Alternatively, when the neural network model mentioned above is a model for predicting regional fitness and performing a binary classification task based on the regional fitness, the neural network model can be obtained by network transformation of the image classification network; specifically, the output layer of the image classification network may include an output channel, and the output channel may include S output dimensions, where S is an integer greater than 2, and the output layer of the neural network model may be obtained by modifying the number of output dimensions in the output layer of the image classification network; Figure 3c As shown, the modified output layer (i.e., the output layer of the neural network model) may include two output dimensions. In this case, different output dimensions in the output layer of the neural network model are used to output different probability values.

[0098] S204 , based on the obtained regional fitness, detecting the matching between the texture map and the three-dimensional model to obtain a matching detection result.

[0099] In a specific implementation, if the computer device obtains the regional fitness between each texture image region and the corresponding texture coordinate region by calling a neural network model, and the neural network model is a model for predicting the regional fitness and performing a binary classification task based on the regional fitness, since the output layer of the neural network model in this case is obtained by modifying the number of output dimensions in the output layer of the image classification network, different output dimensions are used to output different prediction probability values; then, the specific implementation method of step S204 may include the following steps: calling the neural network model to perform the binary classification task based on the obtained regional fitness to obtain a target classification result, and the target classification result includes: a first prediction probability value that there is a match between the texture map and the texture coordinate map, and a second prediction probability value that there is no match between the texture map and the texture coordinate map; after the neural network model outputs the target classification result through the two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

[0100] It can be seen that the embodiment of the present application obtains a neural network model by structurally transforming the image classification network. After calling the neural network model to obtain the regional fitness between each texture image area and the corresponding texture coordinate area, the binary classification task can be directly performed based on the obtained regional fitness, so as to detect the matching between the texture map and the three-dimensional model, thereby realizing the conversion of the matching detection task into a binary classification task of the neural network model in the operation mode. This not only simplifies the detection process, saves processing resources and improves the matching detection efficiency, but also makes use of the powerful computing power and learning ability of the neural network model to make the final matching detection result have higher accuracy.

[0101] In another specific implementation, if the computer device obtains the regional fitness between each texture image region and the corresponding texture coordinate region by calculating the intersection-over-union ratio, or the computer device obtains the regional fitness between each texture image region and the corresponding texture coordinate region by calculating the pixel value gradient, or the computer device obtains the regional fitness between each texture image region and the corresponding texture coordinate region by calling a neural network model, and the neural network model is a model only for predicting regional fitness; then, the specific implementation of step S204 can be any of the following:

[0102] Implementation method one: Compare the obtained fitness of each region with the first fitness threshold respectively, where the first fitness threshold can be set according to actual needs; if the obtained fitness of each region is greater than or equal to the first fitness threshold, it can be considered that each texture image region in the texture map is matched with the corresponding model region, so it can be determined that the matching between the texture map and the three-dimensional model is detected, thereby generating a matching detection result indicating that there is a matching between the texture map and the three-dimensional model; if there is at least one region whose fitness is less than the first fitness threshold, it can be considered that there is at least one texture image region in the texture map that is not matched with the corresponding model region, so it can be determined that the matching between the texture map and the three-dimensional model is not detected, thereby generating a matching detection result indicating that there is no matching between the texture map and the three-dimensional model.

[0103] It can be seen that through this implementation, a single texture image area and a single model area can be used as detection granularity to detect the matching between the texture map and the three-dimensional model, which can help improve the accuracy of matching detection.

[0104] Implementation method 2: Determine the regional weight of each texture image area based on the importance of the model area corresponding to each texture image area; wherein, the importance of each model area can be set according to business needs or experience values. The embodiment of the present application does not limit the specific method of determining the regional weight of each texture image area based on the importance of each model area, as long as the regional weight is proportional to the importance, that is, the greater the importance of the model area, the greater the regional weight of the texture image area corresponding to the model area. After determining the regional weight of each texture image area, the regional weight of each texture image area can be used to perform a weighted summation on the regional fitness corresponding to the corresponding texture image area to obtain the image fitness between the texture map and the texture coordinate map, and based on the image fitness, detect the matching between the texture map and the three-dimensional model to obtain a matching detection result. Specifically, the image fitness is compared with a second fitness threshold, where the second fitness threshold can be set according to actual needs; if the image fitness is greater than or equal to the second fitness threshold, it can be determined that the matching between the texture map and the three-dimensional model is detected, thereby generating a matching detection result indicating that there is a matching between the texture map and the three-dimensional model; if the image fitness is less than the second fitness threshold, it can be determined that the matching between the texture map and the three-dimensional model is not detected, thereby generating a matching detection result indicating that there is no matching between the texture map and the three-dimensional model.

[0105] For example, assuming that the texture map includes three texture image areas, and the importance of the corresponding model areas is 0.2, 0.3 and 0.5 respectively, the importance of each model area can be directly used as the area weight of the corresponding texture image area, that is, the area weights of the three texture image areas are 0.2, 0.3 and 0.5 respectively; if the area fitness corresponding to the three texture image areas is 0.8, 0.7 and 0.4 respectively, the image fitness between the texture map and the texture coordinate map can be calculated to be 0.2×0.8+0.3×0.7+0.5×0.4=0.57. Since 0.57 is less than the second fitness threshold (set as 0.8), it can be determined that there is no match between the texture map and the three-dimensional model. For another example, suppose the texture map includes three texture image regions, and the importance of the corresponding model regions is 3, 3 and 4, respectively. The importance of the three model regions can be normalized, and the three normalized importances are: 3 / (3+3+4)=0.3, 3 / (3+3+4)=0.3, 4 / (3+3+4)=0.4, respectively; thus, the normalized importance of each model region is used as the regional weight of the corresponding texture image region, that is, the regional weights of the three texture image regions are 0.3, 0.3 and 0.4, respectively; if the regional fitness corresponding to the three texture image regions is 0.8, 0.7 and 0.9, respectively, then the image fitness between the texture map and the texture coordinate map can be calculated to be 0.3×0.8+0.3×0.7+0.4×0.9=0.81. Since 0.81 is greater than the second fitness threshold (set as 0.8), it can be determined that the texture map and the three-dimensional model are matched.

[0106] It can be seen that through this implementation, when detecting the matching between the texture map and the three-dimensional model, it is possible to focus on the adaptability between the more important texture image area in the texture map and the corresponding model area, so that the final matching detection result is more accurate.

[0107] The embodiment of the present application can map the three-dimensional model of the object to the texture space to obtain a texture coordinate map, and then detect the matching between the texture map and the three-dimensional model (i.e., detect whether the texture map matches the three-dimensional model) based on the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map. This not only eliminates complex processing processes such as view rendering, thereby improving the efficiency of texture map matching detection and saving hardware resources; it also eliminates the process of manual participation in judgment, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection.

[0108] Based on the above Figure 2The embodiment of the method shown in the figure further proposes a texture map matching detection method; in the embodiment of the present application, the computer device is still used as an example to illustrate the method. Figure 4 As shown, the method may include the following steps S401-S407:

[0109] S401, obtaining a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and one texture image region is applied to one model region.

[0110] S402: Map the three-dimensional model to a texture space to obtain a texture coordinate map, wherein the texture coordinate map includes a plurality of texture coordinate regions, and a texture coordinate region is obtained by mapping a model region included in the three-dimensional model.

[0111] S403: performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect map.

[0112] In one embodiment, the image overlay effect diagram may include a first overlay effect diagram, which may also be referred to as a combined effect diagram, and is an image obtained by performing image overlay on the mask coordinate diagram and the texture map corresponding to the texture coordinate diagram. That is, when the computer device executes step S403, it may obtain the mask coordinate diagram corresponding to the texture coordinate diagram, and perform image overlay on the texture map and the mask coordinate diagram to obtain the first overlay effect diagram. The mask coordinate diagram mentioned here may include multiple mask areas, one mask area corresponds to one texture coordinate area in the texture coordinate diagram, the pixel value of each pixel point located in the mask area is a first value (such as a value of 1), and the pixel value of each pixel point located outside the mask area is a second value (such as a value of 0). Since each texture coordinate area in the texture coordinate map contains more grid lines, and each mask area in the mask coordinate map corresponding to the texture coordinate map does not contain these grid lines, by obtaining the mask coordinate map corresponding to the texture coordinate map, the mask coordinate map and the texture map are superimposed, so that the final first superposition effect map does not contain the grid lines contained in the texture coordinate map. In this way, when the subsequent neural network model predicts the regional adaptability based on the first superposition effect map, the interference of the internal grid lines can be removed, and the neural network model can be guided to focus on the overall regional direction and adaptability, thereby improving the prediction efficiency of the regional adaptability.

[0113] Specifically, when obtaining a mask coordinate map corresponding to a texture coordinate map, the computer device may directly perform mask processing on the texture coordinate map to obtain the mask coordinate map, i.e., the pixel values ​​of the pixels in each texture coordinate region in the texture coordinate map are set to a first value, and the pixel values ​​of the pixels outside each texture coordinate region are set to a second value, thereby obtaining the mask coordinate map. Alternatively, considering that in actual applications, in order to improve the mapping effect, the size of each texture image region in the texture map may be set to be larger than the size of the corresponding model region so that each texture image region can completely cover the corresponding model region, based on this, to avoid the influence of this factor on subsequent matching detection, when obtaining the mask coordinate map corresponding to the texture coordinate map, the computer device may perform dilation processing on each texture coordinate region in the texture coordinate map along the region boundary of each texture coordinate region in the texture coordinate map; specifically, the method for dilating the i-th texture coordinate region includes: sampling multiple points on the region boundary of the i-th texture coordinate region, determining an extension coefficient for each point based on the curvature of each point, and extending the corresponding point along the normal of the corresponding point based on the extension coefficient to dilate the i-th texture coordinate region. Wherein, the value of i is a positive integer and is less than or equal to the number of the texture coordinate regions, and the extension coefficient is proportional to the curvature, that is, the greater the curvature, the greater the extension coefficient; and see Figure 5a As shown, the normal of any point refers to the line 52 perpendicular to the tangent line 51 of the point. It can be seen that this method can realize adaptive expansion of each texture coordinate region and improve the expansion effect. After each texture coordinate region is expanded, each expanded texture coordinate region can be masked in the texture coordinate map to obtain a coordinate mask map, that is, the pixel values ​​of the pixel points of each expanded texture coordinate region in the texture coordinate map are set to the first value, and the pixel values ​​of the pixel points outside each expanded texture coordinate region are set to the second value to obtain a mask coordinate map, as shown in FIG. Figure 5b shown.

[0114] In addition, the computer device may perform image superposition on the texture map and the mask coordinate map through Boolean operations to obtain a first superposition effect map; the so-called Boolean operation may also be called a Boolean operation, which may include multiple operations such as merging and intersection. Among them, the pixel value of the pixel point at the fth position (f is a positive integer) in the first superposition effect map is determined based on the pixel value of the pixel point at the fth position in the mask coordinate map and the pixel value of the pixel point at the fth position in the texture map; if the pixel value of the pixel point at the fth position in the mask coordinate map is a first value (such as value 1), then the pixel value of the pixel point at the fth position in the first superposition effect map is equal to the pixel value of the pixel point at the fth position in the texture map; if the pixel value of the pixel point at the fth position in the mask coordinate map is a second value (such as value 0), then the pixel value of the pixel point at the fth position in the first superposition effect map is equal to the second value (such as value 0). Based on this, the schematic diagram of generating the first superposition effect map can be seen. Figure 5c shown.

[0115] In another embodiment, the image overlay effect diagram may include a second overlay effect diagram, which may also be called a transparent overlay effect diagram, and is an image obtained by adding a certain degree of transparency to the texture coordinate diagram and then overlaying it with the texture map. That is, when the computer device executes step S403, it may use a preset transparency to perform transparent processing on the texture coordinate diagram to obtain a transparently processed texture coordinate diagram; and, perform image overlay on the transparently processed texture coordinate diagram and the texture map to obtain a second overlay effect diagram. Among them, the f-th position in the second overlay effect diagram can simultaneously display the image content at the f-th position in the transparently processed texture coordinate diagram and the image content at the f-th position in the texture map. Based on this, the schematic diagram for generating the second overlay effect diagram can be seen in Figure 5d shown.

[0116] It should be noted that in actual applications, either of the two aforementioned embodiments can be selected for implementation, or both can be implemented simultaneously. When both embodiments are implemented simultaneously, the image overlay effect graph can include both the first overlay effect graph and the second overlay effect graph. Furthermore, in other embodiments, the computer device can directly overlay the texture map and the texture coordinate map to obtain the image overlay effect graph.

[0117] S404, constructing input data of the neural network model by using the image overlay effect diagram.

[0118] In one embodiment, the computer device can directly use the image superposition effect diagram as the input data of the neural network model; that is, in this case, if the image superposition effect diagram includes the first superposition effect diagram, the input data is the first superposition effect diagram; if the image superposition effect diagram includes the second superposition effect diagram, the input data is the second superposition effect diagram; if the image superposition effect diagram includes both the first superposition effect diagram and the second superposition effect diagram, the input data includes the first superposition effect diagram and the second superposition effect diagram.

[0119] In another embodiment, the computer device may use multiple images such as an image overlay effect map, a texture map, and a texture coordinate map to construct the input data of the neural network model, so that the input data can contain enough pixel domain visual information, so that the neural network model can learn more information from the input data, and then more accurately predict the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map, thereby improving the prediction accuracy of the regional fit. Specifically, considering that any image includes N channels of image data, N is a positive integer, for example, N = 3; therefore, the computer device can obtain the image data of each channel in the texture map, the image data of each channel in the texture coordinate map, and the image data of each channel in the image overlay effect map; and stack the obtained M channels of image data in the channel axis direction to obtain the input data of the neural network model, that is, the input data includes M channels of image data. Exemplarily, the computer device can also implement the stacking of M channels of image data based on Pytorch (a deep learning training framework).

[0120] Among them, the M mentioned above is an integer greater than N. It is understandable that if the image overlay effect diagram includes the first overlay effect diagram or the second overlay effect diagram, the number of images used to construct the input data is 3, then M=3N; if the image overlay effect diagram includes the first overlay effect diagram and the second overlay effect diagram, the number of images used to construct the input data is 4, then M=4N. Taking any image including 3 channels of image data, the size of each image data is 224×224, and the image overlay effect diagram includes the first overlay effect diagram as an example, the image data of 3 channels in the texture map, the image data of 3 channels in the texture coordinate map, and the image data of 3 channels in the first overlay effect diagram are stacked to obtain a schematic diagram of the input data (including 9 224×224 image data) can be seen. Figure 5e shown.

[0121] It should be noted that the neural network model mentioned above can be a model for predicting regional fitness and performing a binary classification task based on the regional fitness, and the output layer of the neural network model is obtained by modifying the number of output dimensions in the output layer of the image classification network. The output layer of the neural network model may include two output dimensions. Furthermore, considering that the input layer of the image classification network usually includes N input channels, it usually accepts N-channel (such as three-channel / single-channel) input. If the input data of the neural network model is constructed using multiple images such as image overlay effect images, texture maps, and texture coordinate maps, then since the neural network model utilizes the visual information of multiple images in the process of matching detection, the visual information here refers to the degree of adaptation and matching of multiple images in the pixel domain. Therefore, it is impossible to directly use the image classification network that accepts N-channel (such as three-channel / single-channel) input. In this case, the input layer of the image classification network can be structurally modified, so that the modified image classification network can be used as the neural network model. It can be seen that the neural network model mentioned in the embodiment of the present application can be obtained by structurally modifying the image classification network; specifically, the input layer of the neural network model can be obtained by adding K input channels to the input layer of the image classification network, K=MN, that is, the input layer of the neural network model in this case includes M input channels, such as Figure 5f As shown in FIG. 1 , different input channels in the input layer of the neural network model are used to input image data of different channels in the aforementioned input data. It can be seen that the embodiment of the present application can adapt to the requirements of different input image data by modifying the number of input channels in the input layer of the image classification network. It is understandable that this is merely an example of how the input layer of the image classification network can be modified by adding input channels, but it is not intended to be limiting.

[0122] It should also be noted that the embodiments of the present application do not limit the specific implementation of the image classification network. For example, it can be a ResNet-50 classification network (ResNet represents a residual neural network). In this case, the neural network model mentioned in the embodiments of the present application is obtained by structurally modifying the ResNet-50 classification network, that is, the neural network model is the modified ResNet-50 classification network. Among them, the ResNet-50 classification network is a network composed of an input layer, multiple convolutional layers, multiple pooling layers, and a fully connected layer (output layer). Among them, any convolution layer can be any of the following: a convolution layer composed of 64 7×7 convolution kernels (denoted as 7×7conv,64), a convolution layer composed of 64 1×1 convolution kernels (denoted as 1×1conv,64), a convolution layer composed of 64 3×3 convolution kernels (denoted as 3×3conv,64), a convolution layer composed of 256 1×1 convolution kernels (denoted as 1×1conv,256), a convolution layer composed of 128 1×1 convolution kernels (denoted as 1×1conv,128), a convolution layer composed of 512 1×1 convolution kernels (denoted as 1×1conv,512), a convolution layer composed of 128 3×3 convolution kernels (denoted as 3×3conv,128), and so on; any pooling layer can be any of the following: a 3×3 maximum pooling layer (i.e., 3×3maxpool) and an average pooling layer (i.e., avg pool). Based on this, we take the example of obtaining a neural network model by changing the number of input channels in the input layer of the ResNet-50 classification network to 9 and the number of output dimensions in the output layer (expressed in fc) to 2. That is, the input layer of the neural network model includes 9 input channels and the output layer includes 2 output dimensions. The network structure of the neural network model can be exemplified by referring to Figure 5g As shown. It is understandable that Figure 5g The network structure of the neural network model is only exemplified and is not limited to this. For example, when constructing a neural network model, the main structure of the ResNet network can be changed, and the network depth, width, etc. can be expanded, or other types of networks can be introduced to achieve the same or similar network output. Exemplarily, other types of networks include but are not limited to: classification networks, regression networks, networks based on the Transformer architecture (a deep learning model architecture based on the attention mechanism), and so on.

[0123] Optionally, in other embodiments, when the input data includes image data of M channels, the neural network model can also be composed of a target convolution layer and an image classification network that accepts N channel input. In this case, the output layer of the image classification network is modified to include two output dimensions, while the input layer is not modified, such as Figure 5hAs shown. The target convolution layer mentioned here refers to a convolution layer for converting the number of channels of the input data, which can make the converted input data include image data of N channels. The embodiment of the present application does not limit the structure of the target convolution layer. For example, it can be a convolution layer containing a 1×1 convolution kernel. It can be seen that in this case, the neural network model can convert the number of channels of the input data through the target convolution layer, and then input the converted input data into the input layer of the image classification network contained in the neural network model for a series of subsequent processing. In this way, while supporting the neural network model to use the visual information of multiple images for matching detection, it can avoid structural modification of the input layer of the image classification network, thereby saving the time cost consumed by structural modification.

[0124] S405 , calling a neural network model to predict the regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the input data.

[0125] In a specific implementation, a computer device may directly input input data into a neural network model, causing the neural network model to extract features from the input data, thereby predicting the degree of regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the extracted features. It is understood that if the data size of the input data does not match the input size supported by the neural network model, the input data may be scaled based on the input size supported by the neural network model, and the scaled input data may be input into the neural network model for feature extraction and other processing.

[0126] S406, calling the neural network model to perform the binary classification task based on the obtained regional fitness to obtain the target classification result.

[0127] Among them, the target classification result may include: a first predicted probability value that there is a match between the texture map and the texture coordinate map, and a second predicted probability value that there is no match between the texture map and the texture coordinate map. It is understandable that the sum of the first predicted probability value and the second predicted probability value may be a reference value (such as a value of 1), and the first predicted probability value and the second predicted probability value may be two different probability values. After obtaining the target classification result, the neural network model can output the target classification result through two output dimensions in the output layer, and one output dimension is used to output a predicted probability value in the target classification result (such as a first predicted probability value or a second predicted probability value).

[0128] S407, after the neural network model outputs the target classification result through the two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

[0129] In a specific implementation, the computer device may compare the first predicted probability value and the second predicted probability value in the classification result. If the first predicted probability value is greater than the second predicted probability value, it is determined that a match between the texture map and the three-dimensional model has been detected, thereby generating a matching detection result indicating that the texture map and the three-dimensional model have a match. If the first predicted probability value is less than the second predicted probability value, it is determined that a match between the texture map and the three-dimensional model has not been detected, thereby generating a matching detection result indicating that the texture map and the three-dimensional model have not a match.

[0130] The embodiment of the present application can map the three-dimensional model of the object to the texture space to obtain a texture coordinate map, and then detect the matching between the texture map and the three-dimensional model (i.e., detect whether the texture map matches the three-dimensional model) based on the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map. This not only eliminates complex processing processes such as view rendering, thereby improving the efficiency of texture map matching detection and saving hardware resources; it also eliminates the process of manual participation in judgment, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection.

[0131] Based on the above Figure 2 and Figure 4 The relevant description of the method embodiment shown in the figure, the embodiment of the present application proposes an algorithm for determining whether the texture map matches based on the texture coordinate map of the three-dimensional model, which is used for asset cleaning and recall in the field of map matching. Among them, the algorithm mainly uses the relationship between the texture coordinate map generated after the three-dimensional model is expanded in the texture space and the texture map in terms of geometric shape and texture direction to determine the matching of the texture map with the three-dimensional model; specifically, in principle, it uses whether the texture coordinate areas of the texture map and the texture coordinate map of the three-dimensional model are compatible, that is, whether each texture image area in the texture map is compatible with the corresponding texture coordinate area in the texture coordinate map for one-to-one judgment. See Figure 6a Positive examples of 3D models matching texture maps, and Figure 6b The three-dimensional model shown in the figure does not match the negative sample of the texture map. It can be seen that there are obvious differences between the positive and negative samples in the texture space. This is the theoretical support and basis for the subsequent model design. Figure 6a and Figure 6b , images numbered 1-4 are the aforementioned texture coordinate map, texture map, combined effect map (i.e., the first overlay effect map), and transparent overlay effect map (i.e., the second overlay effect map), and image numbered 5 is a front view rendering obtained by applying the texture map to the three-dimensional model.

[0132] In a specific implementation, the algorithm proposed in the embodiment of the present application mainly includes the following parts:

[0133] (1) Model modification.

[0134] The algorithm proposed in the embodiment of the present application focuses on converting the visual discrimination task of map matching into a binary classification task based on regional similarity. Therefore, the image classification network can be structurally modified to obtain a neural network model. Among them, the embodiment of the present application does not restrict the Backbone (backbone network) part of the image classification network. It is recommended but not limited to using CNN (convolutional neural network) convolution layers and Res-Net network residual connection structure to extract features of input data. After each layer of feature extraction network, the Feature Map (feature map) is updated once.

[0135] The structural modification of the image classification model includes: modifying the number of channels of the input head (input layer) and output head (output layer) of the image classification network to adapt to the input of multi-channel image data and the requirements of the binary classification task. For example, the input channels in the input layer of the Rest-Net network (image classification network) are modified to 9 input channels, and the output dimension in the fully connected layer (i.e., the output layer) for output is modified to 2 output dimensions, corresponding to the classification labels 0 and 1 respectively; these two classification labels semantically correspond to the two types of results of map mismatch and map match, such as label 0 corresponds to mismatch and label 1 corresponds to match. It can be understood that the predicted value of each classification label is the predicted probability value in the corresponding classification result output by the above-mentioned neural network model.

[0136] (2) Model training.

[0137] (1) Dataset acquisition: The dataset acquisition stage is used to prepare the training dataset, and its method is also applicable to the preprocessing of subsequent inference data. First, the 3D model of the Mesh (skinned) object (i.e., the object) is parsed using 3D software to obtain the texture coordinate map (i.e., UV coordinate map) of the 3D model. After the texture coordinate regions in the texture coordinate map are expanded using the CV (computer vision) method, a mask coordinate map is obtained. The texture map (map file) to be identified is obtained, and the mask coordinate map and the texture map are superimposed using Boolean operations to obtain a combined effect map (i.e., the first superimposed effect map mentioned above). Based on this, batch operations can be performed to obtain a training dataset for model training or data for model inference.

[0138] (2) Data preprocessing and enhancement: The data preprocessing stage mainly involves stacking multiple images along the channel axis. For example, the image data of each channel in multiple images such as texture coordinate maps, mask coordinate maps, and combined effect maps can be stacked along the channel axis to obtain input data. It should be noted that the number of images used for stacking in this stage can be adjusted as needed, such as removing the mask coordinate map or adding a transparent overlay effect map (i.e., the second overlay effect map mentioned above). The purpose is to make the constructed input data contain sufficient pixel domain visual information. In addition, the data enhancement stage can increase the diversity of the training data set by introducing Gaussian noise, random flipping, random cropping, etc. in the input data, so as to improve the generalization performance of the subsequent neural network model.

[0139] (3) Model optimization: Based on the input size supported by the neural network model, the input data is uniformly scaled and input into the neural network model, which then performs feature extraction, regional fitness prediction, and binary classification tasks, thereby outputting the corresponding classification results. Based on the classification results, the model parameters of the neural network model are optimized to achieve pre-training of the neural network model and improve the performance of the neural network model.

[0140] (3) Model reasoning. After completing the pre-training of the neural network model, the input data can be constructed according to the same data pre-processing method and input into the neural network model to perform a binary classification task to determine whether the current texture map and the texture coordinate map match. The matching between the current texture map and the 3D model is determined based on the corresponding classification results. For details, see Figure 6c As shown in FIG, the process of model reasoning is as follows: parsing the three-dimensional model to obtain a texture coordinate map (i.e., a UV coordinate map); dilating each texture coordinate region in the texture coordinate map to obtain a mask coordinate map; obtaining the texture map (map file) to be identified, and superimposing the mask coordinate map and the texture map through Boolean operations to obtain a combined effect map; stacking multiple images such as the texture coordinate map, the mask coordinate map, and the combined effect map along the channel axis to obtain input data; after uniformly scaling the input data, inputting it into the pre-trained neural network model, the neural network model extracts features, and predicts regional fitness based on the extracted features and performs a binary classification task to obtain a target classification result, which includes a first predicted probability value (i.e., a predicted value of classification label 1) that the texture map and the texture coordinate map match, and a second predicted probability value (i.e., a predicted value of classification label 0) that the texture map and the texture coordinate map do not match; after the neural network model outputs the target classification result, whether the texture map and the three-dimensional model match can be detected based on the target classification result.

[0141] Based on the relevant description of the above algorithm, the embodiment of the present application also proposes a texture matching method based on deep learning technology. This method can automatically detect whether a three-dimensional model matches a texture map, thereby improving the efficiency and accuracy of texture matching. The method includes the following steps: first, using a deep learning method to train an image classifier (i.e., a neural network model) to determine whether the texture map and the texture coordinate map match; then, the texture coordinate map and texture map of the three-dimensional model to be matched are input into the image classifier to obtain a classification result, and then, based on the classification result, determine whether the three-dimensional model and the texture map match.

[0142] Practice has proven that the algorithms and corresponding methods proposed in the embodiments of this application can have the following beneficial effects:

[0143] (1) Improved the accuracy of texture matching: In actual applications, texture mapping is designed according to the design concept that a texture image area in the texture map is applied to a model area in the three-dimensional model. Therefore, by converting the standard of whether the texture map matches into the degree of adaptability between the texture image area in the texture map and the connected area (i.e., the texture coordinate area) in the texture coordinate map of the three-dimensional model and the rationality of the texture direction, whether the texture map and the three-dimensional model match can be detected, which can effectively improve the accuracy of matching detection. Its judgment principle follows the design concept in the art asset design process, which can solve the problem of low judgment accuracy at the principle level and eliminate the interference of subjective factors.

[0144] (2) Improved the efficiency of map matching: By engineering the entire algorithm and corresponding methods, the feasibility and reliability of the huge map matching task in actual production are greatly increased, and the efficiency of the map matching task is effectively improved.

[0145] (3) Improved applicability and flexibility of texture matching: The training dataset used for model training can be flexibly adjusted according to the requirements of the texture matching task. The corresponding expert model can be trained for the texture matching task in a certain field. For example, different datasets, network architectures, loss functions, etc. can be designed for 3D (three-dimensional) human assets, 3D object assets, etc., thereby improving the applicability and flexibility of the entire texture matching task.

[0146] (4) The repair and enhancement of three-dimensional model assets are realized: the algorithm and corresponding method proposed in the embodiment of the present application can not only repair the three-dimensional model with missing texture maps (i.e., find matching texture maps for the three-dimensional model with missing texture maps), thereby avoiding the waste of three-dimensional model assets, but also find more suitable texture maps for the three-dimensional model in the map library to generate more samples, thereby achieving the purpose of asset enhancement.

[0147] (5) Providing high-quality training data for 3D model assets for the subsequent development of AIGC (Artificial Intelligence Generative Content) pipelines and tools: The algorithm and corresponding method proposed in the embodiment of the present application can be used to recall more 3D model assets with texture matching from a large number of art asset databases with varying asset quality, providing high-quality training data for 3D model assets for the subsequent development of more AIGC pipelines and tools.

[0148] In summary, the algorithm and corresponding method proposed in the embodiments of the present application can automatically detect whether the three-dimensional model matches the texture map, avoid the subjectivity and inconsistency of manual judgment, and improve the efficiency and accuracy of map matching. The algorithm and method have broad application prospects and can provide strong technical support for three-dimensional modeling, texture map matching, AIGC data engineering and other fields.

[0149] Based on the description of the above-mentioned texture map matching detection method embodiment, the present application embodiment further discloses a texture map matching detection device; the texture map matching detection device can be a computer program (including one or more instructions) running on a computer device, and the texture map matching detection device can perform each step of the above-mentioned method embodiment. Figure 7 The texture map matching detection device can run the following units:

[0150] An acquisition unit 701 is configured to acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and each texture image region is used to determine the display effect of each model region;

[0151] The processing unit 702 is configured to map the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, and a texture coordinate region is obtained by mapping a model region included in the three-dimensional model;

[0152] The processing unit 702 is further configured to obtain a degree of regional compatibility between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping the model region determined by the corresponding texture image region;

[0153] The processing unit 702 is further configured to detect the matching between the texture map and the three-dimensional model based on the obtained regional fitness, and obtain a matching detection result.

[0154] In one embodiment, when the processing unit 702 is used to obtain the regional adaptation between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0155] Performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect image; and constructing input data of a neural network model using the image superposition effect image;

[0156] The neural network model is called to predict the regional adaptation between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the input data.

[0157] In another embodiment, the image superposition effect diagram includes: a first superposition effect diagram; accordingly, when the processing unit 702 is used to perform image superposition on the texture map and the texture coordinate map to obtain the image superposition effect diagram, it can be specifically used to:

[0158] Obtaining a mask coordinate map corresponding to the texture coordinate map, wherein the mask coordinate map includes a plurality of mask regions, and one mask region corresponds to one texture coordinate region in the texture coordinate map;

[0159] Perform image superposition on the texture map and the mask coordinate map to obtain a first superposition effect image.

[0160] In another embodiment, when the processing unit 702 is used to obtain the mask coordinate map corresponding to the texture coordinate map, it can be specifically used to:

[0161] dilating each texture coordinate region in the texture coordinate map along its boundary; wherein the dilating process for the i-th texture coordinate region comprises: sampling a plurality of points on the boundary of the i-th texture coordinate region, determining an extension coefficient of the corresponding point according to a curvature of each point, and extending the corresponding point along a normal of the corresponding point according to the extension coefficient to dilate the i-th texture coordinate region, where i is a positive integer and is less than or equal to the number of the texture coordinate regions;

[0162] In the texture coordinate map, mask processing is performed on each expanded texture coordinate region to obtain a coordinate mask map.

[0163] In another embodiment, the image superposition effect diagram includes: a second superposition effect diagram; accordingly, when the processing unit 702 is used to perform image superposition on the texture map and the texture coordinate map to obtain the image superposition effect diagram, it can be specifically used to:

[0164] Performing transparency processing on the texture coordinate map using a preset transparency to obtain a texture coordinate map after transparency processing;

[0165] The texture coordinate image after the transparency processing and the texture map are superimposed to obtain a second superimposed effect image.

[0166] In another embodiment, any image includes N channels of image data, where N is a positive integer; accordingly, when the processing unit 702 is used to construct input data of the neural network model using the image superposition effect diagram, it can be specifically used to:

[0167] Acquire image data of each channel in the texture map, image data of each channel in the texture coordinate map, and image data of each channel in the image overlay effect map;

[0168] In the channel axis direction, the acquired image data of M channels are stacked to obtain the input data of the neural network model; where M is an integer greater than N.

[0169] In another embodiment, the input layer of the neural network model includes M input channels, and different input channels are used to input image data of different channels in the input data.

[0170] In another embodiment, the output layer of the neural network model includes two output dimensions, and different output dimensions are used to output different prediction probability values;

[0171] Accordingly, when the processing unit 702 is used to detect the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtain the matching detection result, it can be specifically used to:

[0172] Calling the neural network model to perform a binary classification task based on the obtained regional fitness to obtain a target classification result; the target classification result includes: a first predicted probability value that there is a match between the texture map and the texture coordinate map, and a second predicted probability value that there is no match between the texture map and the texture coordinate map;

[0173] After the neural network model outputs the target classification result through two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

[0174] In another embodiment, when the processing unit 702 is used to obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0175] For a p-th texture image region in the texture map, calculating an intersection-over-union ratio between the p-th texture image region and a corresponding texture coordinate region in the texture coordinate map, where the value of p is a positive integer and is less than or equal to the number of texture image regions included in the texture map;

[0176] The calculated intersection-over-union ratio is used as the region adaptation degree between the p-th texture image region and the corresponding texture coordinate region.

[0177] In another embodiment, when the processing unit 702 is used to obtain the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, it can be specifically used to:

[0178] For a p-th texture image region in the texture map, Q points are selected as boundary points on the region boundary of the p-th texture coordinate region corresponding to the p-th texture image region, where Q is a positive integer;

[0179] Sampling R pixels in the texture map along the normal of the qth boundary point, where q∈[1,Q], R is an integer greater than 1, the R pixels are all located on the normal of the qth boundary point, and at least two of the R pixels are located on both sides of the region boundary of the pth texture coordinate region;

[0180] Determining pixel values ​​of the R pixel points based on the texture map, and calculating pixel value gradients pixel by pixel in the R pixel points according to the pixel values ​​of the R pixel points; and integrating the calculated pixel value gradients to obtain the pixel value gradient of the qth boundary point;

[0181] The degree of regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region is determined according to the pixel value gradients of the Q boundary points.

[0182] In another embodiment, when the processing unit 702 is configured to detect the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtain the matching detection result, it can be specifically configured to:

[0183] Determine the regional weight of each texture image region based on the importance of the model region corresponding to each texture image region, where the regional weight is proportional to the importance;

[0184] Using the regional weight of each texture image region, weighted summing is performed on the regional fitness corresponding to the acquired corresponding texture image region to obtain the image fitness between the texture map and the texture coordinate map;

[0185] Based on the image fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

[0186] According to another embodiment of the present application, Figure 7 The various units in the texture mapping matching detection device shown can be individually or all combined into one or several other units to form, or one (or some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the texture mapping-based matching detection device can also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0187] According to another embodiment of the present application, a computer program (including one or more instructions) capable of executing each step involved in the above method embodiment can be constructed by running the computer program (including one or more instructions) on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). Figure 7 The texture map matching detection device shown in the , and the various methods proposed in the embodiments of the present application are implemented. The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.

[0188] It is worth noting that, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can include a part of the overall module or unit of the module or unit function.

[0189] The embodiment of the present application can map the three-dimensional model of the object to the texture space to obtain a texture coordinate map, and then detect the matching between the texture map and the three-dimensional model (i.e., detect whether the texture map matches the three-dimensional model) based on the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map. This not only eliminates complex processing processes such as view rendering, thereby improving the efficiency of texture map matching detection and saving hardware resources; it also eliminates the process of manual participation in judgment, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection.

[0190] Based on the description of the above method embodiment and apparatus embodiment, the present application embodiment also provides a computer device. Figure 8 , the computer device at least includes a processor 801, an input interface 802, an output interface 803 and a computer storage medium 804. Among them, the processor 801, input interface 802, output interface 803 and computer storage medium 804 in the computer device can be connected via a bus or other means. The computer storage medium 804 can be stored in the memory of the computer device, and the computer storage medium 804 is used to store a computer program, and the computer program includes one or more instructions. The processor 801 is used to execute one or more instructions in the computer program stored in the computer storage medium 804. The processor 801 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to realize the corresponding method flow or corresponding function.

[0191] In one embodiment, the processor 801 described in the embodiment of the present application can be used to perform a series of detection processes on the texture map, specifically including: obtaining a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes multiple model areas; the texture map includes multiple texture image areas, and one texture image area is used to determine the display effect of a model area; mapping the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes multiple texture coordinate areas, and one texture coordinate area is obtained by mapping a model area included in the three-dimensional model; obtaining the regional adaptability between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map; wherein the texture coordinate area corresponding to the texture image area refers to: the texture coordinate area obtained by mapping the model area determined by the corresponding texture image area; based on the obtained regional adaptability, detecting the matching between the texture map and the three-dimensional model to obtain a matching detection result, and so on.

[0192] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing computer programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space, which stores the operating system of the computer device. In addition, a computer program is also stored in the storage space, which includes one or more instructions suitable for being loaded and executed by the processor 801, and these instructions can be one or more program codes. It should be noted that the computer storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.

[0193] In one embodiment, a processor may load and execute one or more instructions stored in a computer storage medium to implement the corresponding steps in the above method embodiment. In a specific implementation, the processor may load and execute the following steps:

[0194] Acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and each texture image region is used to determine the display effect of each model region;

[0195] Mapping the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, and a texture coordinate region is obtained by mapping a model region included in the three-dimensional model;

[0196] Obtaining a degree of regional fit between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping a model region determined by the corresponding texture image region;

[0197] Based on the obtained regional fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

[0198] In one embodiment, when obtaining the regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, the one or more instructions may be loaded and specifically executed by the processor:

[0199] Performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect image; and constructing input data of a neural network model using the image superposition effect image;

[0200] The neural network model is called to predict the regional adaptation between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the input data.

[0201] In another embodiment, the image overlay effect diagram includes: a first overlay effect diagram; accordingly, when the texture map and the texture coordinate map are overlaid to obtain the image overlay effect diagram, the one or more instructions may be loaded and specifically executed by the processor:

[0202] Obtaining a mask coordinate map corresponding to the texture coordinate map, wherein the mask coordinate map includes a plurality of mask regions, and one mask region corresponds to one texture coordinate region in the texture coordinate map;

[0203] Perform image superposition on the texture map and the mask coordinate map to obtain a first superposition effect image.

[0204] In another embodiment, when obtaining the mask coordinate map corresponding to the texture coordinate map, the one or more instructions may be loaded and specifically executed by the processor:

[0205] dilating each texture coordinate region in the texture coordinate map along its boundary; wherein the dilating process for the i-th texture coordinate region comprises: sampling a plurality of points on the boundary of the i-th texture coordinate region, determining an extension coefficient of the corresponding point according to a curvature of each point, and extending the corresponding point along a normal of the corresponding point according to the extension coefficient to dilate the i-th texture coordinate region, where i is a positive integer and is less than or equal to the number of the texture coordinate regions;

[0206] In the texture coordinate map, mask processing is performed on each expanded texture coordinate region to obtain a coordinate mask map.

[0207] In another embodiment, the image overlay effect diagram includes: a second overlay effect diagram; accordingly, when the texture map and the texture coordinate map are overlaid to obtain the image overlay effect diagram, the one or more instructions may be loaded and specifically executed by the processor:

[0208] Performing transparency processing on the texture coordinate map using a preset transparency to obtain a texture coordinate map after transparency processing;

[0209] The texture coordinate image after the transparency processing and the texture map are superimposed to obtain a second superimposed effect image.

[0210] In another embodiment, any image includes N channels of image data, where N is a positive integer; accordingly, when the image superposition effect diagram is used to construct input data of a neural network model, the one or more instructions may be loaded and specifically executed by the processor:

[0211] Acquire image data of each channel in the texture map, image data of each channel in the texture coordinate map, and image data of each channel in the image overlay effect map;

[0212] In the channel axis direction, the acquired image data of M channels are stacked to obtain the input data of the neural network model; where M is an integer greater than N.

[0213] In another embodiment, the input layer of the neural network model includes M input channels, and different input channels are used to input image data of different channels in the input data.

[0214] In another embodiment, the output layer of the neural network model includes two output dimensions, and different output dimensions are used to output different prediction probability values;

[0215] Accordingly, when detecting the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtaining a matching detection result, the one or more instructions may be loaded and specifically executed by the processor:

[0216] Calling the neural network model to perform a binary classification task based on the obtained regional fitness to obtain a target classification result; the target classification result includes: a first predicted probability value that there is a match between the texture map and the texture coordinate map, and a second predicted probability value that there is no match between the texture map and the texture coordinate map;

[0217] After the neural network model outputs the target classification result through two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

[0218] In another embodiment, when obtaining the regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, the one or more instructions may be loaded and specifically executed by the processor:

[0219] For a p-th texture image region in the texture map, calculating an intersection-over-union ratio between the p-th texture image region and a corresponding texture coordinate region in the texture coordinate map, where the value of p is a positive integer and is less than or equal to the number of texture image regions included in the texture map;

[0220] The calculated intersection-over-union ratio is used as the region adaptation degree between the p-th texture image region and the corresponding texture coordinate region.

[0221] In another embodiment, when obtaining the regional fit between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map, the one or more instructions may be loaded and specifically executed by the processor:

[0222] For a p-th texture image region in the texture map, Q points are selected as boundary points on the region boundary of the p-th texture coordinate region corresponding to the p-th texture image region, where Q is a positive integer;

[0223] Sampling R pixels in the texture map along the normal of the qth boundary point, where q∈[1,Q], R is an integer greater than 1, the R pixels are all located on the normal of the qth boundary point, and at least two of the R pixels are located on both sides of the region boundary of the pth texture coordinate region;

[0224] Determining pixel values ​​of the R pixel points based on the texture map, and calculating pixel value gradients pixel by pixel in the R pixel points according to the pixel values ​​of the R pixel points; and integrating the calculated pixel value gradients to obtain the pixel value gradient of the qth boundary point;

[0225] The degree of regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region is determined according to the pixel value gradients of the Q boundary points.

[0226] In another embodiment, when detecting the matching between the texture map and the three-dimensional model based on the obtained regional fitness and obtaining the matching detection result, the one or more instructions may be loaded and specifically executed by the processor:

[0227] Determine the regional weight of each texture image region based on the importance of the model region corresponding to each texture image region, where the regional weight is proportional to the importance;

[0228] Using the regional weight of each texture image region, weighted summing is performed on the regional fitness corresponding to the acquired corresponding texture image region to obtain the image fitness between the texture map and the texture coordinate map;

[0229] Based on the image fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

[0230] The embodiment of the present application can map the three-dimensional model of the object to the texture space to obtain a texture coordinate map, and then detect the matching between the texture map and the three-dimensional model (i.e., detect whether the texture map matches the three-dimensional model) based on the regional fit between each texture image area in the texture map and the corresponding texture coordinate area in the texture coordinate map. This not only eliminates complex processing processes such as view rendering, thereby improving the efficiency of texture map matching detection and saving hardware resources; it also eliminates the process of manual participation in judgment, thereby saving labor costs and objectively detecting whether the texture map and the three-dimensional model match, avoiding the problem of low detection accuracy due to human subjective intentions, thereby improving the accuracy of texture map matching detection.

[0231] It should be noted that, according to one aspect of the embodiments of the present application, a computer program product or computer program is also provided, and the computer program product or computer program includes one or more instructions, and the one or more instructions are stored in a computer storage medium. The processor of the computer device reads the one or more instructions from the computer storage medium, and the processor executes the one or more instructions, so that the computer device executes the method provided in various optional ways in the above method embodiments. It should be understood that what is disclosed above is only a preferred embodiment of the present application, and it is certainly not used to limit the scope of rights of the present application. Therefore, equivalent changes made in accordance with the claims of the present application are still within the scope covered by the present application.

Claims

1. A texture map matching detection method, characterized in that: include: Acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and one texture image region is applied to one model region; Mapping the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, and a texture coordinate region is obtained by mapping a model region included in the three-dimensional model; Obtaining a degree of regional fit between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping the model region to which the corresponding texture image region is applied; Based on the obtained regional fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

2. The method according to claim 1, wherein The obtaining of the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map comprises: Performing image superposition on the texture map and the texture coordinate map to obtain an image superposition effect image; and constructing input data of a neural network model using the image superposition effect image; The neural network model is called to predict the regional adaptation between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map based on the input data.

3. The method according to claim 2, wherein The image superposition effect diagram includes: a first superposition effect diagram; the image superposition effect diagram is obtained by superimposing the texture map and the texture coordinate map, including: Obtaining a mask coordinate map corresponding to the texture coordinate map, wherein the mask coordinate map includes a plurality of mask regions, and one mask region corresponds to one texture coordinate region in the texture coordinate map; Perform image superposition on the texture map and the mask coordinate map to obtain a first superposition effect image.

4. The method according to claim 3, wherein The obtaining of the mask coordinate map corresponding to the texture coordinate map includes: dilating each texture coordinate region in the texture coordinate map along its boundary; wherein the dilating process for the i-th texture coordinate region comprises: sampling a plurality of points on the boundary of the i-th texture coordinate region, determining an extension coefficient of the corresponding point according to a curvature of each point, and extending the corresponding point along a normal of the corresponding point according to the extension coefficient to dilate the i-th texture coordinate region, where i is a positive integer and is less than or equal to the number of the texture coordinate regions; In the texture coordinate map, mask processing is performed on each expanded texture coordinate region to obtain a coordinate mask map.

5. The method according to claim 2, wherein The image superposition effect diagram includes: a second superposition effect diagram; the image superposition effect diagram is obtained by superimposing the texture map and the texture coordinate map, including: Performing transparency processing on the texture coordinate map using a preset transparency to obtain a texture coordinate map after transparency processing; The texture coordinate image after the transparency processing and the texture map are superimposed to obtain a second superimposed effect image.

6. The method according to claim 2, wherein Any image includes N channels of image data, where N is a positive integer; the input data for constructing a neural network model by using the image superposition effect diagram includes: Acquire image data of each channel in the texture map, image data of each channel in the texture coordinate map, and image data of each channel in the image overlay effect map; In the channel axis direction, the acquired image data of M channels are stacked to obtain the input data of the neural network model; where M is an integer greater than N.

7. The method according to claim 6, wherein The input layer of the neural network model includes M input channels, and different input channels are used to input image data of different channels in the input data.

8. The method according to claim 2, wherein The output layer of the neural network model includes two output dimensions, and different output dimensions are used to output different prediction probability values; The detecting the matching between the texture map and the three-dimensional model based on the obtained regional fitness, and obtaining a matching detection result, includes: Calling the neural network model to perform a binary classification task based on the obtained regional fitness to obtain a target classification result; the target classification result includes: a first predicted probability value that there is a match between the texture map and the texture coordinate map, and a second predicted probability value that there is no match between the texture map and the texture coordinate map; After the neural network model outputs the target classification result through two output dimensions in the output layer, the matching between the texture map and the three-dimensional model is detected based on the target classification result to obtain a matching detection result.

9. The method according to claim 1, wherein The obtaining of the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map comprises: For a p-th texture image region in the texture map, calculating an intersection-over-union ratio between the p-th texture image region and a corresponding texture coordinate region in the texture coordinate map, where the value of p is a positive integer and is less than or equal to the number of texture image regions included in the texture map; The calculated intersection-over-union ratio is used as the region adaptation degree between the p-th texture image region and the corresponding texture coordinate region.

10. The method according to claim 1, wherein The obtaining of the regional adaptability between each texture image region in the texture map and the corresponding texture coordinate region in the texture coordinate map comprises: For a p-th texture image region in the texture map, Q points are selected as boundary points on the region boundary of the p-th texture coordinate region corresponding to the p-th texture image region, where Q is a positive integer; Sampling R pixels in the texture map along the normal of the qth boundary point, where q∈[1,Q], R is an integer greater than 1, the R pixels are all located on the normal of the qth boundary point, and at least two of the R pixels are located on both sides of the region boundary of the pth texture coordinate region; Determining pixel values ​​of the R pixel points based on the texture map, and calculating pixel value gradients pixel by pixel in the R pixel points according to the pixel values ​​of the R pixel points; and integrating the calculated pixel value gradients to obtain the pixel value gradient of the qth boundary point; The degree of regional adaptation between the p-th texture image region and the corresponding p-th texture coordinate region is determined according to the pixel value gradients of the Q boundary points.

11. The method according to claim 9 or 10, wherein: The detecting the matching between the texture map and the three-dimensional model based on the obtained regional fitness, and obtaining a matching detection result, includes: Determine the regional weight of each texture image region based on the importance of the model region corresponding to each texture image region, where the regional weight is proportional to the importance; Using the regional weight of each texture image region, weighted summing is performed on the regional fitness corresponding to the acquired corresponding texture image region to obtain the image fitness between the texture map and the texture coordinate map; Based on the image fitness, the matching between the texture map and the three-dimensional model is detected to obtain a matching detection result.

12. A texture map matching detection device, characterized in that: include: An acquisition unit, configured to acquire a three-dimensional model and a texture map of an object, wherein the three-dimensional model includes a plurality of model regions; the texture map includes a plurality of texture image regions, and a texture image region is used to determine a display effect of a model region; a processing unit configured to map the three-dimensional model to a texture space to obtain a texture coordinate map; the texture coordinate map includes a plurality of texture coordinate regions, where a texture coordinate region is obtained by mapping a model region included in the three-dimensional model; The processing unit is further configured to obtain a degree of regional fit between each texture image region in the texture map and a corresponding texture coordinate region in the texture coordinate map; wherein the texture coordinate region corresponding to the texture image region is a texture coordinate region obtained by mapping a model region determined by the corresponding texture image region; The processing unit is further configured to detect the matching between the texture map and the three-dimensional model based on the acquired regional fitness, and obtain a matching detection result.

13. A computer device comprising an input interface and an output interface, characterized in that: Also includes: processors and computer storage media; The processor is adapted to implement one or more instructions, the computer storage medium stores one or more instructions, and the one or more instructions are adapted to be loaded by the processor and executed by the texture map matching detection method according to any one of claims 1 to 11.

14. A computer storage medium, characterized in that The computer storage medium stores one or more instructions, and the one or more instructions are suitable for being loaded by a processor and executing the texture map matching detection method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The computer program product includes one or more instructions; when the one or more instructions in the computer program are executed by a processor, the texture map matching detection method according to any one of claims 1 to 11 is implemented.