Graphic image processing system for computer
By building a GANs model and deep learning technology to generate realistic 3D models, the problem of scarcity of high-quality data is solved, especially in the medical and aerospace fields, providing the system with rich data resources and improving animation quality and processing speed.
Patent Information
- Application Number
- CN202510694194.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, high-quality and diverse 3D animation data is difficult to obtain in specific fields such as medicine and aerospace, resulting in limited system performance, reduced animation quality or slower processing speed.
The GANs model is combined with TensorFlow to build the generator and discriminator, and realistic three-dimensional models are generated through adversarial training. The physical lighting model and environment mapping technology are used to increase the realism and diversity of the data, and deep learning is combined for dynamic effect processing and comprehensive evaluation.
Rich and diverse 3D character model data is generated to meet the data needs of specific fields, improve the system's data resources and processing efficiency, and achieve customized style transfer effects.
Smart Images

Figure CN120672915A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of graphics and image processing, and in particular to a graphics and image processing system for a computer. Background Art
[0002] With the rapid development of computer technology, graphics and image processing techniques are constantly innovating. The integration of AI technology has brought unprecedented changes to 3D animation graphics and image processing. AI technology can improve the efficiency and accuracy of graphics and image processing, making 3D animation production more realistic and vivid. 3D animation has a wide range of applications in film and television production, game development, advertising design, architectural representation, and other fields, and market demand continues to grow.
[0003] After searching, the invention patent with Chinese patent number CN118918224A discloses a 3D animation graphics image processing system based on AI artificial intelligence, which belongs to the field of computer graphics technology. It includes a 3D image scene determination module, a 3D scene image information acquisition module, a data standardization processing module, a 3D image text processing module, a 3D image text comprehensive processing module, an image processing comprehensive judgment module and an interactive feedback module. The present invention realizes comprehensive monitoring of the target area through precise area division, and covers the key information of the three dimensions of picture quality, rendering effect and efficiency by collecting frame rate, QP value, number of grids and textures, pixel accuracy, error, rendering time, memory occupancy and image quality indicators (SSIM, PSNR), etc. A comprehensive mathematical model is established to comprehensively evaluate the comprehensive performance of the 3D animation graphics image processing system, thereby improving the work efficiency and quality of 3D animation graphics image processing.
[0004] The above-mentioned system aims to improve the efficiency and quality of 3D animation production. To achieve this goal, the system relies on a large amount of 3D model, image, and video data to train and optimize its algorithms. Although a large amount of 3D model, image, and video data exists online, high-quality and diverse data is relatively scarce. High-quality data is particularly difficult to obtain in certain specific fields or industries, such as medicine and aerospace. Without sufficient high-quality data to train the algorithm, the system's performance may be limited. This can lead to reduced animation quality or slower processing speeds. Therefore, a computer graphics and image processing system is proposed. Summary of the Invention
[0005] The present invention aims to address the relative scarcity of high-quality, diverse data, a problem currently faced by existing technologies. In particular, high-quality data is particularly difficult to obtain in certain fields or industries, such as healthcare and aerospace. Without sufficient high-quality data to train algorithms, system performance may be limited. This can lead to reduced animation quality or slower processing speeds. The present invention proposes a computer graphics and image processing system.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A computer graphics image processing system, comprising:
[0008] 3D image scene determination module: automatically identifies and analyzes animation scenes, determines objects, light sources, materials and other elements in the scene, and generates a preliminary 3D model based on the scene information;
[0009] GANs model module: Utilizes TensorFlow to build a GANs model, which includes a generator and a discriminator. Adopting an adversarial training strategy, the generator learns the distribution characteristics of the data. The generator is responsible for generating realistic 3D models. Based on a physical lighting model and environment mapping technology, it simulates different lighting conditions and background environments to increase the fidelity and diversity of the generated data. The discriminator distinguishes between real data and generated data. New 3D models are generated by adjusting the input noise vector of the GANs. By adjusting the structures and loss functions of the generator and discriminator, specific artistic styles can be transferred to the generated 3D models.
[0010] Information collection module: collects multi-dimensional information including frame rate, QP value, number of grids, pixel accuracy and rendering time, and monitors various key performance indicators of animation production. The information collection module collects preliminary 3D models and new 3D model data. The multi-dimensional information collected by the information collection module is passed to the data standardization processing module;
[0011] Data standardization processing module: cleans, converts and standardizes the collected data to ensure data consistency and accuracy. The standardized data is input into the image processing module;
[0012] Image processing module: uses deep learning to process dynamic effects and image details. The image data processed by the image processing module is passed to the comprehensive judgment module;
[0013] Comprehensive judgment module: Build a comprehensive mathematical model to evaluate rendering effect, image quality and efficiency;
[0014] User interaction module: provides users with an operation interface and result display.
[0015] The above technical solution further includes:
[0016] Furthermore, the specific steps of the 3D image scene determination module generating a preliminary 3D model are:
[0017] Scene capture and preprocessing: Capture animation scenes through a camera or other image acquisition device, and perform preprocessing on the captured images, such as denoising and contrast enhancement;
[0018] Feature extraction: Image processing and computer vision techniques are used to extract key features from pre-processed images, such as objects, light sources, and materials. These features may include edges, corners, textures, and colors.
[0019] 3D modeling: Based on the extracted features, a preliminary 3D model is generated using 3D modeling technology, which involves steps such as point cloud generation, mesh construction, and texture mapping.
[0020] Furthermore, the GANs model module includes a generator, a discriminator, a physical lighting model and environment mapping technology unit and a style transfer unit. The generator is responsible for generating a three-dimensional model from random noise. The discriminator is responsible for distinguishing whether the input data is real data (from a real data set) or fake data generated by the generator, and outputting a probability value indicating the confidence that the data is real data. The physical lighting model and environment mapping technology unit are used to simulate different lighting conditions and background environments to increase the realism and diversity of the generated data. The physical lighting model calculates the lighting effect based on physical principles. The environment mapping technology uses pre-made environment maps to simulate the background environment. The style transfer unit is responsible for migrating a specific artistic style to the generated three-dimensional model. Style transfer is achieved by adjusting the input noise vector of GANs and the structure, loss function and other parameters of the generator and discriminator.
[0021] Furthermore, the structure of the generator includes:
[0022] Input layer: Receives a random noise vector that follows a Gaussian or uniform distribution. This vector serves as the starting point of the generation process and contains all the potential information needed to generate an image or model, including but not limited to shape and structure information, texture and material information, motion and dynamic information, lighting and shadow information, and style and emotion information.
[0023] Convolutional transpose layer: used to upsample the input vector to a higher spatial resolution. Through the inverse process of the convolution operation, it maps the low-dimensional features to a high-dimensional space to generate an image or model with a more detailed structure. The calculation formula is expressed as Y = ConvTranspose(X,W,b,s,p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the stride, which controls the spatial resolution of the output feature map; p is padding, which is used to add additional zero values around the input feature map to control the size of the output feature map.
[0024] Batch Normalization Layer: Normalizes the output of the convolutional transpose layer so that the input data of each batch has the same distribution, which helps to speed up the training process and improve the convergence speed and stability of the model. The normalization process is expressed as Where X is the input feature map; μ and σ 2 are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a small positive number used to avoid division by zero errors;
[0025] Activation function layer: introduces nonlinear factors to enable the generator to learn complex feature representations. Common activation functions include ReLU, Leaky ReLU, Sigmoid, etc.
[0026] Output layer: Generate a 3D model.
[0027] Furthermore, the structure of the discriminator includes:
[0028] Input layer: receives data from the real dataset or generator as input;
[0029] Convolution layer: extracts features of input data through convolution operation. The convolution kernel slides on the input data, calculates local features at each position, and generates a feature map. Multiple convolution layers are stacked together to extract deeper features. The calculation formula is expressed as Y = Conv(X, W, b, s, p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the stride, which controls the sliding distance of the convolution kernel on the input feature map; p is padding, which is used to add extra zero values around the input feature map to control the size of the output feature map;
[0030] Batch Normalization Layer: The batch normalization layer normalizes the output of the convolutional layer so that the input data of each batch has the same distribution. The normalization process is expressed as Where X is the input feature map; μ and σ 2are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a small positive number used to avoid division by zero errors;
[0031] Activation function layer: The activation function layer introduces nonlinear factors, allowing the discriminator to learn complex feature representations;
[0032] Fully connected layer: The fully connected layer converts the output of the convolutional layer and the activation function layer into a fixed-length feature vector, which is used for the final classification decision;
[0033] Output layer: The output layer is a single neuron that uses the Sigmoid activation function to output a probability value, which indicates the confidence that the input data is real data. The closer the probability value is to 1, the more real the input data is; the closer the probability value is to 0, the more false the input data is.
[0034] Furthermore, the physical illumination model calculates the illumination effect based on physical principles, and the environment mapping technology uses a pre-made environment map to simulate the background environment, including the following steps:
[0035] Light source definition: define the position, color, intensity and type of light source (such as point light source, parallel light source, etc.);
[0036] Surface material properties: define the reflectivity, refractive index, roughness and other properties of the object surface;
[0037] Lighting calculation: Diffuse reflection: random reflection of light on a rough surface. The formula is I diffuse =k d I s max(0,n·l), where k d is the diffuse reflectance, I s is the intensity of the light source, n is the surface normal, and l is the direction vector from the surface to the light source; Specular reflection: directional reflection of light on a smooth surface, the formula is I pecular =k s I s ·(r·v) d , where k s is the specular reflection coefficient, r is the reflection direction, v is the viewing direction, and d is the specular index; Ambient light: global illumination from the surrounding environment, usually expressed as a constant;
[0038] Final color calculation: Add diffuse reflection, specular reflection and ambient light to get the final color value;
[0039] Environment map production: Use panoramic camera or 3D modeling software to produce environment maps;
[0040] Texture Mapping: Map the environment map onto a sphere or other geometry to simulate the surrounding environment;
[0041] Reflection calculation: During the rendering process, the color value is sampled from the environment map according to the surface properties of the object and the viewing direction as the color of the reflected light.
[0042] Furthermore, the comprehensive judgment module includes a mathematical model construction unit, an evaluation index calculation unit and a result feedback generation unit. The mathematical model construction unit constructs a suitable mathematical model according to the specific requirements of image processing, which is used to analyze and evaluate the image data; the evaluation index calculation unit calculates the evaluation index of the image according to the mathematical model, and the result feedback generation unit generates comprehensive feedback on the image processing effect according to the results of the evaluation index calculation unit.
[0043] Furthermore, the evaluation indicators include resolution, color fidelity, contrast, frame rate and processing time;
[0044] The resolution is represented by the dimensions of the image (such as width and height);
[0045] The color fidelity is measured using the peak signal-to-noise ratio, and the formula for the peak signal-to-noise ratio is: Wherein, MAX represents the maximum possible pixel value in the image (for an 8-bit image, usually 255), and MSE represents the mean square error of the image. The formula for the mean square error of the image is: Where I(i, j) represents the pixel value at position (i, j) of the original image, K(i, j) represents the pixel value at the same position of the processed image, M and N represent the width and height of the image respectively;
[0046] The contrast ratio measures the difference between bright and dark parts of an image using a contrast ratio formula, where σ represents the standard deviation of the image pixel values, and μ represents the mean of the image pixel values;
[0047] The frame rate is directly obtained through measurement or statistics. For video processing, the frame rate refers to the number of frames displayed per second;
[0048] The processing time is measured by a timer and represents the time required to process one image or one video.
[0049] The present invention has the following beneficial effects:
[0050] In this paper, by adjusting the input noise vector of GANs, this technical solution can generate new and diverse 3D character model data. This solves the problem of scarce high-quality data, especially in specific fields such as medicine and aerospace, providing the system with a rich data resource. By adjusting the parameters and structure of GANs, customized style transfer effects can be achieved to meet the diverse needs of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a system block diagram of a computer graphics image processing system proposed by the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] See also Figure 1 As shown, the present invention is a system for computer graphics image processing, comprising:
[0054] 3D image scene determination module: automatically identifies and analyzes animation scenes, determines objects, light sources, materials and other elements in the scene, and generates a preliminary 3D model based on the scene information;
[0055] GANs model module: Utilizes TensorFlow to build a GANs model, which includes a generator and a discriminator. Adopting an adversarial training strategy, the generator learns the distribution characteristics of the data. The generator is responsible for generating realistic 3D models. Based on a physical lighting model and environment mapping technology, it simulates different lighting conditions and background environments to increase the fidelity and diversity of the generated data. The discriminator distinguishes between real data and generated data. New 3D models are generated by adjusting the input noise vector of the GANs. By adjusting the structures and loss functions of the generator and discriminator, specific artistic styles can be transferred to the generated 3D models.
[0056] Information collection module: collects multi-dimensional information including frame rate, QP value, number of grids, pixel accuracy and rendering time, and monitors various key performance indicators of animation production. The information collection module collects preliminary 3D models and new 3D model data. The multi-dimensional information collected by the information collection module is passed to the data standardization processing module;
[0057] Data standardization processing module: cleans, converts and standardizes the collected data to ensure data consistency and accuracy. The standardized data is input into the image processing module;
[0058] Image processing module: uses deep learning to process dynamic effects and image details. The image data processed by the image processing module is passed to the comprehensive judgment module;
[0059] Comprehensive judgment module: Build a comprehensive mathematical model to evaluate rendering effect, image quality and efficiency;
[0060] User interaction module: provides users with an operation interface and result display.
[0061] In one embodiment, the specific steps of the 3D image scene determination module generating a preliminary 3D model are:
[0062] Scene capture and preprocessing: Capture animation scenes through a camera or other image acquisition device, and perform preprocessing on the captured images, such as denoising and contrast enhancement;
[0063] Feature extraction: Image processing and computer vision techniques are used to extract key features from pre-processed images, such as objects, light sources, and materials. These features may include edges, corners, textures, and colors.
[0064] 3D modeling: Based on the extracted features, a preliminary 3D model is generated using 3D modeling technology, which involves steps such as point cloud generation, mesh construction, and texture mapping.
[0065] In one embodiment, the GANs model module includes a generator, a discriminator, a physical lighting model and environment mapping technology unit, and a style transfer unit. The generator is responsible for generating a three-dimensional model from random noise. The discriminator is responsible for distinguishing whether the input data is real data (from a real data set) or fake data generated by the generator, and outputting a probability value indicating the confidence that the data is real data. The physical lighting model and environment mapping technology unit are used to simulate different lighting conditions and background environments to increase the realism and diversity of the generated data. The physical lighting model calculates the lighting effect based on physical principles. The environment mapping technology uses pre-made environment maps to simulate the background environment. The style transfer unit is responsible for migrating a specific artistic style to the generated three-dimensional model. Style transfer is achieved by adjusting the input noise vector of GANs and the structure, loss function and other parameters of the generator and discriminator.
[0066] In one embodiment, the structure of the generator includes:
[0067] Input layer: Receives a random noise vector that follows a Gaussian or uniform distribution. This vector serves as the starting point of the generation process and contains all the potential information needed to generate an image or model, including but not limited to shape and structure information, texture and material information, motion and dynamic information, lighting and shadow information, and style and emotion information.
[0068] Convolutional transpose layer: used to upsample the input vector to a higher spatial resolution. Through the inverse process of the convolution operation, it maps the low-dimensional features to a high-dimensional space to generate an image or model with a more detailed structure. The calculation formula is expressed as Y = ConvTranspose(X,W,b,s,p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the stride, which controls the spatial resolution of the output feature map; p is padding, which is used to add additional zero values around the input feature map to control the size of the output feature map.
[0069] Batch Normalization Layer: Normalizes the output of the convolutional transpose layer so that the input data of each batch has the same distribution, which helps to speed up the training process and improve the convergence speed and stability of the model. The normalization process is expressed as Where X is the input feature map; μ and σ 2 are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a small positive number used to avoid division by zero errors;
[0070] Activation function layer: introduces nonlinear factors to enable the generator to learn complex feature representations. Common activation functions include ReLU, Leaky ReLU, Sigmoid, etc.
[0071] Output layer: Generate a 3D model.
[0072] In one embodiment, the structure of the discriminator includes:
[0073] Input layer: receives data from the real dataset or generator as input;
[0074] Convolution layer: extracts features of input data through convolution operation. The convolution kernel slides on the input data, calculates local features at each position, and generates a feature map. Multiple convolution layers are stacked together to extract deeper features. The calculation formula is expressed as Y = Conv(X, W, b, s, p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the stride, which controls the sliding distance of the convolution kernel on the input feature map; p is padding, which is used to add extra zero values around the input feature map to control the size of the output feature map;
[0075] Batch Normalization Layer: The batch normalization layer normalizes the output of the convolutional layer so that the input data of each batch has the same distribution. The normalization process is expressed as Where X is the input feature map; μ and σ 2 are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a small positive number used to avoid division by zero errors;
[0076] Activation function layer: The activation function layer introduces nonlinear factors, allowing the discriminator to learn complex feature representations;
[0077] Fully connected layer: The fully connected layer converts the output of the convolutional layer and the activation function layer into a fixed-length feature vector, which is used for the final classification decision;
[0078] Output layer: The output layer is a single neuron that uses the Sigmoid activation function to output a probability value, which indicates the confidence that the input data is real data. The closer the probability value is to 1, the more real the input data is; the closer the probability value is to 0, the more false the input data is.
[0079] In one embodiment, the physical lighting model calculates lighting effects based on physical principles, and the environment mapping technology uses a pre-made environment map to simulate the background environment, including the following steps:
[0080] Light source definition: define the position, color, intensity and type of light source (such as point light source, parallel light source, etc.);
[0081] Surface material properties: define the reflectivity, refractive index, roughness and other properties of the object surface;
[0082] Lighting calculation: Diffuse reflection: random reflection of light on a rough surface. The formula is I diffuse =k d I s max(0,n·l), where k d is the diffuse reflectance, I s is the intensity of the light source, n is the surface normal, and l is the direction vector from the surface to the light source; Specular reflection: directional reflection of light on a smooth surface, the formula is I pecular =k s I s ·(r·v) d , where k s is the specular reflection coefficient, r is the reflection direction, v is the viewing direction, and d is the specular index; Ambient light: global illumination from the surrounding environment, usually expressed as a constant;
[0083] Final color calculation: Add diffuse reflection, specular reflection and ambient light to get the final color value;
[0084] Environment map production: Use panoramic camera or 3D modeling software to produce environment maps;
[0085] Texture Mapping: Map the environment map onto a sphere or other geometry to simulate the surrounding environment;
[0086] Reflection calculation: During the rendering process, the color value is sampled from the environment map according to the surface properties of the object and the viewing direction as the color of the reflected light.
[0087] In one embodiment, the comprehensive judgment module includes a mathematical model construction unit, an evaluation index calculation unit and a result feedback generation unit. The mathematical model construction unit constructs a suitable mathematical model according to the specific requirements of image processing, which is used to analyze and evaluate image data; the evaluation index calculation unit calculates the evaluation index of the image according to the mathematical model, and the result feedback generation unit generates comprehensive feedback on the image processing effect according to the results of the evaluation index calculation unit.
[0088] In one embodiment, the evaluation metrics include resolution, color fidelity, contrast, frame rate, and processing time;
[0089] The resolution is represented by the dimensions of the image (such as width and height);
[0090] The color fidelity is measured using the peak signal-to-noise ratio, and the formula for the peak signal-to-noise ratio is: Wherein, MAX represents the maximum possible pixel value in the image (for an 8-bit image, usually 255), and MSE represents the mean square error of the image. The formula for the mean square error of the image is: Where I(i, j) represents the pixel value at position (i, j) of the original image, K(i, j) represents the pixel value at the same position of the processed image, M and N represent the width and height of the image respectively;
[0091] The contrast ratio measures the difference between bright and dark parts of an image using a contrast ratio formula, where σ represents the standard deviation of the image pixel values, and μ represents the mean of the image pixel values;
[0092] The frame rate is directly obtained through measurement or statistics. For video processing, the frame rate refers to the number of frames displayed per second;
[0093] The processing time is measured by a timer and represents the time required to process one image or one video.
[0094] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A computer graphics image processing system, characterized in that: include: 3D image scene determination module: automatically identifies and analyzes animation scenes, identifies elements in the scene, and generates a preliminary 3D model based on the scene information; GANs model module: Builds a GANs model, which includes a generator and a discriminator. Adaptive training strategies are used to enable the generator to learn the distribution characteristics of the data. The generator is responsible for generating realistic 3D models, simulating different lighting conditions and background environments based on physical lighting models and environment mapping technology. The discriminator is responsible for distinguishing between real data and generated data. New 3D models are generated by adjusting the input noise vector of the GANs. A specific artistic style is transferred to the generated new 3D models by adjusting the structures and loss functions of the generator and discriminator. Information collection module: collects multi-dimensional information including frame rate, QP value, number of grids, pixel accuracy and rendering time, and monitors various key performance indicators of animation production. The information collection module collects preliminary 3D models and new 3D model data. The multi-dimensional information collected by the information collection module is passed to the data standardization processing module; Data standardization processing module: cleans, converts and standardizes the collected data, and the standardized data is input into the image processing module; Image processing module: uses deep learning to process dynamic effects and image details. The image data processed by the image processing module is passed to the comprehensive judgment module; Comprehensive judgment module: Build a comprehensive mathematical model to evaluate rendering effect, image quality and efficiency; User interaction module: provides users with an operation interface and result display.
2. A computer graphics image processing system according to claim 1, characterized in that: The specific steps of the 3D image scene determination module generating a preliminary 3D model are as follows: Scene capture and preprocessing: Capture animation scenes and preprocess the captured images; Feature extraction: Extract key elements from pre-processed images using image processing and computer vision techniques; 3D modeling: Based on the extracted features, a preliminary 3D model is generated using 3D modeling technology.
3. A computer graphics image processing system according to claim 1, characterized in that: The GANs model module includes a generator, a discriminator, a physical lighting model and environment mapping technology unit, and a style transfer unit. The generator is responsible for generating a three-dimensional model from random noise. The discriminator is responsible for distinguishing whether the input data is real data or false data generated by the generator, and outputting a probability value indicating the confidence that the data is real data. The physical lighting model and environment mapping technology unit are used to simulate different lighting conditions and background environments. The physical lighting model calculates lighting effects based on physical principles. The environment mapping technology uses pre-made environment maps to simulate the background environment. The style transfer unit is responsible for migrating a specific artistic style to the generated three-dimensional model. Style transfer is achieved by adjusting the input noise vector of GANs and the parameters of the generator and discriminator.
4. A computer graphics image processing system according to claim 3, characterized in that: The structure of the generator includes: Input layer: Receives a random noise vector that follows a Gaussian distribution or a uniform distribution. This vector serves as the starting point of the generation process and contains all the potential information required to generate an image or model. Convolutional transpose layer: used to upsample the input vector to the spatial resolution, and map the low-dimensional features to the high-dimensional space through the inverse process of the convolution operation to generate an image or model. The calculation formula is expressed as Y = ConvTranspose(X,W,b,s,p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the step size, which controls the spatial resolution of the output feature map; p is padding, which is used to add additional zero values around the input feature map to control the size of the output feature map; Batch normalization layer: Normalize the output of the convolutional transpose layer so that the input data of each batch has the same distribution. The normalization process is expressed as Where X is the input feature map; μ and σ 2 are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a positive number; Activation function layer: introduces nonlinear factors to enable the generator to learn complex feature representations; Output layer: Generate a 3D model.
5. The computer graphics image processing system according to claim 3, characterized in that: The structure of the discriminator includes: Input layer: receives data from the real dataset or generator as input; Convolution layer: extracts the features of the input data through convolution operation. The convolution kernel slides on the input data, calculates the local features of each position, and generates a feature map. The calculation formula is expressed as Y = Conv(X, W, b, s, p), where X is the input feature map, W is the convolution kernel weight; b is the bias term; s is the step size, which controls the sliding distance of the convolution kernel on the input feature map; p is padding, which is used to add extra zero values around the input feature map to control the size of the output feature map; Batch Normalization Layer: The batch normalization layer normalizes the output of the convolutional layer so that the input data of each batch has the same distribution. The normalization process is expressed as Where X is the input feature map; μ and σ 2 are the mean and variance of the input feature map respectively; γ and β are learnable parameters used to adjust the normalized feature map; ε is a positive number; Activation function layer: The activation function layer introduces nonlinear factors, allowing the discriminator to learn complex feature representations; Fully connected layer: The fully connected layer converts the output of the convolutional layer and the activation function layer into a fixed-length feature vector, which is used for the final classification decision; Output layer: The output layer is a single neuron that uses the Sigmoid activation function to output a probability value, which indicates the confidence that the input data is real data. The closer the probability value is to 1, the more real the input data is; the closer the probability value is to 0, the more false the input data is.
6. A computer graphics image processing system according to claim 3, characterized in that: The physical lighting model calculates lighting effects based on physical principles, and the environment mapping technology uses pre-made environment maps to simulate the background environment, including the following steps: Light source definition: define the position, color, intensity and type of light source; Surface material properties: define the properties of the object surface; Lighting calculation: Diffuse reflection: random reflection of light on a rough surface, the formula is I diffuse =k d I s max(0,n·l), where k d is the diffuse reflectance, I s is the intensity of the light source, n is the surface normal, and l is the direction vector from the surface to the light source; Specular reflection: directional reflection of light on a smooth surface, the formula is I pecular =k s I s ·(r·v) d , where k s is the specular reflection coefficient, r is the reflection direction, v is the viewing direction, and d is the specular index; Ambient light: global illumination from the surrounding environment; Final color calculation: Add diffuse reflection, specular reflection and ambient light to get the final color value; Environment map production: Make environment map; Texture Mapping: Map the environment map onto a sphere or other geometry to simulate the surrounding environment; Reflection calculation: During the rendering process, the color value is sampled from the environment map according to the surface properties of the object and the viewing direction as the color of the reflected light.
7. The computer graphics image processing system according to claim 1, characterized in that: The comprehensive judgment module includes a mathematical model construction unit, an evaluation index calculation unit and a result feedback generation unit. The mathematical model construction unit constructs a mathematical model for analyzing and evaluating image data; the evaluation index calculation unit calculates the image evaluation index from three dimensions of rendering effect, picture quality and efficiency based on the mathematical model; the result feedback generation unit generates comprehensive feedback on the image processing effect based on the results of the evaluation index calculation unit.
8. The computer graphics image processing system according to claim 7, characterized in that: The evaluation metrics include resolution, color fidelity, contrast, frame rate, and processing time; Said resolution is represented by the size of the image; The color fidelity is measured using the peak signal-to-noise ratio, and the formula for the peak signal-to-noise ratio is: Wherein, MAX represents the maximum possible pixel value in the image, MSE represents the mean square error of the image, and the formula of the mean square error of the image is: Where I(i, j) represents the pixel value at position (i, j) of the original image, K(i, j) represents the pixel value at the same position of the processed image, M and N represent the width and height of the image respectively; The contrast ratio measures the difference between bright and dark parts of an image using a contrast ratio formula, where σ represents the standard deviation of the image pixel values, and μ represents the mean of the image pixel values; The frame rate is directly obtained through measurement or statistics. For video processing, the frame rate refers to the number of frames displayed per second; The processing time is measured by a timer and represents the time required to process one image or one video.
Citation Information
Patent Citations
Three-dimensional animation graphic image processing system based on AI artificial intelligence
CN118918224A