Active defense AI self-aiming method based on micro rendering
By injecting adversarial perturbations into the three-dimensional model texture of the game character and using a micro-renderer optimization, the problem of poor defense of AI self-targeting cheating in the prior art is solved, and an efficient, concealed and performance-free defense effect is achieved.
Patent Information
- Application Number
- CN202510639241.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing technology is difficult to effectively defend against AI self-targeting cheating, traditional anti-cheating methods are difficult to detect, and the existing anti-attack methods have problems such as visual interference, poor robustness, insufficient migration and performance overhead.
By injecting carefully designed adversarial perturbations directly into the three-dimensional model texture of the game characters, optimized with a micro-renderer, the character image finally rendered onto the screen can effectively deceive or interfere with the target detection model that the visual self-image depends on.
It realizes effective defense against visual self-image cheating, significantly reduces the detection success rate of AI self-image model, and has the least impact on player visual experience and zero performance overhead.
Smart Images

Figure CN120154899A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine vision technology, and in particular relates to an active defense AI self-aiming method based on differentiable rendering. Background Art
[0002] In recent years, first-person shooter (FPS) games have continued to be popular around the world, with a huge player base. However, as games become more popular, cheating problems have become increasingly serious, especially vision-based aiming robots (referred to as "AI self-aiming" or "plug-ins"). This type of plug-in captures the game screen, uses target detection models in computer vision (such as YOLO, Faster R-CNN, etc.) to automatically identify and lock enemy characters, and then simulates mouse operations to aim and shoot, which seriously undermines the fair competition environment of the game and damages the player experience and the interests of game manufacturers.
[0003] Traditional anti-cheating methods, such as memory scanning to detect the behavior of modifying game memory, or post-analysis based on player behavior data (such as VAC, BattleEye, etc.), are difficult to effectively deal with AI aiming. Because AI aiming usually runs outside the game process, does not directly modify game files or memory, and may simulate human operation modes, making detection very difficult and lagging.
[0004] To address this challenge, researchers began to explore active defense strategies. Adversarial sample technology was introduced into this field as a method to interfere with the judgment of deep learning models. By adding tiny perturbations that are difficult for the human eye to detect to the input image, the target detection model can produce incorrect recognition results. A small number of works have attempted to apply adversarial samples to FPS game anti-cheating, such as superimposing two-dimensional perturbations on the screen or modifying the texture of the scene (such as the wall). However, these methods have some inherent defects: (1) Visual interference: Adding disturbances directly to the screen or background may significantly affect the visual experience of normal players and the aesthetics of the game screen.
[0005] (2) Lack of robustness: Two-dimensional perturbations are sensitive to dynamically changing factors such as lighting, viewing angle, and distance within the game, and the defensive effect is unstable.
[0006] (3) Limited transferability: Perturbations generated for a specific model may be ineffective for other types of visual self-aiming models, and the type of model used by the attacker is unknown.
[0007] (4) Performance overhead: Some methods that require real-time generation or adjustment of disturbances may impose additional computational burden on the client, affecting the game frame rate.
[0008] Therefore, there is an urgent need for a technical solution that can actively, effectively, robustly defend against visual aimbot cheating with the least impact on the player experience. Summary of the Invention
[0009] Object of the Invention: Aiming at the problems in the prior art that the anti-cheating means have poor effect on visual aimbot and the existing adversarial attack methods have visual interference, poor robustness, insufficient transferability, and affect performance, the present invention aims to provide an active defense AI aimbot method based on differentiable rendering.
[0010] The present invention directly injects a carefully designed adversarial perturbation into the three-dimensional model texture of the game character and optimizes it using a differentiable renderer, so that the final character image rendered on the screen can effectively deceive or interfere with the target detection model relied on by visual aimbot, realizing effective defense against visual aimbot cheating.
[0011] The method of the present invention includes the following steps: Step 1, Initialize adversarial perturbation and texture fusion: Obtain the original three-dimensional model mesh of the game character and the corresponding two-dimensional texture map T, where , R is the real number space, and represent the length and width of the texture map respectively; Initialize an optimizable perturbation , , using random values uniformly distributed in the interval for initialization; Superimpose the perturbation onto the two-dimensional texture map T to generate an initial adversarial texture : , where function is used to limit the pixel value within the range; Step 2, Set the differentiable rendering environment and randomize the parameters: Select a differentiable renderer R, and the differentiable renderer R is used to calculate the gradient of the rendered output image with respect to the input parameters (texture); To simulate the variable observation conditions in the game, randomize the rendering parameters, and the rendering parameters include camera parameters and lighting parameters ; Step 3, Perform differentiable rendering and background fusion: Apply the initial adversarial texture generated in Step 1 to the surface of the original three-dimensional model mesh through UV mapping (the process of mapping two-dimensional texture coordinates to the three-dimensional model surface, UV mapping is the core technology for texture mapping in 3D modeling, by unfolding the 3D model surface onto a 2D plane (U and V axes), so that the 2D texture can be accurately projected onto the model surface), and then use the randomly generated camera parameters in Step 2 and lighting parameters , through a differentiable renderer render the model with perturbed texture into a 2D image . During the rendering process, the texture sampling operation calculates the pixel color based on the vertex color affected by the perturbation, thus transferring the perturbation to the rendered image. Randomly select a background image from a pre-prepared background image dataset , and combine the 2D image with the background image to generate a composite image closer to the real game scene ; Step 4, perform image enhancement and multi-model evaluation: Perform image preprocessing operations on the composite image to simulate the processing steps that an attacker may adopt, and obtain the preprocessed image ; Input the image in parallel into a set of pre-selected general object detection models representing the prior art , , denotes the k-th object detection model (mainly using Faster RCNN and YOLO series), and obtain the detection results of each object detection model for the image ; Step 5, calculate the multi-object joint adversarial loss function ; Step 6, backpropagation optimization and perturbation update; Step 7, repeat Steps 2 to 6 for iterative optimization until the loss function converges, and the finally obtained optimized perturbation is added to the 2D texture map T to generate the final texture with the ability to deceive AI auto-aim; Step 8, deployment and application: Replace the original texture image of the corresponding character in the game client with the generated adversarial texture image. When the game engine runs, it will automatically load and use the adversarial texture image to render the character. When the AI auto-aim system observes a character carrying this adversarial texture, its detection performance will be significantly disturbed and it will be difficult to accurately lock and identify the target, thus achieving the purpose of active defense.
[0012] In Step 2, the camera parameters include distance , pitch angle and azimuth angle . The lighting parameters include light source intensity. Randomly generate the following camera parameters at each iteration: , , , where s is the size of the game character model, rand(1) represents a uniformly distributed random number in the interval [0, 1), represents the angle uniformly sampled from the
[0013] In step 3, the generated mask m of the two-dimensional image is used for fusion to obtain the synthesized image : , where the mask m identifies the area of the rendered character in the two-dimensional image : that is, the mask area shows the rendered character, and the rest shows the background; ⊙ represents element-wise multiplication; step 3 aims to let the optimization process consider background interference so that the perturbation is still effective in complex scenarios.
[0014] In step 4, the preprocessing operations include random color jittering (randomly changing brightness, contrast, saturation, hue), cropping (simulating the incomplete situation of the target in the picture), and scaling (simulating the target size at different distances).
[0015] In step 4, the detection result of the k-th object detection model for the image is expressed as: .
[0016] In step 4, the detection result includes the detected object category, confidence score, and bounding box information.
[0017] Step 5 includes: designing a comprehensive loss function to guide the optimization of the perturbation , and the loss function includes the category loss , the total variation loss , and the perturbation amplitude constraint loss : , , , where represents the value of the perturbation at the i-th row, the -th column, and the c-th channel, represents the perturbation The maximum value of the absolute values of all elements in is the confidence score of the k-th model's i-th detection of the target category, and N is the number of detections. is the confidence score of the k-th model's i-th detection of the human category. is the activation function. The category loss aims to reduce the confidence of the surrogate model in the true target category (such as "person") and mislead it to an irrelevant category, making it unable to detect the target.
[0018] The total variation loss is used to penalize the differences between adjacent pixel values in to improve the spatial smoothness of the perturbation and make it look more natural.
[0019] The perturbation magnitude constraint loss limits the maximum absolute value of the perturbation over all pixels and channels, ensuring that each component of the perturbation is not too large, thus more directly controlling the imperceptibility of the perturbation.
[0020] The losses of each part are weighted and combined as: ; Step 6 includes: Using the automatic differentiation function of the deep learning framework, calculate the total loss function with respect to the optimizable perturbation gradient , the calculation of this gradient uses automatic differentiation technology to achieve end-to-end backpropagation, quantifying the sensitivity of the loss function to the final optimization goal (perturbation ). This process specifically includes: The gradient signal starts from the loss , passes through the object detection model , image preprocessing operations and background fusion in sequence, and is transmitted to the rendered image . Using the differentiable renderer R, the gradient in the image space ( ) is backpropagated and mapped to the texture space on the surface of the 3D model to obtain the gradient with respect to the adversarial texture ( ); Finally, through the (differentiable) reverse calculation of the texture overlay operation, the final gradient required to guide the update of the perturbation is obtained ; represents the partial derivative.
[0021] Update the perturbation according to the gradient : , where is the learning rate, is a constant, is the loss function Partial derivative with respect to the perturbation , denotes the perturbation after the (t + 1)-th iteration update, and denotes the perturbation at the t-th iteration.
[0022] Using the above steps, a natural adversarial patch in the physical world can be generated. After applying the natural adversarial patch to the attack target, the target detector can be made to make misjudgments.
[0023] The present invention also provides an electronic device, including a processor and a memory. The memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the method.
[0024] The present invention also provides a storage medium storing a computer program or instruction, and when the computer program or instruction runs on a computer, the steps of the method are executed.
[0025] The present invention has the following beneficial effects: (1) High-efficiency active defense ability: This method directly acts on the input (rendered character image) on which the target detection model, the core of the AI auto-aiming tool, depends. By introducing adversarial perturbations at the texture level, it can effectively interfere with its recognition process before cheating behaviors occur. Experiments show that this method can significantly reduce the detection success rate (DSR) of various AI auto-aiming models (such as YOLO series, RTMDet, etc.), and can even reach more than 99% on some models (such as for YOLOv5s).
[0026] (2) Excellent visual concealment: By strictly restricting the perturbation amplitude ( norm constraint) and using the total variation loss ( ), the smoothness of the perturbation is guaranteed, and the generated adversarial texture is almost indistinguishable from the original texture to the human eye. Subjective perception research and objective image quality evaluation indicators both confirm that the adversarial texture generated by this method has minimal impact on the visual experience of players and does not affect the character appearance recognition and game immersion.
[0027] (3) Strong robustness and generalization ability: By jointly optimizing proxy models with various different architectures, the perturbation patterns learned through adversarial texture are universal and can effectively defend against other vision-based auto-aim models that were not involved in training. Experiments have shown that adversarial textures trained using YOLOv5x and RTMDet also exhibit good defense effects against models such as Faster R-CNN, YOLOv8n, and YOLOv3n. Secondly, during the generation process, by randomizing camera perspectives, distances, lighting conditions, fusing diverse background images, and simulating image preprocessing, the adversarial textures can maintain a high defense success rate in various complex game environments (different maps, open / closed spaces) and under different observation conditions.
[0028] (4) Seamless integration and zero performance overhead: The generation process of this method is completely completed offline. During game runtime, the rendering engine only needs to load the replaced adversarial texture without any additional real-time calculations. Therefore, this method has no negative impact on the game's frame rate (FPS) and client performance, ensuring a smooth gaming experience, and is particularly suitable for FPS games with extremely high real-time requirements.
[0029] (5) Easy to deploy and expand: For game developers, deploying the adversarial textures generated by this method only requires replacing the texture files in the game resource package, which is simple to operate. This method can be applied to models of different characters and factions in the game and has good scalability.
[0030] In summary, the present invention provides a novel, efficient, stealthy, robust, and performance-loss-free active defense solution, effectively addressing the deficiencies of existing technologies in dealing with vision-based auto-aim cheating and providing strong technical support for maintaining a fair competitive environment in FPS games. Brief Description of the Drawings
[0031] Figure 1 is a schematic diagram of the AI auto-aim workflow.
[0032] Figure 2 is a schematic diagram of the relationship between game character textures and AI auto-aim.
[0033] Figure 3 is the overall flowchart of the present invention.
[0034] Figure 4 is a qualitative comparison diagram of the defense effect of this method.
[0035] Figure 5 is a curve graph showing the influence of the learning rate on the defense success rate (DSR) of this method.
[0036] Figure 6 is a curve graph showing the influence of the perturbation amplitude on the defense success rate (DSR) of this method. Detailed implementation manners
[0037] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.
[0038] An embodiment of the present invention provides an active defense AI self-aiming method based on differentiable rendering and demonstrates its application effect in a real game scenario. Although this method has wide applicability and can be applied to various first-person shooting (FPS) games, the present invention will take a certain first-person shooting game as a specific example for illustration.
[0039] Figure 1 The implementation scenario of the present invention is shown, which is a typical AI self-aiming work process. The core of AI self-aiming is a target detection model. Therefore, the fundamental strategy for defending against AI self-aiming lies in interfering with or deceiving this target detection model. The conventional AI self-aiming process is as follows: First, a real-time game screen is obtained through a screen capture software (such as common OBS, etc.); subsequently, this screen is input into the target detection model to identify and locate enemy characters; finally, the operating system-provided API (such as Windows API) is used to control the mouse, precisely move it to the target position and trigger a shot. The entire process can usually be completed within 10 milliseconds, far exceeding the reaction speed of ordinary players, which is the reason why AI self-aiming poses a serious threat.
[0040] Figure 2 explains how in-game texture images are presented from the player's perspective. The visual presentation of game characters mainly depends on three-dimensional meshes and textures. The mesh defines the basic geometric shape of the model, while the texture endows the model surface with rich details and colors. Essentially, visual AI self-aiming works by identifying the texture features of the character model. In order to endow the model surface with detailed images, UV unwrapping is required during the character creation process, that is, mapping the three-dimensional model surface to a two-dimensional UV coordinate space. Then, art designers draw details such as skin and clothing on this two-dimensional texture image. As Figure 2 shown, these two-dimensional texture maps are "fitted" to the three-dimensional model surface through UV mapping technology, that is, the pixels (texels) on the texture image establish a correspondence with the model vertices. Finally, the three-dimensional model with textures is processed through the rendering pipeline, combined with parameters such as the model position, lighting conditions, and camera perspective, and projected onto a two-dimensional screen to form the image that the player finally sees. This means that the pixel colors and details displayed on the screen directly come from the texture information on the model surface. Visual AI self-aiming captures the screen image and uses the target detection model to analyze this texture information to further identify enemy characters.
[0041] The following combines Figure 3 to specifically describe the specific steps of the method proposed by the present invention in detail: First, regarding the dataset, obtain the 3D model (mesh M) of the game character and its corresponding original texture map T. For example, use a dataset (CAT) containing 66 models of character 1 and a dataset (TAT) containing 73 models of character 2. Then, capture high-quality screenshots of different scenes from multiple game maps (such as Dust2, Mirage, Anubis, etc.), crop them to a unified resolution (e.g., 640x640 pixels), and construct a background dataset containing approximately 1800 images ( ).
[0042] Step 1, for the selected original texture of the character , initialize a perturbation map of the same size as , and fill its values with uniform random numbers within the range of . Generate the initial adversarial texture , ensuring that the texture values are within the valid range.
[0043] Step 2, randomly generate the following camera parameters at each iteration: , , , where s is the size of the game character model, rand(1) represents a uniform distribution random number in the interval [0, 1), represents an angle uniformly sampled from the interval. To simulate the changing perspectives in the game. The lighting parameter mainly refers to the lighting intensity.
[0044] Step 3, apply the current adversarial texture to the 3D model M. Using the random parameters generated in Step 2, render the foreground image containing the adversarial model through a differentiable renderer implemented using the Pytorch3D framework. At the same time, generate a mask according to (with pixel values of 1 where non-zero and 0 where zero). Randomly select a background image from the background dataset, and use the mask to fuse with to obtain the blended image .
[0045] Step 4, for the fused image Apply data augmentation operations such as random scaling and color jitter to obtain the final image input to the model . Feed into the proxy models YOLOv5x and RTMDet to obtain the detection results , .
[0046] Step 5, calculate the joint adversarial loss , where and are hyperparameters, set to 0.2 and 0.05 by experimental results. The class loss , is the number of models, is the number of detection boxes, is the confidence of human detection, is the confidence of non-human detection, used to reduce the model's detection confidence for the "human" category. The total variation loss calculates the total variation of the perturbation to control the smooth change of the perturbation. The perturbation amplitude constraint loss is used to control the maximum value of the perturbation.
[0047] Step 6, optimization and update. Use the Adam optimizer to update according to . According to the ablation study results, the learning rate η = 0.005 gives better results. After the update, clip : , and keep the perturbation in the black area zero.
[0048] Step 7, repeat Steps 2 - 6 for a fixed number of iterations (2000 times) until the loss converges. Obtain the final optimized perturbation . Generate the final adversarial texture .
[0049] Next, combined with Figure 4 , Figure 5 , Figure 6 , demonstrate the effectiveness of the method of the present invention in experimental and real game scenarios: Figure 4 Intuitively shows the comparison of the detection effects of AI aimbot before and after applying the method. Without using the method, the AI aimbot can accurately identify game characters, thus achieving cheating. After generating the adversarial texture using the method, the AI aimbot will misidentify the target as other categories (it only attacks targets identified as "human") or completely fail to detect the target, thus effectively defending against the AI aimbot cheat.
[0050] Figure 5Shows the influence of different learning rates on the final generated adversarial texture effect. The experimental results show that when the learning rate is set around 0.005, the adversarial texture obtained by training has the best defense effect against AI aimbot.
[0051] Figure 6 Discussed the influence of the perturbation magnitude ( ), on the attack success rate (i.e., the defense effect). The results show that the larger the perturbation magnitude, the stronger the adversarial property and the better the defense effect, but the more obvious the visual change of the texture. To achieve a balance between the defense effect and the player's visual experience, the present invention selects as the final setting. Under this setting, the visual impact on the player caused by the perturbation is minimized, while still maintaining excellent defense performance against AI aimbot.
[0052] To further evaluate the performance of the present invention, Table 1 in the appendix compares it with two benchmark methods: random perturbation (adding unoptimized noise to the texture) and AdvMap (adding perturbations to the screen image). Two key metrics, the defense success rate (DSR) and the structural similarity (SSIM), are used for the evaluation. DSR quantifies the degree of failure of the AI aimbot (i.e., the proportion of frames in which the AI aimbot fails to recognize the target as a "person"), and a higher value represents more effective defense; SSIM measures the visual similarity between the adversarial texture and the original texture, and the closer it is to 1, the smaller the impact on the player's visual experience. The experimental results (tested on multiple AI aimbot cheating models for the CAT and TAT datasets) clearly show that the method of the present invention is significantly superior to the two benchmark methods in terms of DSR. At the same time, its SSIM score is extremely close to 1, proving that this method can achieve a powerful defense ability against AI aimbot with almost no impact on the original visual effect of the game.
[0053] Table 1
[0054]
[0055] The present invention provides an active defense method against AI aimbot based on differentiable rendering. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.
Claims
1. An active defense AI self-aiming method based on differentiable rendering, characterized in that: The following steps are involved: Step 1: Initialize adversarial perturbation and texture fusion: Get the original 3D model mesh of the game character and the corresponding two-dimensional texture map T, where , R is the real number space, and Represent the length and width of the texture map respectively; initialize an optimizable perturbation with the same dimension as the two-dimensional texture map T , , used in Initialize with a random value uniformly distributed in the interval; Superimposed on the two-dimensional texture map T to generate the initial adversarial texture : ,in The function is used to limit the pixel value to within the scope; Step 2, set up a differentiable rendering environment and perform parameter randomization: select a differentiable renderer R, which is used to calculate the gradient of the rendered output image relative to the input parameters; randomize the rendering parameters, which include camera parameters and lighting parameters ; Step 3: Perform differentiable rendering and background fusion: The initial adversarial texture generated in step 1 Apply to the original 3D model mesh via UV mapping Then use the randomly generated camera parameters in step 2 and lighting parameters , through a differentiable renderer Rendering a model with a perturbed texture into a 2D image During rendering, the texture sampling operation calculates the pixel color based on the vertex color affected by the disturbance, thereby Passed to the rendered image, a background image is randomly selected from the pre-prepared background image dataset , the two-dimensional image With background image Fusion to generate a synthetic image that is closer to the real game screen ; Step 4: Image enhancement and multi-model evaluation: Perform image preprocessing operations to obtain the preprocessed image ; The image Input to a set of object detection models in parallel middle, , Represents the kth target detection model, and obtains each target detection model for the image The test results; Step 5: Calculate the multi-objective joint adversarial loss function ; Step 6: Back propagation optimization and perturbation update; Step 7: Repeat steps 2 to 6 for iterative optimization until the loss function Convergence, the final optimized perturbation It is added to the 2D texture map T to generate the final texture that can deceive AI self-aiming; Step 8, deploy the application: replace the original texture image of the corresponding character in the game client with the generated adversarial texture image. The game engine will automatically load and use the adversarial texture image to render the character during runtime.
2. The method according to claim 1, characterized in that In step 2, the camera parameters Including distance, pitch angle and azimuth, lighting parameters Including light intensity.
3. The method according to claim 2, characterized in that In step 3, use the two-dimensional image The generated mask Fusion to obtain a composite image : , The mask m identifies the two-dimensional image The area where the character is rendered in ; ⊙ represents element-wise multiplication.
4. The method according to claim 3, characterized in that In step 4, the preprocessing operations include random color jittering, cropping, and scaling.
5. The method according to claim 4, characterized in that In step 4, the kth target detection model For images Test results It is expressed as: 。 6. The method according to claim 5, characterized in that In step 4, the detection result includes the detected target category, confidence score and bounding box information.
7. The method according to claim 6, characterized in that Step 5 includes: designing a comprehensive loss function , used to guide the perturbation The loss function is Including category loss , total variational loss and the perturbation amplitude constraint loss : , , , in Represents disturbance In the Row, No. Column, No. The value at the channel, Represents disturbance The maximum absolute value of all elements in ; is the confidence score of the k-th model for the i-th detection of the target category, N is the number of detections, is the confidence score of the k-th model for the i-th detection of the human category, is the activation function; The weighted combination of each part of the loss is: , in and is the weight parameter.
8. The method according to claim 7, characterized in that Step 6 includes: using the automatic differentiation function of the deep learning framework to calculate the total loss function Relative to the optimizable disturbance Gradient , the gradient signal is derived from the loss Start by passing through the target detection model , image preprocessing operations and background fusion, passed to the rendered image , using the differentiable renderer R, the gradient of the image space Back propagation and mapping to the texture space of the 3D model surface to obtain the adversarial texture Gradient Finally, the guidance perturbation is obtained through the reverse calculation of the texture superposition operation Update the final gradient required ; represents partial differential; Update the perturbation according to the gradient : , in is the learning rate, is a constant, represents the disturbance after the t+1th iteration update, represents the perturbation at the tth iteration.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 8.
10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.
Citation Information
Patent Citations
Game picture generation method and device, equipment and storage medium
CN112652046A
Method and device for generating confrontation chartlet for resisting AI self-aiming cheating
CN116993893A
Anti-cheating method for instant shooting online game
CN118416493A
Physical confrontation coating generation method and device for vehicle target detector
CN119540158A
Anti-peek system for video games
US11623145B1
Cited By
Transferable anti-texture coating generation method and system
CN121600150A