Physical confrontation attack patch generation method and device for night monitoring imaging system
By designing a physical adversarial attack method based on reflective tape, the problem of insufficient robustness of near-infrared images in the night monitoring system is solved, and a high success rate attack in real physical scenarios is achieved.
Patent Information
- Application Number
- CN202510056361.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for night surveillance camera systems to obtain clear visible light images under extremely poor lighting conditions, and the near-infrared spectral images are less robust and are susceptible to image confrontation attacks, which poses safety risks.
A physical confrontation attack method is proposed. By setting up a new physical mapping scheme, an attack method based on reflective tape is designed, and the materials printed with the masking paper and reflective tape are combined to achieve confrontation attacks, which improves the practicality and success rate of attacks.
It effectively improves the success rate of physical attacks, makes its attack effect significant in real physical scenarios, and solves the problem of insufficient robustness of near-infrared images in night monitoring systems.
Smart Images

Figure CN119995948A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and counterattack, and in particular relates to a physical counterattack method for a nighttime monitoring camera system. Background Art
[0002] Nighttime surveillance camera systems are widely used in the security field. During the day, infrared filters are typically used to filter out infrared components and capture conventional RGB images. However, in extremely poor lighting conditions (such as at midnight), obtaining clear visible light images becomes very difficult due to the limited sensitivity of the camera. To address this issue, near-infrared (NIR) LEDs are often used as auxiliary lighting. Their central wavelength is typically around 850nm, which is within the spectral response range of silicon-based image sensors. Since NIR light is not perceptible to the human eye, it helps reduce light pollution and can extend the monitoring range without being conspicuous.
[0003] Despite the widespread adoption of NIR capture technology, the potential security vulnerabilities associated with it have received relatively little in-depth research. While attempts have been made to obscure faces and evade surveillance cameras by wearing electronic devices that emit near-infrared light, few studies have explored the potential threats posed by NIR image-based detectors. Given the widespread application of NIR technology in security, particularly surveillance systems, vulnerabilities in NIR imaging could present opportunities for attackers, posing significant security risks.
[0004] The imaging principles of NIR nighttime surveillance camera systems and the installation of their auxiliary lighting devices reveal inherent vulnerabilities that make NIR-based AI algorithms less robust than RGB-based algorithms. Compared to visible light, NIR LED lighting near 850nm exhibits a significant drawback: color and texture details of objects in the scene captured by the camera are easily lost. This phenomenon is particularly pronounced with dyed fabrics, especially the texture of human clothing. This is due to two fundamental principles of near-infrared imaging. First, the spectral sensitivities of the three color filters of silicon-based cameras tend to overlap in the near-infrared band (specifically, the wavelength range of 850nm to 1000nm). As a result, although the sensor can capture color images, it only produces monochrome images in the NIR band, preventing the reproduction of vibrant colors. Second, the reflectance spectra of different materials in the near-infrared band also tend to be consistent, resulting in a significant loss of texture information in the image. Consequently, the intensity variation range of near-infrared images is very limited, and AI models trained on these data often lack robustness. This means that any changes in the intensity distribution of the input image directly expose the AI model's vulnerabilities to image adversarial attacks.
[0005] To save installation space and reduce obstruction, NIR LEDs are typically installed around the camera lens, placing the light source and camera in almost the same geometric position. While this approach simplifies system design in terms of hardware, it also creates potential for physical attacks. Summary of the Invention
[0006] To address the aforementioned issues, this paper proposes a novel physical countermeasure method, establishes a reasonable near-infrared dataset for nighttime surveillance, and specifically analyzes the feasibility and concealment of physical attacks. Leveraging the imaging principles of nighttime camera systems, a new physical mapping scheme is developed to rationally map the size, angle, and quantity of digital images. A targeted attack method based on reflective tape is designed. A novel material combination of printed masking paper and reflective tape is employed to implement a physical countermeasure attack, improving the practicality of the physical attack and effectively reproducing the digital attack, with no significant reduction in the success rate of attacks in real physical scenarios.
[0007] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0008] The present invention proposes a physical attack method for a nighttime monitoring system, which specifically includes the following steps:
[0009] S1: First, a near-infrared spectral image I∈R is defined 3×H×W , where H and W represent the height and width of the image respectively; using the pre-trained human detector F, we can obtain the bounding box position Y pos and the corresponding confidence Y conf The output label Y is expressed as:
[0010] Y=F(I)=[Y pos , Y conf ]. (1)
[0011] Our goal is to develop a method to minimize the confidence Y of the human detector conf , thereby deceiving the detector so that it cannot detect human bodies or other objects in the image, expressed as:
[0012] argminY conf (2)
[0013] Design of attack algorithm In order to achieve the above goals, we designed an attack algorithm f atk, the algorithm accepts a near-infrared image I as input and outputs an attacked image I that minimizes the confidence of the human detector, expressed as:
[0014] argminF(f atk (I)). (3)
[0015] S2: Define patch coverage model We define the patch M∈{0, 1} h×w , used to describe the shape and position of the reflective patch on the target object;
[0016] The patch coverage process can be expressed as:
[0017]
[0018] Where ⊙ represents the Hadamard product, M is the mask matrix, and x nir ∈I h×w is the original cover image, and is the generated attack image, whose value is obtained by taking a picture in the physical world by a near-infrared camera.
[0019] 1. According to the method for describing the physical adversarial attack problem in claim 1, the patch generation scheme comprises the following steps:
[0020] S3: Define a series of anchor points on the image as the basic units for shape construction. By adjusting the positions and coordinates of the anchor points, the contour of the patch can be flexibly changed.
[0021] We use centripetal Catmull-Rom spline curves to connect anchor points. To ensure the natural transition and smoothness of the patch shape, we use centripetal Catmull-Rom spline curves to connect anchor points. This method not only avoids the generation of circular or self-intersecting curves, but also ensures that there are no inflection points in the curve segment. i and P i+1 The curve segment C between i , we use four anchor points P i-1 , P i , P i+1 , P i+2 And the centripetal Catmull-Rom spline function CCRS. The generation of the curve segment Ci can be expressed by the mathematical formula:
[0022] C i =CCRS(P i-1 , P i , P i+1 , P i+2 ) (5)
[0023] S4: Representation of patch contours when all n curve segments C i When connected, the contour of the patch can be represented as a mask matrix M∈{0, 1} h×w , which can be expressed mathematically as:
[0024] M={C i |0≤i≤n-1} (6)
[0025] The mask matrix M defines the shape and position of the patch, indicating which areas in the image are covered by the patch;
[0026] By setting the brightness value of the mask position area to pure white to simulate the reflected white light in the real world, the definition of the coverage mask can be described by the mathematical formula:
[0027]
[0028] S5: During the shape optimization process, there may be undesirable situations such as boundary line crossing, anchor points stuck in narrow positions, or anchor points located outside the human body area. To solve these problems, we define the feasible area of the anchor points; Figure 2 (a) is the normal anchor point area, and the other three figures are illegal situations that should be restricted. We construct the feasible area B of the anchor point P through two adjacent equidistant lines (the edge of the inner circle and the boundary of the outer circle). This is an acute-angle fan-shaped area. The position of the anchor point P is restricted by the following formula:
[0029]
[0030] And ensure that the anchor point P is located within the acute angle fan-shaped area B surrounded by two adjacent straight lines;
[0031] In order to prevent the curve from crossing, we use the straight line function to determine whether the anchor point P is in its feasible area. Only when the product of the two is less than 0, the anchor point P j It is located inside the region. It can be expressed as follows:
[0032]
[0033] To prevent the anchor point from falling into a narrow position, we also need to set an inner circle with a radius of r within the initial circle and restrict the anchor point from moving into the inner circle:
[0034]
[0035] To ensure that the patch is within the valid area of the human body, we define the bounding box area of the target detector and use it as the outer boundary of the patch generation to ensure that the generated patch is not outside the valid area of the human body:
[0036] O∈{(xn ,y n )|x l <x<x r ,y d <y<y u} (11)
[0037] S6: Through the patch optimization process described above, we set the initial population for adversarial patch generation. This population consists of a series of carefully designed circular patches, laying the foundation for the subsequent evolution process.
[0038] S7: Building on this, we implemented mutation and crossover operations, and considered boundary handling strategies, to generate new generations of subpopulations. By continuously performing these operations on the seeds, the patches in the population achieved a higher level of balance and adversarial resistance. We also employed a black-box query-based attack strategy, which analyzes the confidence scores output by the object detector and uses this as a basis for fitness optimization.
[0039] Figure 3 is the complete framework of this process, defining the population as a set of anchor points {P j |j=1,2,...n} represents. Given a population size Q, the fitness function of the kth generation is defined as:
[0040]
[0041] Among them S i (k) is the shape of the i-th patch, S ij (k) is the kth generation S i (k) the j-th anchor point. and Together they constitute the feasible region B j , that is, the moving range of the jth anchor point in each patch shape, in the k+1th iteration, based on S(k), the solution S is obtained through crossover, mutation and selection i (k+1), for S i (k) Apply the fitness function to evaluate the attack effect. The fitness function only uses the confidence score of the target detector; i The fitness score of (k) is expressed as J(s i ). In order to evaluate J(s i ) value. Fitness function J(s i ) can be expressed as:
[0042]
[0043] Where λ is a weighting factor, Reflects the progress of the current attack success (the larger the value, the closer to success). It can be formalized as:
[0044]
[0045] f nir (x nir ) is the confidence score of pedestrian detector in near-infrared, The confidence score of the image covered with the adversarial patch in the near-infrared pedestrian detector;
[0046] S8: After multiple iterations, the optimal solution of a single theory is locked and selected, and the corresponding mask and digital mask image are retained.
[0047] 2. The patch generation scheme described in claim 2, wherein the actual physical implementation scheme comprises the following steps:
[0048] S9: Using a printer, print the masked image on white paper in the appropriate proportions of the physical world;
[0049] S10: Cut out the mask area and fix the reflective tape on the back with tape to ensure the reusability of the reflective tape;
[0050] S11: The patch is attached to human clothing or other attackable targets. The local light intensity of the camera is controlled by high exposure value on the front or shape optimization on the side to achieve a physical adversarial attack, successfully preventing the camera from recognizing the target object.
[0051] The present invention has at least the following beneficial effects
[0052] Compared to existing physical adversarial attacks, the method presented in this paper significantly improves the difficulty of mapping the digital world to the physical world and increases the success rate of attacks. By exploring multi-shape optimization in a large search space, utilizing a rational genetic selection evolutionary algorithm, coupled with a rational multi-patch and size scaling strategy, and leveraging the ability of reflective tape to accurately fit real physical patches, this method provides a new and effective approach to physical adversarial attacks against nighttime surveillance systems.
[0053] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0055] Figure 1 This is a framework diagram of the method of the present invention.
[0056] Figure 2 Optimize rules for patch shapes.
[0057] Figure 3 Implementation flowchart for digital iteration. DETAILED DESCRIPTION
[0058] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0059] See also Figure 1 、 Figure 2 and Figure 3 The present invention provides a physical counter-attack method for a nighttime monitoring system, which specifically includes the following steps:
[0060] S1: Identify the human detector in the nighttime surveillance camera system as the attack target, and the goal is to deceive the detector so that it cannot detect human beings in the image;
[0061] S2: Define a series of anchor points on the image as the basic units for constructing patch shapes;
[0062] S3: Change the contour of the patch by moving the anchor points to ensure the naturalness and smoothness of the shape;
[0063] S4: Use centripetal Catmull-Rom spline curves to connect anchor points and create a smooth patch outline. Connect all curve segments and generate anchor points P. i and P i+1 The curve segment C between i , we use four anchor points P i-1 ,P i ,P i+1 ,P i+2 The centripetal Catmull-Rom spline function (CCRS) is used to form a patch shape. At the same time, it is necessary to avoid boundary line crossing, anchor points stuck in narrow positions, or anchor points outside the human body area. The feasible area and rules of the anchor points are defined. The movement range of the anchor points is restricted by setting the boundaries of the inner and outer circles. Finally, the straight line function is used to divide the area to determine whether the anchor point is within the feasible area.
[0064] S5: Set the normalized mask matrix. You need to define the position and number of patches. These patches each have independent anchor coordinates and shape parameters. They are initialized within the preset center point. The setting method is to evenly divide the detection box into different areas and obtain the center point from each area.
[0065] S6: Set the brightness value of the mask position area to pure white to simulate the physical reflection effect;
[0066] S7: Ensure that the patch is within the valid area and define the bounding box area of the target detector as the outer boundary of the patch generation;
[0067] S8: Create a series of circular patches as the initial random population. Then generate a new generation of subpopulations through mutation, crossover, and selection operations. Select the optimal anchor point location area in each round and cross-combine it with other points to further optimize the shape of the patch.
[0068] S10: Use the confidence score of the target detector as the fitness function to evaluate the attack effect of the patch;
[0069] S11: Lock the optimal solution in multiple iterations, output and retain the corresponding digital mask and digital image;
[0070] S12: Print the mask image on white paper, cut out the mask area, and apply reflective tape to the back. Attach the patch to clothing or other objects. Ensure the quality and reflectivity of the reflective tape. The patch should be firmly attached to prevent it from falling off during movement. Adjust the patch's position and shape based on camera angle and lighting conditions.
[0071] S13: While ensuring the patch is within the camera's field of view, the camera's local light intensity is controlled by using a high exposure value or optimizing the side shape. This attack prevents the camera from recognizing the target object.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. Physical adversarial attack method for nighttime surveillance camera system, the method mainly includes the following steps: S1: First, a near-infrared spectral image I∈R is defined 3×H×W , where H and W represent the height and width of the image respectively; using the pre-trained human detector F, we are able to obtain the bounding box position Y pos and the corresponding confidence Y conf The output label Y is expressed as: Y=F(I)=[Y pos ,Y conf ]. (1) Our goal is to develop a method to minimize the confidence Y of the human detector conf , thereby deceiving the detector so that it cannot detect human bodies or other objects in the image, expressed as: argminY conf . (2) Design of attack algorithm To achieve the above goals, we designed an attack algorithm f atk , the algorithm accepts a near-infrared image I as input and outputs an attacked image I that minimizes the confidence of the human detector, expressed as: argminF(f atk (I)). (3) S2: Define patch coverage model We define a patch M∈{0, 1} h×w , used to describe the shape and position of the reflective patch on the target object; The patch coverage process can be expressed as: where ⊙ represents the Hadamard product, M is the mask matrix, and x nir ∈I h×w is the original cover image, and is the generated attack image, whose value is obtained by taking a picture in the physical world by a near-infrared camera.
2. According to the method for describing a physical adversarial attack problem according to claim 1, the patch generation scheme comprises the following steps: S3: Define a series of anchor points on the image as the basic units for shape construction. The contour of the patch can be flexibly changed by adjusting the position and coordinates of the anchor points. We use centripetal Catmull-Rom spline curves to connect anchor points. To ensure the natural transition and smoothness of the patch shape, we use centripetal Catmull-Rom spline curves to connect anchor points. This method not only avoids the generation of circular or self-intersecting curves, but also ensures that there are no inflection points in the curve segment. i and P i+1 The curve segment C between i , we use four anchor points P i-1 ,P i ,P i+1 ,P i+2 And the centripetal Catmull-Rom spline function CCRS. The generation of the curve segment Ci can be expressed by a mathematical formula: C i =CCRS(P i-1 ,P i ,P i+1 ,P i+2 ) (5) S4: Representation of patch contours When all n curve segments C i When connected, the contour of the patch can be represented as a mask matrix M∈{0,1} h×w , which can be expressed mathematically as: M={C i |0≤i≤n-1} (6) The mask matrix M defines the shape and position of the patch, indicating which areas of the image are covered by the patch; By setting the brightness value of the mask position area to pure white to simulate the reflected white light in the real world, the definition of the coverage mask can be described by a mathematical formula: S5: In the process of shape optimization, there may be undesirable situations such as boundary line crossing, anchor points stuck in narrow positions or anchor points outside the human body area. In order to solve these problems, we define the feasible area of the anchor point; Figure 2 (a) is the normal anchor point area, and the other three figures are illegal situations that should be restricted. We construct the feasible area B of the anchor point P through two adjacent equidistant lines (the edge of the inner circle and the boundary of the outer circle). This is an acute-angle fan-shaped area. The position of the anchor point is restricted by the following formula: And ensure that the anchor point P is located within the acute-angle fan-shaped area B enclosed by two adjacent straight lines; In order to prevent the curve from crossing, we use the straight line function to determine whether the anchor point P is in its feasible area. Only when the product of the two is less than 0, the anchor point P j It is located inside the region. It can be expressed as: In order to prevent the anchor point from falling into a narrow position, we also need to set an inner circle with a radius of r inside the initial circle and restrict the anchor point from moving into the inner circle: In order to ensure that the patch is within the valid area of the human body, we define the bounding box area of the target detector as the outer boundary of the patch generation to ensure that the generated patch is not outside the valid area of the human body: O∈{(x n ,and n )|x l <x<x r ,and d <and<and u } (11) S6: Through the above patch optimization process, we set the initial population for adversarial patch generation. This population consists of a series of carefully designed circular patches, laying the foundation for the subsequent evolution process; S7: Based on this, we implemented mutation and crossover operations and considered boundary handling strategies to generate new generations of subpopulations. Let the seed continue to perform the above operations so that the patches in the population can reach a higher level of balance and confrontation. A black-box query-based attack strategy is adopted, which analyzes the output confidence scores of the target detector and optimizes the fitness based on them; Figure 3 is a complete framework of this process, defining the population as a set of anchor points {P j |j=1,2,...n}. Given a population size Q, the fitness function of the kth generation is defined as: Where S i (k) is the shape of the i-th patch, S ij (k) is the kth generation S i (k) is the j-th anchor point. and Together they constitute the feasible region B j , that is, the moving range of the jth anchor point in each patch shape. In the k+1th iteration, based on S(k), the solution S is obtained through crossover, mutation and selection. i (k+1), for S i (k) Apply the fitness function to evaluate the attack effect. The fitness function only uses the confidence score of the target detector. i The fitness score of (k) is expressed as J(s i ). In order to evaluate J(s i ) value. The fitness function J(s i ) can be expressed as: where λ is a weighting factor, Reflects the progress of the current attack success (the larger the value, the closer to success). It can be formalized as: f nir (x nir ) is the confidence score in the near-infrared pedestrian detector, The confidence score of the image covered with the adversarial patch in the near-infrared pedestrian detector; S8: Lock and select the optimal solution in multiple iterations, and keep the corresponding mask and digital cover image.
3. The patch generation scheme described in claim 2, wherein the actual physical implementation scheme comprises the following steps S9: Using a printer, print the masked image on white paper in an appropriate proportion to the physical world; S10: Cut out the mask area, and fix the reflective tape on the back with tape to ensure the reusability of the reflective tape; S11: Paste the patch on human clothing or other attackable targets, and control the local light intensity of the camera through high exposure value on the front or shape optimization on the side to achieve physical adversarial attack, successfully making the camera unable to recognize the target object.
Citation Information
Patent Citations
Anti-attack method and system for cheating thermal infrared detector by using cold and hot patches
CN115906090A
Cross-modal adversarial sample generation method in physical environment
CN116596052A
Physically realizable face depth image confrontation sample generation method and system
CN118116046A