Image processing method, related apparatus, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING REALAI TECH CO LTD
- Filing Date
- 2023-06-20
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]在相关技术中,对于信息的干扰主要是通过在图像中添加干扰图像来实现的,然而,相关技术中添加干扰图像的方式较为简单,难以模拟实际场景中对象被遮挡的情况,因此对抗扰动的干扰作用较为失真,难以体现出模型对于实际场景中被遮挡对象的识别能力
[0046]由上述技术方案可以看出,在获取包括待检测对象的待处理图像后,根据待检测对象对应的第一位置即可以确定待检测对象对应的第二位置,在根据所述第二位置的指示将初始对抗扰动添加到所述待处理图像中的待检测对象上后,一旦添加了对抗扰动的待处理图像被采集,就会形成对抗样本(即将数字世界的对抗扰动转换至真实的物理世界中),并作为对象检测模型的输入。一方面,由于第二位置较为贴合在实际场景中待检测对象所对应的实际空间位置,当待检测对象在三维空间中的位置发生变动时,该第二位置可以相应的进行改变,进而在基于第二位置添加对抗扰动时,该对抗扰动在图像中的位置贴合待检测对象的实际三维空间位置,当待检测对象在三维空间中的位置发生变动时,该对抗扰动能够准确跟随待检测对象在三维空间中的位置变化而变化,因此具备时间一致特性,不容易被检测出,隐蔽性较强;另一方面,由于本申请是基于第二位置添加对抗扰动,而第二位置是基于三维空间位置确定的,因此在添加对抗扰动的过程中可以结合待检测对象的三维空间位置进行空间形变、角度改变等操作,使对抗扰动更加贴合待检测对象在三维空间中的姿态,并基于待检测对象在三维空间中的姿态变化而变化,从而能够在待处理图像中更加隐秘的与待检测对象进行贴合,具有空间上的一致性。由此可见,本申请添加对抗扰动的方式能够使对抗扰动具有较强的时间一致性和空间一致性,从而使对抗扰动与待检测对象的贴合度更高,扰动隐秘性更强,因此模型设计者可以通过该方式处理后的图像更方便地在数字世界仿真测试对象检测模型在面临物理的贴图对抗攻击时的鲁棒性表现,同时使处理后的图像能够在对抗攻击中作为对抗样本对模型具有较强的攻击性,此外,本申请的图像添加方式能够较为高效的产出高质量的对抗样本,对于对抗攻击技术领域具有较为显著的优化效果,有助于加速模型的迭代速度,缩短模型的迭代周期。
Smart Images

Figure CN116721317B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method, related apparatus, storage medium, and program product. Background Technology
[0002] To accurately identify various objects in image and video information, multiple models exist for object recognition in related technologies. During model training or other scenarios, to verify the model's object recognition capabilities, some distorted image or video information can be input into the model to verify whether it can accurately identify objects in this challenging information.
[0003] In related technologies, information interference is mainly achieved by adding interfering images to the image. However, the method of adding interfering images in related technologies is relatively simple and it is difficult to simulate the situation where objects are occluded in real-world scenes. Therefore, the interference effect against disturbances is relatively distorted and it is difficult to reflect the model's ability to recognize occluded objects in real-world scenes. Summary of the Invention
[0004] This application provides an image processing method, related apparatus, storage medium, and program product. The adversarial perturbations generated by this method are more closely related to object occlusion in real-world scenarios.
[0005] The embodiments of this application disclose the following technical solutions:
[0006] In a first aspect, embodiments of this application disclose an image processing method, the method comprising:
[0007] Acquire an image to be processed, wherein the image to be processed includes an object to be detected;
[0008] Based on the first position corresponding to the object to be detected, the second position corresponding to the object to be detected is determined, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed;
[0009] Based on the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed.
[0010] Secondly, embodiments of this application disclose an image processing apparatus, the apparatus comprising an acquisition unit and a processing unit:
[0011] The acquisition unit is used to acquire an image to be processed, wherein the image to be processed includes an object to be detected;
[0012] The processing unit is configured to determine a second position corresponding to the object to be detected based on a first position corresponding to the object to be detected, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed.
[0013] Based on the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed.
[0014] In one possible implementation, the processing unit is specifically used for:
[0015] Based on the first position, a projection transformation is performed on the target plane to determine the second position corresponding to the object to be detected. The target plane is the two-dimensional plane of the calibration camera corresponding to the device that generates the image to be processed.
[0016] In one possible implementation, the processing unit is specifically used for:
[0017] Determine the target image corresponding to the object to be detected in the image to be processed, wherein the target image is the visible surface of the object to be detected in the image to be processed;
[0018] The second position is determined by projecting the object to be detected onto the target plane based on its actual three-dimensional spatial position corresponding to the target image.
[0019] In one possible implementation, the processing unit is specifically used for:
[0020] Based on the planar pose corresponding to the first position and the target plane, the target visible surface of the object to be detected in the image to be processed with a display area greater than a preset threshold is determined as the target image.
[0021] In one possible implementation, the processing unit is specifically used for:
[0022] The three-dimensional vertex coordinates of the region to be interfered with are determined based on the three-dimensional vertex coordinates of the target image in the first position. The three-dimensional vertex coordinates of the region to be interfered with are located in the target image, and the region to be interfered with is proportional to the target surface.
[0023] The second position is determined by projecting the coordinates of the three-dimensional vertices corresponding to the region to be disturbed onto the target plane.
[0024] In one possible implementation, the processing unit is specifically used for:
[0025] Based on the second position, the initial counter-perturbation is spatially deformed to obtain the target counter-perturbation.
[0026] The target adversarial perturbation is added at the second position, and the area occupied by the target adversarial perturbation in the image to be processed is consistent with the area divided by the second position in the image to be processed.
[0027] In one possible implementation, the acquisition unit is further configured to:
[0028] Acquire a target image, wherein the target image is the image to be processed with target adversarial perturbation added;
[0029] The processing unit is also used for:
[0030] Input the target image into the object detection model;
[0031] If the loss value corresponding to the target image does not meet the preset conditions, the target adversarial perturbation is updated to obtain a candidate adversarial perturbation. The target adversarial perturbation is obtained by updating based on the initial adversarial perturbation and the candidate adversarial perturbation.
[0032] The candidate adversarial perturbation is converted into the target adversarial perturbation until the loss value corresponding to the target detection model satisfies the preset condition.
[0033] In one possible implementation, the target adversarial perturbation is located at a first target position on the image to be processed, and the candidate adversarial perturbation is located at a second target position on the image to be processed, wherein the first target position and the second target position are located in the second position, and the processing unit is specifically used for:
[0034] If the loss value corresponding to the target detection model does not meet the preset conditions, the position of the target adversarial perturbation on the image to be processed is updated to obtain the candidate adversarial perturbation.
[0035] In one possible implementation, the video to be processed comprises N video frame images, wherein the image to be processed is the i-th video frame image among the N video frame images, and the first position is determined by the following method:
[0036] Obtain the third position and the fourth position corresponding to the object to be detected. The third position is the actual three-dimensional spatial position of the object to be detected in the (i-1)th video frame of the N video frames. The fourth position is the actual three-dimensional spatial position of the object to be detected in the (i+1)th video frame of the N video frames.
[0037] The first position is obtained by interpolating the third and fourth positions.
[0038] In one possible implementation, the video to be processed comprises N video frame images, wherein the image to be processed is the i-th video frame image among the N video frame images, and the first position is determined by the following method:
[0039] Obtain the M actual three-dimensional spatial positions corresponding to the object to be detected. The M actual three-dimensional spatial positions are the actual three-dimensional spatial positions corresponding to the positions of the object to be detected in the M video frame images respectively. The M video frame images are the M video frame images located before the i-th video frame image among the N video frame images.
[0040] The first position is determined based on the M actual three-dimensional spatial positions.
[0041] Thirdly, embodiments of this application disclose a computer device, which includes a processor and a memory:
[0042] The memory is used to store program code and transmit the program code to the processor;
[0043] The processor is configured to execute the image processing method described in any one of the first aspects according to the instructions in the program code.
[0044] Fourthly, embodiments of this application disclose a computer-readable storage medium for storing a computer program for executing the image processing method described in any one of the first aspects.
[0045] Fifthly, embodiments of this application disclose a computer program product including instructions that, when run on a computer, cause the computer to perform the image processing method described in any one of the first aspects.
[0046] As can be seen from the above technical solution, after acquiring the image to be processed including the object to be detected, the second position corresponding to the object to be detected can be determined according to the first position corresponding to the object to be detected. After adding the initial adversarial perturbation to the object to be detected in the image to be processed according to the indication of the second position, once the image to be processed with the adversarial perturbation is acquired, an adversarial sample will be formed (that is, the adversarial perturbation in the digital world is converted into the real physical world) and used as the input of the object detection model. On the one hand, since the second position closely matches the actual spatial position of the object to be detected in the real scene, when the position of the object to be detected changes in three-dimensional space, the second position can be changed accordingly. Therefore, when adversarial perturbation is added based on the second position, the position of the adversarial perturbation in the image matches the actual three-dimensional spatial position of the object to be detected. When the position of the object to be detected changes in three-dimensional space, the adversarial perturbation can accurately follow the change in the position of the object to be detected in three-dimensional space, thus possessing time consistency characteristics, making it difficult to detect and highly concealed. On the other hand, since this application adds adversarial perturbation based on the second position, and the second position is determined based on the three-dimensional spatial position, spatial deformation, angle changes, and other operations can be performed in conjunction with the three-dimensional spatial position of the object to be detected during the process of adding adversarial perturbation, making the adversarial perturbation more closely match the posture of the object to be detected in three-dimensional space and change based on the change in the posture of the object to be detected in three-dimensional space. Thus, it can more covertly fit with the object to be detected in the image to be processed, possessing spatial consistency. Therefore, the adversarial perturbation method of this application can make the adversarial perturbation have strong temporal and spatial consistency, thereby making the adversarial perturbation fit the object to be detected more closely and the perturbation more covert. Thus, model designers can use the processed images to more easily test the robustness of the model when facing physical texture adversarial attacks in the digital world simulation test object. At the same time, the processed images can be used as adversarial examples in adversarial attacks, which have strong offensive power against the model. In addition, the image addition method of this application can produce high-quality adversarial examples more efficiently, which has a significant optimization effect on the adversarial attack technology field, helps to accelerate the model iteration speed and shorten the model iteration cycle. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1A schematic diagram illustrating an image processing method in a practical application scenario provided by an embodiment of this application;
[0049] Figure 2 A flowchart illustrating an image processing method provided in an embodiment of this application;
[0050] Figure 3 A flowchart illustrating an image processing method in a practical application scenario provided in this application embodiment;
[0051] Figure 4 A schematic diagram illustrating an image processing method in a practical application scenario provided by an embodiment of this application;
[0052] Figure 5 A schematic diagram illustrating an image processing method in a practical application scenario provided by an embodiment of this application;
[0053] Figure 6 A structural block diagram of an image processing apparatus provided in an embodiment of this application;
[0054] Figure 7 A structural diagram of a terminal provided in an embodiment of this application;
[0055] Figure 8 This is a structural diagram of a server provided in an embodiment of this application. Detailed Implementation
[0056] The embodiments of this application will now be described with reference to the accompanying drawings.
[0057] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. For example, the first position refers to the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed, and the second position is the position used to guide the addition of anti-perturbation measures in the image to be processed. There is no sequential relationship between the first position and the second position. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The module divisions appearing in the embodiments of this application are merely logical divisions; in actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not performed. In addition, the shown or discussed mutual couplings or direct couplings or communication connections may be through some interfaces, and the indirect couplings or communication connections between modules may be electrical or other similar forms, none of which are limited in the embodiments of this application. Moreover, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed across multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.
[0058] The technical solution provided in this application can be applied to the technical scenario of robustness verification of a model. Specifically, it is used to generate a variety of test images that can verify the robustness of the model when performing robustness verification. For example, it can be used to add adversarial perturbations to multiple video frames included in a video or to add adversarial perturbations to a single image.
[0059] It is understandable that this method can be applied to computer devices capable of image processing, such as terminal devices or servers with image processing functions. This method can be executed independently by the terminal device or server, or it can be applied to network scenarios where the terminal device and server communicate, executing in cooperation. When the computer device is a server, it can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. When the computer device is a terminal, it can include, but is not limited to, smart terminals with multimedia data processing functions (e.g., video playback and music playback functions), such as smartphones, tablets, laptops, desktop computers, smart TVs, smart speakers, personal digital assistants (PDAs), desktop computers, and smartwatches.
[0060] The solutions in this application can be implemented based on artificial intelligence technology, specifically involving computer vision technology in artificial intelligence technology and cloud computing, cloud storage and database in cloud technology, which will be described separately below.
[0061] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0062] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning. This application's embodiments mainly relate to machine learning and computer vision technologies.
[0063] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0064] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0065] For example, embodiments of this application can utilize computer vision technology to identify objects to be detected in images, and when adjusting adversarial perturbations based on the identification results of object detection models, machine learning technology can be used to automatically complete multiple adjustments to the adversarial perturbations, thereby improving the adjustment efficiency of adversarial perturbations.
[0066] To facilitate understanding of the technical solutions provided in the embodiments of this application, the following will introduce an image processing method provided in the embodiments of this application in conjunction with a practical application scenario.
[0067] See Figure 1 , Figure 1 This is a schematic diagram of an image processing method in a practical application scenario provided by an embodiment of this application. In this practical application scenario, the computer device can be a terminal device 101, which can be, for example, a computer.
[0068] As shown in the figure, the image to be processed acquired by terminal device 101 is an image recording vehicle driving conditions, and the object to be detected can be a vehicle. Terminal device 101 can first determine the first position corresponding to the object to be detected. This first position is the actual three-dimensional spatial position of the object to be detected in the image to be processed, which can reflect the positional characteristics of the object to be detected in three-dimensional space. Based on the first position, terminal device 101 can determine the second position corresponding to the object to be detected. This second position is used to indicate the addition of adversarial perturbation to the object to be detected in the image to be processed.
[0069] Since the second position is determined based on the first position, and the first position accurately reflects the vehicle's positional characteristics in three-dimensional space, on the one hand, the second position closely matches the actual spatial position of the vehicle in the actual three-dimensional scene. When the vehicle's position in three-dimensional space changes, the second position can be changed accordingly. Therefore, when adding adversarial perturbations based on the second position, it can ensure that the adversarial perturbations and the vehicle's three-dimensional spatial position are consistent in time. On the other hand, during the process of adding adversarial perturbations based on the second position, spatial deformation, angle changes, and other operations can be combined with the vehicle's three-dimensional spatial position, making the adversarial perturbations more closely match the vehicle's posture in three-dimensional space. They can change based on the vehicle's posture changes in three-dimensional space, thus more subtly fitting the vehicle in the image to be processed, exhibiting spatial consistency. Next, various technical solutions provided by the embodiments of this application will be described in conjunction with the accompanying drawings.
[0070] See Figure 2 , Figure 2 A flowchart of an image processing method provided in this application embodiment, the method comprising:
[0071] S101: Obtain the image to be processed.
[0072] The image to be processed can be any image that requires adversarial perturbation. This image includes the object to be detected, which is the object to be detected during object detection. The adversarial perturbation interferes with the object recognition of this object, thus interfering with the detection results. The image to be processed can be a video frame from a video recording, and the object to be detected can be a specific vehicle, etc., within that video frame. Object detection can be performed by an object detection model. The adversarial perturbation can be used to launch an adversarial attack against the object detection model. The object detection model identifies objects in an image based on image information and detects the object's location, category, and other information based on the recognition results. The adversarial perturbation interferes with the object detection model's ability to recognize objects by affecting the image information corresponding to the object (e.g., changing the pixel value of the pixel corresponding to the object in the image). This reduces the accuracy of the object detection model's detection results regarding the object's location, category, etc., thereby verifying whether the object detection model has strong anti-interference capabilities, i.e., whether it has good robustness.
[0073] S102: Determine the second position of the object to be detected based on the first position of the object to be detected.
[0074] Understandably, in related technologies, the three-dimensional spatial position of the object to be detected is not considered when adding adversarial perturbations. Instead, the perturbations are added based solely on the object's position in the image. This results in the adversarial perturbations not conforming to the actual three-dimensional spatial characteristics of the object, leading to poor realism and effectiveness of the perturbations, and consequently, poor interference with the recognition of the object to be detected.
[0075] Based on this, in this embodiment of the application, the computer device can add adversarial perturbations to the object to be detected based on its actual three-dimensional spatial position in three-dimensional space. The computer device can obtain a first position corresponding to the object to be detected, which is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed.
[0076] Then, the computer device can determine the second position of the object to be detected, used to add adversarial perturbations, based on the first position corresponding to the object to be detected. Since the first position is the actual three-dimensional spatial position of the object to be detected, it can reflect the positional characteristics of the object to be detected in the image to be processed. Therefore, the second position determined based on the first position can fit the actual three-dimensional spatial position characteristics of the object to be detected.
[0077] S103: According to the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed.
[0078] Since the second position is determined based on the first position, it closely matches the actual spatial position of the object to be detected in the real scene. When the position of the object to be detected changes in three-dimensional space, the second position can be changed accordingly. Therefore, when adding adversarial perturbations based on the second position, it ensures that the adversarial perturbation and the three-dimensional spatial position of the object to be detected are consistent in time. Furthermore, during the process of adding adversarial perturbations based on the second position, spatial deformation and angle changes can be performed in conjunction with the three-dimensional spatial position of the object to be detected. This makes the adversarial perturbation more closely match the pose of the object to be detected in three-dimensional space, changing according to the pose changes of the object. This allows it to more subtly fit the object in the image being processed, achieving spatial consistency.
[0079] By achieving consistency in time and space, computer devices can make adversarial perturbations added based on a second location fit the object to be detected better, thus making the adversarial perturbations more concealed in the image to be processed and less likely to be detected by the object detection model. This allows for an effective test of the robustness of the object detection model without requiring personnel to collect samples in real-world scenarios.
[0080] Based on this, according to the indication of the second position, the computer device can add an initial adversarial perturbation to the object to be detected in the image to be processed. This initial adversarial perturbation can be added directly to the second position, or it can be added to a specific location within the second position; no limitation is made here.
[0081] As can be seen from the above technical solution, after acquiring the image to be processed including the object to be detected, the second position corresponding to the object to be detected can be determined according to the first position corresponding to the object to be detected. After adding the initial adversarial perturbation to the object to be detected in the image to be processed according to the indication of the second position, once the image to be processed with the adversarial perturbation is acquired, an adversarial sample will be formed (that is, the adversarial perturbation in the digital world is converted into the real physical world) and used as the input of the object detection model. On the one hand, since the position of the adversarial perturbation in the image matches the actual three-dimensional spatial position of the object to be detected, it can accurately follow the changes in the position of the object to be detected in three-dimensional space. Therefore, it has the characteristic of temporal consistency, is not easy to be detected, and has strong concealment. On the other hand, since this application adds the adversarial perturbation based on the second position, and the second position is determined based on the three-dimensional spatial position, the process of adding the adversarial perturbation can be combined with the three-dimensional spatial position of the object to be detected to perform spatial deformation, angle change and other operations, so that the adversarial perturbation fits the posture of the object to be detected in three-dimensional space more closely, and changes based on the changes in the posture of the object to be detected in three-dimensional space. Thus, it can fit more secretly with the object to be detected in the image to be processed, and has spatial consistency. Therefore, the adversarial perturbation method of this application can make the adversarial perturbation have strong temporal and spatial consistency, thereby making the adversarial perturbation fit the object to be detected more closely and the perturbation more covert. Thus, model designers can use the processed images to more easily test the robustness of the model when facing physical texture adversarial attacks in the digital world simulation test object. At the same time, the processed images can be used as adversarial examples in adversarial attacks, which have strong offensive power against the model. In addition, the image addition method of this application can produce high-quality adversarial examples more efficiently, which has a significant optimization effect on the adversarial attack technology field, helps to accelerate the model iteration speed and shorten the model iteration cycle.
[0082] The methods for determining the second position based on the three-dimensional spatial position of the object to be detected can include various approaches. For example, in one possible implementation, when determining the second position of the object to be detected based on the first position, the computer device can perform a projection. The computer device can project the first position onto a target plane to determine the second position of the object to be detected. This target plane is the calibration camera's two-dimensional plane corresponding to the device that generates the image to be processed. This device can be, for example, an image capturing device such as a camera, or an image generating device capable of automatically generating images (such as a computer or mobile phone). The calibration camera's two-dimensional plane refers to the two-dimensional plane corresponding to the image to be processed when the device generates the image; for example, it can be the shooting plane corresponding to a camera.
[0083] This projection transformation allows for the accurate mapping of the second position on the two-dimensional plane corresponding to the first position, based on the positional relationship between the three-dimensional space corresponding to the image to be processed and the two-dimensional plane of the calibration camera corresponding to the device that generated the image.
[0084] Understandably, an object to be detected can have multiple faces. In the image being processed, due to factors such as the image generation angle, not all of these faces may be visible; that is, not all of them are visible in the image. Object detection is primarily based on the objects displayed in the image. These objects are actually composed of their visible faces within the image. Therefore, to add highly concealed adversarial perturbations, these perturbations should closely match the position and orientation characteristics of the visible faces of the object in the image within three-dimensional space.
[0085] Based on this, in one possible implementation, in order to improve the fit of the adversarial perturbation to the object to be detected, and thus further improve the concealment of the added adversarial perturbation, the computer device can add adversarial perturbation based on the visible surface of the object to be detected.
[0086] When determining the second position corresponding to the object to be detected by projecting the object to the target plane based on the first position, the computer device can first determine the target screen corresponding to the object to be detected in the image to be processed, wherein the target screen is the visible surface of the object to be detected in the image to be processed.
[0087] Then, the computer device can project and transform the object to be detected onto the target plane according to the actual three-dimensional spatial position of the object to be detected on the target screen to determine the second position. Thus, when adding initial adversarial perturbation according to the indication of the second position, the initial adversarial perturbation can be made to better match the characteristics of the visible surface of the object to be detected in the three-dimensional space of the image to be processed, thereby further improving the interference capability of the initial adversarial perturbation for object detection.
[0088] It is understandable that the relative relationship between the position of the object to be detected in three-dimensional space and the two-dimensional plane corresponding to the image to be processed determines the visible surface of the object in the image to be processed. For example, the side of the object to be detected facing the two-dimensional plane corresponding to the image to be processed is usually the visible surface. Based on this, in one possible implementation, when determining the target image corresponding to the object to be detected in the image to be processed, the computer device can determine the visible surface of the object to be detected based on the three-dimensional spatial position of the object to be detected and the two-dimensional plane corresponding to the image to be processed.
[0089] The computer device can determine the target image based on the first position and the planar pose corresponding to the target plane, identifying the visible surface of the object to be detected in the image to be processed as having a display area greater than a preset threshold. Here, planar pose refers to the orientation of the target plane in three-dimensional space, such as its position or rotation angle. Based on the position and orientation of the object to be detected in three-dimensional space as indicated by the first position, and the orientation of the image to be processed in the corresponding two-dimensional plane based on the planar pose, the computer device can determine the display status of each face of the object to be detected in the image to be processed, thereby determining the display area of each face in the image to be processed.
[0090] It is understandable that since the model detects objects based on the information displayed by the objects in the image being processed, the visible surfaces with a small display area have little effect on object detection. Adding adversarial perturbations to these visible surfaces will not have a significant impact. Therefore, computer devices can only add adversarial perturbations to target visible surfaces with a display area greater than a preset threshold.
[0091] Specifically, the actual three-dimensional spatial position can be constituted by the three-dimensional vertex coordinates of the object to be detected in three-dimensional space. In one possible implementation, when determining the second position by projecting the object to be detected onto the target plane based on its actual three-dimensional spatial position corresponding to the target image, the computer device can determine the three-dimensional vertex coordinates corresponding to the region to be interfered with based on the three-dimensional vertex coordinates of the target image corresponding to the first position. The three-dimensional vertex coordinates of the region to be interfered with are located on the target image, and the region to be interfered with is proportional to the target surface. That is, the region to be interfered with can be the entire target image or a part of the target image.
[0092] The computer device can project the three-dimensional vertex coordinates corresponding to the region to be interfered with onto the target plane to determine the second position. These three-dimensional vertex coordinates can be the vertex coordinates corresponding to the faces of the object being detected itself, or the vertex coordinates corresponding to faces that characterize the three-dimensional spatial position of the object being detected, such as the three-dimensional vertex coordinates corresponding to the faces of the smallest circumscribed cuboid of the object being detected.
[0093] Since the second position is determined based on the three-dimensional spatial position of the object to be detected, the shape of the adversarial perturbation is not considered during determination. Therefore, there may be a situation where the shape of the initial adversarial perturbation does not match the shape corresponding to the second position. In one possible implementation, in order to make the adversarial perturbation fit the second position better and thus achieve a better interference effect, when adding the initial adversarial perturbation to the object to be detected in the image to be processed according to the indication of the second position, the computer device can first perform spatial deformation on the initial adversarial perturbation according to the second position to obtain the target adversarial perturbation, which has a good fit with the second position in spatial shape.
[0094] The computer device can add the target adversarial perturbation at the second location, and the area occupied by the target adversarial perturbation in the image to be processed is consistent with the area divided at the second location in the image to be processed. The spatial deformation can be performed in various ways, such as through perspective transformation.
[0095] When introducing step S101, it was mentioned that the role of adversarial perturbation is to interfere with the model's object recognition of the object to be detected. Therefore, in one possible implementation, in order to ensure the interference effect of the added adversarial perturbation on the object detection model, the computer device can analyze the interference effect of the adversarial perturbation through the object detection model capable of object detection, and adjust the adversarial perturbation based on the analysis results to obtain an adversarial perturbation that can effectively interfere with the detection effect of the object detection model, so that the adversarial perturbation can have a relatively effective testing effect on the robustness of the object detection model.
[0096] The computer device can first acquire a target image, which is the image to be processed with added adversarial perturbations. The adversarial perturbations continuously change based on the detection performance of the object detection model. The adversarial perturbation added to the target image is the latest modified version, while the adversarial perturbation added to the earliest acquired target image is the initial perturbation mentioned above. Furthermore, the adversarial perturbations are added based on the second position determined above.
[0097] A computer device can input the target image into an object detection model and obtain a loss value corresponding to the target image. This loss value reflects the difference between the detection result obtained by the object detection model in detecting objects in the target image and the actual detection result of the object to be detected in the target image. The larger the loss value, the greater the difference between the object detection model's detection result for the object to be detected in the target image and the actual result corresponding to the object to be detected. In other words, the weaker the object detection model's ability to detect objects in the target image, the stronger the interference effect of adversarial perturbations on the object detection model's object detection ability.
[0098] Therefore, the computer device can set a preset condition to determine whether the target adversarial perturbation can effectively interfere with the object detection model. If the loss value corresponding to the target image meets the preset condition, it indicates that the adversarial perturbation added to the target image can effectively interfere with the object detection model's detection performance against the target object. If the loss value corresponding to the target image does not meet the preset condition, it indicates that the target adversarial perturbation has not had a sufficient interference effect on the object detection model. In this case, the computer device can update the target adversarial perturbation to obtain a candidate adversarial perturbation. This candidate adversarial perturbation is updated based on the initial adversarial perturbation and the candidate adversarial perturbation. That is, after adjusting and obtaining the candidate adversarial perturbation, the candidate adversarial perturbation will be used as the target adversarial perturbation for the next input object detection model. In the next model detection, the computer device will use the image to be processed with the added candidate adversarial perturbation as the target image.
[0099] The computer device can convert the candidate adversarial actions into the target adversarial perturbation until the loss value corresponding to the target detection model meets the preset condition. Since the adjusted adversarial perturbation can make the loss value corresponding to the target image meet the preset condition, the target adversarial perturbation added to the target image can effectively interfere with the detection effect of the object detection model on the object to be detected. Therefore, the computer device can add the target adversarial action as the final adversarial perturbation to the image to be processed for subsequent effective verification of the robustness of the object detection model.
[0100] The object detection model can be of various types. For example, it can be a location detection model, which is used to identify the position of objects in an image. The loss value can be based on the difference between the actual position attribute information of the object to be detected in the image and the position attribute information detected by the location detection model. The position attribute information describes the position of the object to be detected in the image. For example, when the image to be processed is a video frame obtained by shooting a vehicle, the location detection model can be used to detect the position of the vehicle in three-dimensional space. The loss value corresponding to the image can be determined by the vehicle's three-dimensional spatial position detected by the location detection model and the vehicle's actual position in three-dimensional space (i.e., the first position mentioned above).
[0101] In another possible implementation, the object detection model can be a category detection model, which is used to detect the object category of objects in an image. The loss value can be based on the difference between the actual category attribute information of the object to be detected in the image and the category attribute information detected by the category detection model. The category attribute information is used to describe the object category of the object to be detected in the image. For example, when the image to be processed is a video frame obtained by shooting a video of a vehicle, the category detection model can be used to detect the vehicle type, vehicle model, etc. The difference between the vehicle type determined by the category detection model and the actual type of the vehicle can determine the loss value corresponding to the image.
[0102] When adjusting target adversarial perturbations based on loss values, computer equipment can adjust both the information composition of the target adversarial perturbation itself and its position in the image to be processed, so as to achieve diversified adversarial perturbation adjustment.
[0103] In one possible implementation, the target adversarial perturbation can be located at a first target position on the image to be processed, and the candidate adversarial perturbation can be located at a second target position on the image to be processed, with the first target position and the second target position located within the second position. That is, when the target adversarial perturbation needs to be adjusted, the computer device will only adjust the target adversarial perturbation at the second position to ensure the degree of fit of the target adversarial perturbation to the three-dimensional spatial position of the device to be detected. When updating the target adversarial perturbation, if the loss value corresponding to the target detection model does not meet the preset conditions, the computer device can update the position of the target adversarial perturbation on the image to be processed to obtain the candidate adversarial perturbation.
[0104] This position adjustment method allows for more diversified adjustment of the target countermeasures, while maintaining the degree of fit of the countermeasures to the three-dimensional position, ultimately resulting in a countermeasures with better interference effect.
[0105] Furthermore, the determination of the first position corresponding to the object to be detected can be achieved in various ways. For example, the first position can be collected manually. Alternatively, to further improve the efficiency of adding anti-perturbation measures, computer equipment can automatically determine the first position corresponding to the object to be detected through various methods.
[0106] For example, in one possible implementation, the video to be processed includes N video frame images, and the image to be processed can be the i-th video frame image among the N video frame images, the first position of which can be determined by the following method:
[0107] The computer device can acquire the third and fourth positions corresponding to the object to be detected. The third position is the actual three-dimensional spatial position of the object to be detected in the (i-1)th video frame out of N video frames, and the fourth position is the actual three-dimensional spatial position of the object to be detected in the (i+1)th video frame out of N video frames. Therefore, through these third and fourth positions, the position of the object to be detected before and after the corresponding moment in the video to be processed can be reflected, thus demonstrating the positional change of the object to be detected at the corresponding moment in the video to be processed. Here, i is a positive integer greater than 1 and not exceeding N-1, and N is a positive integer greater than 2.
[0108] The computer device can perform interpolation processing on the third position and the fourth position to obtain the first position. The first position can more accurately reflect the three-dimensional spatial position of the object to be detected at the corresponding moment in the image to be processed when the object changes from the third position to the fourth position.
[0109] In another possible implementation, the computer device can also use multiple positional information prior to the time corresponding to the image to be processed to analyze the pattern of positional changes and predict the first position. For example, the video to be processed includes N video frame images, and the image to be processed is the i-th video frame image among the N video frame images. The first position can be determined by:
[0110] The computer device can acquire M actual three-dimensional spatial positions corresponding to the object to be detected. These M positions are the actual three-dimensional spatial positions of the object to be detected in M video frames, where each of the N video frames is located before the i-th video frame. In other words, these M actual three-dimensional spatial positions represent the movement position of the object to be detected before the time corresponding to the image being processed, thus revealing the positional changes of the object to be detected.
[0111] The computer device can determine the first position based on the M actual three-dimensional spatial positions. For example, it can fit and predict the position change path of the object to be detected based on the M actual three-dimensional spatial positions to determine the three-dimensional spatial position of the object to be detected at the corresponding moment in the image to be processed, thereby obtaining the first position. Here, i is a positive integer greater than M and not exceeding N, N is a positive integer greater than 1, and M is a positive integer less than N.
[0112] The above methods can be used to obtain the three-dimensional spatial position of the object to be detected in each frame of an image without manual processing, which greatly reduces the manual processing burden and improves the efficiency of adding anti-disturbance measures.
[0113] Furthermore, in practical image processing, depending on the data format of the video to be processed, the computer device can process the video frame by frame. For example, the video to be processed may include N video frame images, and the image to be processed is the i-th video frame image among the N video frame images. In response to determining the adversarial perturbation corresponding to the (i-1)-th video frame image, the computer device can further determine the adversarial perturbation corresponding to the i-th video frame, thereby realizing the addition of adversarial perturbations to the N video frame images frame by frame. Here, i is a positive integer greater than 1 and not exceeding N, and N is a positive integer greater than 1.
[0114] To facilitate understanding of the technical solutions provided in the embodiments of this application, the following will introduce an image processing method provided in the embodiments of this application in conjunction with a practical application scenario.
[0115] See Figure 3 , Figure 3 A flowchart illustrating an image processing method in a practical application scenario provided by this application embodiment. The method may include the following steps:
[0116] S201: Obtain a complete video segment and the corresponding 3D annotation information for each frame, and begin traversing this video segment.
[0117] The video clip may include multiple video frames. The 3D annotation information is used to annotate the actual three-dimensional spatial position of the object to be detected in each video frame, that is, the first position in each video frame.
[0118] S202: In the current video frame, select the 3D bounding box annotation of the object to be detected, calculate the occlusion of the four sides of the object, and determine the observable visible surface.
[0119] In this embodiment, the 3D annotation information can be 3D bounding box annotations, which refer to the spatial position information corresponding to the smallest bounding box of the object to be detected, such as... Figure 4As shown, the 3D bounding box annotation can include the three-dimensional spatial coordinates of each vertex of the vehicle's smallest circumscribed cuboid. Figure 4 The image shows the 3D spatial coordinate system corresponding to this video frame, the minimum bounding cube corresponding to the vehicle, and some vertices on the cube. The vehicle is the object to be detected in this image.
[0120] like Figure 5 As shown, computer equipment can determine the visible surface of each video frame in each scene of the vehicle based on the actual three-dimensional spatial position identified by the 3D bounding box annotation and the pose of the camera calibration two-dimensional plane corresponding to the camera used to generate the video clip in three-dimensional space (such as camera angle).
[0121] S203: For the observable target image, calculate the 3D coordinates of the four corner points of the patch on the surface according to the required patch size, project the four corner points onto the image plane of each camera, and calculate the 2D coordinates.
[0122] Patch refers to the adversarial perturbation applied in this practical application scenario. In this scenario, the computer device does not add adversarial perturbations to every visible surface. Instead, based on the display area of each visible surface in the current frame, it determines the target image whose display area is greater than a preset threshold and adds adversarial perturbations to the target image. For example, in Figure 5 In the system, the top, front, and sides of the vehicle are all visible surfaces, but only the side surfaces have a large display area that meets the preset threshold. Therefore, the computer equipment will add anti-disturbance measures to the sides of the vehicle.
[0123] The computer device can determine the coordinates of the four corner points of the adversarial disturbance in three-dimensional space based on the coordinates of the four corner points of the target image in three-dimensional space. For example, the computer device can determine the four corner points of the area to be interfered with based on the size of the adversarial disturbance to be added, and project the four corner points onto the two-dimensional plane of the camera calibration corresponding to the current frame to obtain 2D coordinates.
[0124] The specific coordinate projection method can be as follows:
[0125] For an object in a single frame image, an adversarial perturbation can be added to a specified 3D location of the object to be detected. In the real world, the size of this 3D location is H. p ×W p Given the annotation information of the 3D bounding box of an object, the positions of the four corners at which adversarial perturbations are added can be calculated in 3D space, denoted by (p1, p2, p3, p4) in a 3D coordinate system, where each point is a 3D coordinate. For each camera view, there is a projection matrix. Based on camera intrinsics and a 3D coordinate system p = (px, py, pz) T Points that can be projected onto the camera calibration two-dimensional plane as follows:
[0126]
[0127] S204: Transform the anti-aliasing graphic based on the 2D coordinates of the corner points and add it to the image.
[0128] The computer device can obtain the 2D coordinates of the corner points of the area to be interfered with based on the projection, and add the adversarial map to the current frame. The adversarial map is the adversarial perturbation in the actual application scenario.
[0129] For example, a computer device can digitally apply an adversarial patch to a target quadrilateral defined by the projection corner coordinates (p′1, p′2, p′3, p′4) using perspective transformation, using a projection coefficient vector. This allows the pixel (tx, ty) to be projected back to its source in the target quadrilateral. p ×W p The patch's position (sx, sy) can be fitted to the actual area in 3D space where anti-perturbation measures need to be added.
[0130]
[0131] The coefficients can be solved using the corresponding angular coordinates. Then, the pixel color at (tx, ty) on the 2D image can be interpolated to the neighboring pixels around (sx, sy) on the patch. This process is differentiable and can therefore be optimized using gradient-based methods.
[0132] S205: Input the image with the added adversarial map into the object detector to obtain the loss function, calculate the gradient of the loss function, and update the adversarial map.
[0133] The target detector is as follows Figure 5 As shown, the target detector carries an object detection model. Image features can be extracted through the backbone network of the model. Combined with the analysis of 3D spatial features, the detection results of the objects to be detected in the image can be obtained. For example, the position of a vehicle in a video frame can be detected.
[0134] Consider an object detection model with a loss function of L. m Then the loss function against the attack is: Where x is the input image and p is the adversarial perturbation. This embodiment of the application can optimize the adversarial patching in consecutive image frames and images captured by multiple cameras; therefore, the loss function of this invention is: Here, the expectation with respect to t represents the expectation over consecutive image frames, and the expectation with respect to c represents the expectation over images captured by multiple cameras.
[0135] S206: Enter the next frame of the video segment and repeat steps S202-S205 until the segment traversal is complete.
[0136] Through the above steps, this embodiment of the application can perform object detection on each video frame with added adversarial maps based on the target detector, determine the loss value corresponding to each video frame through the target detector, and adjust the adversarial map based on the loss value corresponding to each video frame. This allows the adversarial map to effectively interfere with the object detection function of the target detector in each video frame. Thus, the interference effect of the adversarial map in multiple video frames can be combined to adjust the adversarial map as a whole, so that the final adversarial map can play a better interference effect on the video segment as a whole.
[0137] Figure 6 This is a schematic block diagram of an image processing apparatus provided in an embodiment of this application. Figure 6 As shown, corresponding to the above image processing methods, this application embodiment also provides an image processing apparatus 500. The image processing apparatus 500 is used to perform image processing, adding anti-perturbation elements to the image to be processed. The image processing apparatus 500 includes... Figure 6 Acquisition unit 501 and processing unit 502:
[0138] The acquisition unit 501 is used to acquire an image to be processed, the image to be processed including an object to be detected;
[0139] The processing unit 502 is used to determine the second position corresponding to the object to be detected based on the first position corresponding to the object to be detected, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed.
[0140] Based on the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed.
[0141] In one possible implementation, the processing unit 502 is specifically used for:
[0142] Based on the first position, a projection transformation is performed on the target plane to determine the second position corresponding to the object to be detected. The target plane is the two-dimensional plane of the calibration camera corresponding to the device that generates the image to be processed.
[0143] In one possible implementation, the processing unit 502 is specifically used for:
[0144] Determine the target image corresponding to the object to be detected in the image to be processed, wherein the target image is the visible surface of the object to be detected in the image to be processed;
[0145] The second position is determined by projecting the object to be detected onto the target plane based on its actual three-dimensional spatial position corresponding to the target image.
[0146] In one possible implementation, the processing unit 502 is specifically used for:
[0147] Based on the planar pose corresponding to the first position and the target plane, the target visible surface of the object to be detected in the image to be processed with a display area greater than a preset threshold is determined as the target image.
[0148] In one possible implementation, the processing unit 502 is specifically used for:
[0149] The three-dimensional vertex coordinates of the region to be interfered with are determined based on the three-dimensional vertex coordinates of the target image in the first position. The three-dimensional vertex coordinates of the region to be interfered with are located in the target image, and the region to be interfered with is proportional to the target surface.
[0150] The second position is determined by projecting the coordinates of the three-dimensional vertices corresponding to the region to be disturbed onto the target plane.
[0151] In one possible implementation, the processing unit 502 is specifically used for:
[0152] Based on the second position, the initial counter-perturbation is spatially deformed to obtain the target counter-perturbation.
[0153] The target adversarial perturbation is added at the second position, and the area occupied by the target adversarial perturbation in the image to be processed is consistent with the area divided by the second position in the image to be processed.
[0154] In one possible implementation, the acquisition unit 501 is further configured to:
[0155] Acquire a target image, wherein the target image is the image to be processed with target adversarial perturbation added;
[0156] The processing unit 502 is further configured to:
[0157] Input the target image into the object detection model;
[0158] If the loss value corresponding to the target image does not meet the preset conditions, the target adversarial perturbation is updated to obtain a candidate adversarial perturbation. The target adversarial perturbation is obtained by updating based on the initial adversarial perturbation and the candidate adversarial perturbation.
[0159] The candidate adversarial perturbation is converted into the target adversarial perturbation until the loss value corresponding to the target detection model satisfies the preset condition.
[0160] In one possible implementation, the target adversarial perturbation is located at a first target position on the image to be processed, and the candidate adversarial perturbation is located at a second target position on the image to be processed, wherein the first target position and the second target position are located in the second position, and the processing unit 502 is specifically used for:
[0161] If the loss value corresponding to the target detection model does not meet the preset conditions, the position of the target adversarial perturbation on the image to be processed is updated to obtain the candidate adversarial perturbation.
[0162] In one possible implementation, the video to be processed comprises N video frame images, wherein the image to be processed is the i-th video frame image among the N video frame images, and the first position is determined by the following method:
[0163] Obtain the third position and the fourth position corresponding to the object to be detected. The third position is the actual three-dimensional spatial position of the object to be detected in the (i-1)th video frame of the N video frames. The fourth position is the actual three-dimensional spatial position of the object to be detected in the (i+1)th video frame of the N video frames.
[0164] The first position is obtained by interpolating the third and fourth positions.
[0165] In one possible implementation, the video to be processed comprises N video frame images, wherein the image to be processed is the i-th video frame image among the N video frame images, and the first position is determined by the following method:
[0166] Obtain the M actual three-dimensional spatial positions corresponding to the object to be detected. The M actual three-dimensional spatial positions are the actual three-dimensional spatial positions corresponding to the positions of the object to be detected in the M video frame images respectively. The M video frame images are the M video frame images located before the i-th video frame image among the N video frame images.
[0167] The first position is determined based on the M actual three-dimensional spatial positions.
[0168] As can be seen from the above technical solution, after acquiring the image to be processed including the object to be detected, the second position corresponding to the object to be detected can be determined according to the first position corresponding to the object to be detected. After adding the initial adversarial perturbation to the object to be detected in the image to be processed according to the indication of the second position, once the image to be processed with the adversarial perturbation is acquired, an adversarial sample will be formed (that is, the adversarial perturbation in the digital world is converted into the real physical world) and used as the input of the object detection model. On the one hand, since the position of the adversarial perturbation in the image matches the actual three-dimensional spatial position of the object to be detected, it can accurately follow the changes in the position of the object to be detected in three-dimensional space. Therefore, it has the characteristic of temporal consistency, is not easy to be detected, and has strong concealment. On the other hand, since this application adds the adversarial perturbation based on the second position, and the second position is determined based on the three-dimensional spatial position, the process of adding the adversarial perturbation can be combined with the three-dimensional spatial position of the object to be detected to perform spatial deformation, angle change and other operations, so that the adversarial perturbation fits the posture of the object to be detected in three-dimensional space more closely, and changes based on the changes in the posture of the object to be detected in three-dimensional space. Thus, it can fit more secretly with the object to be detected in the image to be processed, and has spatial consistency. Therefore, the adversarial perturbation method of this application can make the adversarial perturbation have strong temporal and spatial consistency, thereby making the adversarial perturbation fit the object to be detected more closely and the perturbation more covert. Thus, model designers can use the processed images to more easily test the robustness of the model when facing physical texture adversarial attacks in the digital world simulation test object. At the same time, the processed images can be used as adversarial examples in adversarial attacks, which have strong offensive power against the model. In addition, the image addition method of this application can produce high-quality adversarial examples more efficiently, which has a significant optimization effect on the adversarial attack technology field, helps to accelerate the model iteration speed and shorten the model iteration cycle.
[0169] This application also provides a computer device; please refer to [link to relevant documentation]. Figure 7 As shown, the computer device can be a terminal device; for example, a mobile phone can be used as a terminal device.
[0170] Figure 7 This diagram illustrates a partial structural representation of a mobile phone related to the terminal device provided in this embodiment. (Reference) Figure 7 The mobile phone includes components such as a radio frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless Fidelity (Wi-Fi) module 770, a processor 780, and a power supply 790. Those skilled in the art will understand that... Figure 7The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0171] The following is combined with Figure 7 A detailed introduction to each component of a mobile phone:
[0172] RF circuit 710 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 780; additionally, it transmits uplink data to the base station. Typically, RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), and a duplexer. Furthermore, RF circuit 710 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).
[0173] The memory 720 can be used to store software programs and modules. The processor 780 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0174] The input unit 730 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 731), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 780, and can also receive and execute commands sent by the processor 780. In addition, the touch panel 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0175] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741, which may optionally be configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel. Further, a touch panel 731 may cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it transmits the information to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides corresponding visual output on the display panel 741 based on the type of touch event. Although in Figure 7 In this embodiment, the touch panel 731 and the display panel 741 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.
[0176] The mobile phone may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can turn off the display panel 741 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0177] Audio circuit 760, speaker 761, and microphone 762 provide an audio interface between the user and the mobile phone. Audio circuit 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. On the other hand, microphone 762 converts collected sound signals into electrical signals, which are received by audio circuit 760, converted into audio data, and then processed by processor 780 before being transmitted via RF circuit 710 to, for example, another mobile phone, or the audio data can be output to memory 720 for further processing.
[0178] Wi-Fi is a short-range wireless transmission technology. Through the Wi-Fi module 770, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 7 The Wi-Fi module 770 is shown, but it is understood that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.
[0179] The processor 780 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 720, and calls data stored in the memory 720 to perform various functions and process data, thereby performing overall detection of the phone. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 780.
[0180] The mobile phone also includes a power supply 790 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 780 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0181] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0182] In this embodiment, the processor 780 included in the terminal device also has the following functions:
[0183] Acquire an image to be processed, wherein the image to be processed includes an object to be detected;
[0184] Based on the first position corresponding to the object to be detected, the second position corresponding to the object to be detected is determined, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed;
[0185] Based on the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed.
[0186] This application also provides a server; please refer to [link / reference]. Figure 8 As shown, Figure 8 This is a structural diagram of a server 800 provided in an embodiment of this application. The server 800 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 822 (e.g., one or more processors) and a memory 832, and one or more storage media 830 (e.g., one or more mass storage devices) for storing application programs 842 or data 844. The memory 832 and storage media 830 can be temporary or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server. Furthermore, the CPU 822 may be configured to communicate with the storage media 830 and execute the series of instruction operations in the storage media 830 on the server 800.
[0187] Server 800 may also include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows Server. TM Mac OS X TM Unix TM Linux TM FreeBSDTM etc.
[0188] The steps performed by the server in the above embodiments can be based on Figure 8 The server structure shown.
[0189] This application also provides a computer-readable storage medium for storing a computer program that executes any one of the image processing methods described in the foregoing embodiments.
[0190] This application also provides a computer program product including a computer program, which, when run on a computer device, causes the computer device to perform any of the image processing methods described in the above embodiments.
[0191] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disk, or optical disk, etc., and other media capable of storing program code.
[0192] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0193] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image processing method, characterized by, The method includes: Acquire an image to be processed, wherein the image to be processed includes an object to be detected; Based on the first position corresponding to the object to be detected, the second position corresponding to the object to be detected is determined, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed; According to the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed; The step of determining the second position corresponding to the object to be detected based on the first position includes: Based on the first position, a projection transformation is performed on the target plane to determine the second position corresponding to the object to be detected. The target plane is the two-dimensional plane of the calibration camera corresponding to the device that generates the image to be processed. The step of projecting the object to be detected onto the target plane based on the first position to determine the second position includes: Determine the target image corresponding to the object to be detected in the image to be processed, wherein the target image is the visible surface of the object to be detected in the image to be processed; The second position is determined by projecting the actual three-dimensional spatial position of the object to be detected onto the target plane based on the target image; The step of determining the second position by projecting the object to be detected onto the target plane based on the actual three-dimensional spatial position of the object in the target image includes: The three-dimensional vertex coordinates of the region to be interfered with are determined based on the three-dimensional vertex coordinates of the target image in the first position. The three-dimensional vertex coordinates of the region to be interfered with are located in the target image, and the region to be interfered with is proportional to the target image. The second position is determined by projecting the coordinates of the three-dimensional vertices corresponding to the region to be disturbed onto the target plane.
2. The method of claim 1, wherein, Determining the target image corresponding to the object to be detected in the image to be processed includes: Based on the planar pose corresponding to the first position and the target plane, the target visible surface of the object to be detected in the image to be processed with a display area greater than a preset threshold is determined as the target image.
3. The method of claim 1, wherein, The step of adding an initial adversarial perturbation to the object to be detected in the image to be processed according to the indication of the second position includes: Based on the second position, the initial counter-perturbation is spatially deformed to obtain the target counter-perturbation. The target adversarial perturbation is added at the second position, and the area occupied by the target adversarial perturbation in the image to be processed is consistent with the area divided by the second position in the image to be processed.
4. The method according to claim 1, characterized in that, The method further includes: Acquire a target image, wherein the target image is the image to be processed with target adversarial perturbation added; Input the target image into the object detection model; If the loss value corresponding to the target image does not meet the preset conditions, the target adversarial perturbation is updated to obtain a candidate adversarial perturbation. The target adversarial perturbation is obtained by updating based on the initial adversarial perturbation and the candidate adversarial perturbation. The candidate adversarial perturbation is converted into the target adversarial perturbation until the loss value corresponding to the target detection model satisfies the preset condition.
5. The method according to claim 4, characterized in that, The target adversarial perturbation is located at a first target position on the image to be processed, and the candidate adversarial perturbation is located at a second target position on the image to be processed. The first target position and the second target position are located in the second position. If the loss value corresponding to the target detection model does not meet the preset conditions, the target adversarial perturbation is updated to obtain the candidate adversarial perturbation, including: If the loss value corresponding to the target detection model does not meet the preset conditions, the position of the target adversarial perturbation on the image to be processed is updated to obtain the candidate adversarial perturbation.
6. The method according to claim 1, characterized in that, The video to be processed includes N video frame images, and the image to be processed is the i-th video frame image among the N video frame images. The first position is determined by the following method: Obtain the third position and the fourth position corresponding to the object to be detected. The third position is the actual three-dimensional spatial position of the object to be detected in the (i-1)th video frame of the N video frames. The fourth position is the actual three-dimensional spatial position of the object to be detected in the (i+1)th video frame of the N video frames. The first position is obtained by interpolating the third and fourth positions.
7. The method according to claim 1, characterized in that, The video to be processed includes N video frame images, and the image to be processed is the i-th video frame image among the N video frame images. The first position is determined by the following method: Obtain the M actual three-dimensional spatial positions corresponding to the object to be detected. The M actual three-dimensional spatial positions are the actual three-dimensional spatial positions corresponding to the positions of the object to be detected in the M video frame images respectively. The M video frame images are the M video frame images located before the i-th video frame image among the N video frame images. The first position is determined based on the M actual three-dimensional spatial positions.
8. An image processing apparatus, characterized in that, The device includes an acquisition unit and a processing unit: The acquisition unit is used to acquire an image to be processed, wherein the image to be processed includes an object to be detected; The processing unit is configured to determine a second position corresponding to the object to be detected based on a first position corresponding to the object to be detected, wherein the first position is the actual three-dimensional spatial position corresponding to the position of the object to be detected in the image to be processed. According to the indication of the second position, an initial adversarial perturbation is added to the object to be detected in the image to be processed; The processing unit is specifically used to: perform a projection transformation on the target plane based on the first position to determine the second position corresponding to the object to be detected, wherein the target plane is the two-dimensional plane of the calibration camera corresponding to the device that generates the image to be processed; The processing unit is specifically used to: determine the target image corresponding to the object to be detected in the image to be processed, wherein the target image is the visible surface of the object to be detected in the image to be processed; and perform a projection transformation on the target plane according to the actual three-dimensional spatial position of the object to be detected corresponding to the target image to determine the second position. The processing unit is specifically used to: determine the three-dimensional vertex coordinates of the area to be interfered with based on the three-dimensional vertex coordinates of the target image in the first position, wherein the three-dimensional vertex coordinates of the area to be interfered with are located in the target image, and the area to be interfered with is proportional to the target image; The second position is determined by projecting the coordinates of the three-dimensional vertices corresponding to the region to be disturbed onto the target plane.
9. A computer device, characterized in that, The computer device includes a processor and memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the image processing method according to any one of claims 1-7 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the image processing method according to any one of claims 1-7.
11. A computer program product comprising instructions that, when run on a computer, causes the computer to perform the image processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Confrontation disturbance generation method and device and storage medium
CN114387647A
Confrontation image generation method and device, electronic equipment and storage medium
CN116152613A