Camera attitude estimation model-based adversarial robust evaluation method and camera attitude estimation model-based adversarial robust evaluation system
By calculating the projection direction consistency loss and constructing disk background images of complex texture structures, the problem of the anti-rooted robust evaluation method in the prior art ignores different perspectives and environmental conditions, improving the robustness of the anti-attack and the reliability evaluation of the camera pose estimation model.
Patent Information
- Application Number
- CN202510130239.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing adversarial robustness evaluation methods rely on a single adversarial sample, neglecting the comprehensive assessment of model robustness under different perspectives, different backgrounds and different environmental conditions, resulting in poor generalization ability and inability to show good attack effects in a changeable real-life environment.
By calculating the consistency loss of projection direction, the fragment images are optimized, so that the fragment images at different perspectives have a high degree of consistency, and the spatial information cannot be accurately estimated, which can cause misleading in the camera estimation process and improve the effect of counterattacks. A disc background image with complex texture structures is constructed and used as an adversarial sample for a robust evaluation of the camera pose estimation model.
It improves the robustness of the attack, enables the disc background image to have significant interference capabilities in both digital and physical environments, and provides an accurate evaluation method for camera pose estimation at different perspectives, ensuring the reliability evaluation of the model under variable perspectives and environmental conditions.
Smart Images

Figure CN120104445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an adversarial robust evaluation method and system based on a camera pose estimation model. Background Art
[0002] The adversarial robustness evaluation method based on the camera pose estimation model is a method that evaluates the robustness of the camera pose estimation model in the face of harsh environments or uncertain factors by simulating attacks (such as image perturbations). This method performs adversarial perturbations on the camera input image and analyzes the performance of the model under different attack conditions, thereby verifying the model's resistance to these perturbations.
[0003] With the widespread application of deep learning models in fields such as autonomous driving and robot navigation, the accuracy and stability of camera pose estimation are crucial to the security of the system. Designing an effective adversarial robustness evaluation method can comprehensively test the performance of the model in complex and dynamic environments, improve the security and reliability of the camera pose estimation model, and provide a theoretical basis for protection mechanisms in practical applications.
[0004] However, existing adversarial robustness evaluation methods rely on a single adversarial sample, ignoring the comprehensive evaluation of model robustness under different perspectives, different backgrounds, and different environmental conditions. As a result, the generalization ability of existing attack methods is poor and they cannot show good attack effects in a changing real-world environment. When conducting adversarial attacks, they mainly focus on adversarial attacks in traditional visual tasks, resulting in the robustness of camera pose estimation models in practical applications not being effectively verified, and the reliability and security of the models cannot be fully guaranteed. Summary of the invention
[0005] In order to solve the technical problems that the existing adversarial robustness evaluation methods rely on a single adversarial sample and ignore the comprehensive evaluation of the model robustness under different perspectives, different backgrounds and different environmental conditions, resulting in poor generalization ability of the existing attack methods and failure to show good attack effects in a changeable real environment, and the adversarial attacks mainly focus on adversarial attacks in traditional visual tasks, resulting in the robustness of the camera pose estimation model in practical applications not being effectively verified and the reliability and security of the model being unable to be fully guaranteed, the present invention provides an adversarial robustness evaluation method and system based on a camera pose estimation model.
[0006] The technical solution provided by the embodiment of the present invention is as follows:
[0007] First aspect:
[0008] An embodiment of the present invention provides an adversarial robust evaluation method based on a camera pose estimation model, comprising:
[0009] S1: Initialize the fragment image to obtain the initial fragment image;
[0010] S2: Randomly select 3D objects and background environment;
[0011] S3: Projecting the 3D object at different projection angles to obtain projection images at different angles;
[0012] S4: Calculate the projection direction consistency loss based on the projection images at different viewing angles;
[0013] S5: Based on the projection direction consistency loss, the projection direction consistency of the initial segment image is optimized to obtain an optimized segment image;
[0014] S6: constructing a disk background image according to the optimized segment image;
[0015] S7: Use the disk background image as an adversarial sample to perform adversarial robust evaluation on the camera pose estimation model.
[0016] Second aspect:
[0017] An embodiment of the present invention provides an adversarial robust evaluation system based on a camera pose estimation model, comprising:
[0018] processor;
[0019] A memory having computer-readable instructions stored therein, wherein when the computer-readable instructions are executed by the processor, the adversarial robust evaluation method based on the camera pose estimation model according to the first aspect is implemented.
[0020] The third aspect:
[0021] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the adversarial robust evaluation method based on a camera pose estimation model as described in the first aspect is implemented.
[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0023] In an embodiment of the present invention, by calculating the projection direction consistency loss and optimizing the fragment images, the fragment images at different viewing angles have a high degree of consistency, and the spatial information cannot be accurately estimated, thereby causing misleading in the camera estimation process, improving the adversarial attack effect, and making the disk background image have significant interference capabilities in both digital and physical environments, providing an accurate evaluation method for camera pose estimation at different viewing angles, ensuring the reliability evaluation of the model under variable viewing angles and environmental conditions, and constructing a disk background image with a complex texture structure and using the disk background as an adversarial sample, which can effectively interfere with the camera pose estimation model, causing errors when processing images from different viewing angles, effectively interfering with the output of the camera pose estimation model, and enhancing the robustness of the adversarial attack. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 A schematic diagram of a flow chart of an adversarial robust evaluation method based on a camera pose estimation model provided by an embodiment of the present invention;
[0026] Figure 2 A schematic diagram of calculating the coordinate flow direction and the average flow direction provided by an embodiment of the present invention;
[0027] Figure 3 A schematic diagram of projection direction consistency loss provided by an embodiment of the present invention;
[0028] Figure 4 A schematic diagram of a disk background image provided by an embodiment of the present invention;
[0029] Figure 5 A schematic diagram of the structure of an adversarial robust evaluation system based on a camera pose estimation model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0031] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0032] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0033] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.
[0034] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0035] Reference Manual Attached Figure 1 , shows a flow chart of an adversarial robust evaluation method based on a camera pose estimation model provided by an embodiment of the present invention.
[0036] An embodiment of the present invention provides an adversarial robust evaluation method based on a camera pose estimation model, the method comprising:
[0037] S1: Initialize the segment image to obtain an initial segment image.
[0038] In a possible implementation manner, S1 specifically includes:
[0039] The segment image is initialized using uniformly distributed random noise to obtain an initial segment image.
[0040] Among them, uniformly distributed random noise refers to randomly generated values within a certain range, and the probability of occurrence of each value is equal. Mathematically, uniformly distributed random noise means that the probability distribution density of each value is the same.
[0041] It should be noted that using uniformly distributed random noise to initialize the fragment image can ensure that the generated initial image does not have any preset structure or bias, breaking the prior assumptions in the adversarial sample generation process and providing a completely free starting point for the optimization process.
[0042] S2: Randomly select 3D objects and background environment.
[0043] It should be noted that randomly selecting 3D objects and background environments can increase the diversity and complexity of generated images, thereby making the generated adversarial samples more universal, effectively avoiding overfitting of the model to a specific environment, and enhancing the interference effect of adversarial samples in various scenarios.
[0044] S3: Project the 3D object at different projection viewing angles to obtain projection images at different viewing angles.
[0045] In a possible implementation, S3 is specifically:
[0046] Under different projection perspectives, the 3D object is projected using a differentiable renderer to obtain projection images under different perspectives.
[0047] Among them, the Differentiable Renderer is a rendering technology that can calculate the gradient information of image pixels during the rendering process, making the image generation process differentiable.
[0048] It should be noted that this step can simulate the process of a camera shooting 3D objects from multiple different angles, generate images with diverse perspectives, increase the complexity and diversity of adversarial samples, and make the model misjudge images from different angles, thereby enhancing the effect of adversarial attacks. At the same time, the differentiable renderer can accurately control the changes in the image, making the adversarial optimization process more efficient and flexible.
[0049] Reference Manual Attached Figure 2 , showing a schematic diagram of calculating the coordinate flow direction and the average flow direction provided by an embodiment of the present invention.
[0050] like Figure 2 , the left figure selects two sub-areas M from the disk area M 1 and M 2 , and is divided by straight line l, the figure shows m 1 and m 2 The coordinate change and coordinate flow direction are calculated. The figure on the right shows the three directions of flow calculated after considering the influence of occlude, and the coordinate flow direction is calculated based on these directions. Figure 2 It shows how to calculate the coordinate change from the pixel coordinates of the local area and finally get the flow direction, so as to calculate the average flow direction and the projection direction consistency loss.
[0051] Reference Manual Attached Figure 3 , showing a schematic diagram of projection direction consistency loss provided by an embodiment of the present invention.
[0052] like Figure 3 The upper part shows the images rendered from different perspectives (viewa and viewb). The corresponding pointmaps are generated by the DUST3R model. The bottom part shows the three directions of flow τ 1 , τ 2 and τ 3 Corresponding to the coordinate flow direction τ under different viewing angles a and τ b , these flow directions are used to measure the projection consistency of the image under different viewing angles. a and τ b It represents the coordinate flow direction of the corresponding perspective. When calculating the projection direction consistency loss, the image is optimized by comparing the cosine similarity of these flow directions. Each small image in the figure represents the flow direction in a different direction to help with consistency calculation and ensure that images generated from different perspectives have similar directionality.
[0053] S4: Calculate the projection direction consistency loss based on the projection images at different viewing angles.
[0054] Among them, the projection direction consistency loss is a loss function that measures whether the direction of the projected image from different perspectives is consistent. It calculates the consistency of images from different perspectives in the spatial direction. The optimization goal is to make images from different perspectives have similar flow directions.
[0055] It should be noted that by calculating the projection direction consistency loss, the consistency of images generated from different perspectives in the spatial direction can be ensured, thereby increasing the interference effect of adversarial samples, and effectively reducing the model's dependence on specific perspectives. The camera pose estimation model performs consistently under different perspectives, thereby enhancing the robustness of the model.
[0056] In a possible implementation, S4 specifically includes:
[0057] S401: extracting a point map of the projection image through the DUST3R model.
[0058] Among them, the DUST3R model is a deep learning model that is specifically used to extract three-dimensional spatial information from images. By inputting images from different perspectives, DUST3R can generate corresponding pointmaps, that is, the position of each pixel in three-dimensional space. Pointmaps are an image representation method that contains spatial position information, and each pixel represents the position of a point in three-dimensional space. Pointmaps describe the spatial layout of objects by mapping the pixels of two-dimensional images into a three-dimensional coordinate system.
[0059] In the present invention, the DUSt3R model inputs H×W×3 image pairs from two viewpoints a and b into the model and regresses the point map O in the camera coordinate system a. a and O b , and then all point maps are fused and regressed into the same coordinate system through a global alignment strategy, and the camera pose is estimated.
[0060] S402: Mapping pixel coordinates of the point image to camera coordinates:
[0061]
[0062] in, represents the camera coordinates, and Both represent pixel coordinates, Φ 1 , Φ 2 and Φ 3 Both represent mapping functions, O i represents the i-th channel of the dot graph, H and W represent the height and degree of the dot graph respectively.
[0063] S403: Determine the coordinate change of the projected image in the disk area.
[0064] S404: Determine the direction of the coordinate flow according to the coordinate change:
[0065]
[0066] Among them, δ i represents the coordinate change of the ith channel, l represents the straight line dividing the disk area, l′ represents the straight line perpendicular to the straight line l, τ i represents the coordinate flow direction of the i-th channel, represents the normal vector of line l, Represents the normal vector of the line l′.
[0067] S405: Calculate the average flow direction according to the coordinate flow direction:
[0068]
[0069] in, represents the average flow direction of the ith channel, and Respectively represent l j and l j ′, l j represents the jth straight line, l j ′ indicates perpendicular to the line l j The straight line is j=1,2,3.
[0070] S406: Calculate the projection direction consistency loss based on the average flow direction:
[0071]
[0072] Among them, L poc represents the projection direction consistency loss function, and They represent the average flow directions of view a and view b in the i-th channel, i=1,2,3 respectively.
[0073] In a possible implementation, the calculation formula of the coordinate change is specifically:
[0074]
[0075] Among them, δ i (l) indicates M 1 and M 2 The average coordinate difference between i (l′) represents the average coordinate difference between the regions divided by the line l′, Φ i represents the i-th component of Φ, m 1 and m 2 They respectively represent the region M 1 and region M 2 Pixels, M 1 and M 2 It indicates that the straight line l divides the disk area M into two areas, and l indicates the straight line dividing the disk area.
[0076] In the present invention, a point graph O of any viewing angle is selected. i For example, the coordinates of all pixels in the disk area are defined as a set M. A straight line l passing through the center of the disk divides M into two parts, namely M 1 and M 2 , define M 1 and M 2 The coordinate change between i (l), for another straight line l' that is perpendicular to l in the disk, its coordinate change δi(l') can be calculated similarly.
[0077] S5: Based on the projection direction consistency loss, the projection direction consistency of the initial segment image is optimized to obtain an optimized segment image.
[0078] It should be noted that by optimizing the initial fragment image based on the projection direction consistency loss, it can ensure that the image exhibits consistent directionality under multiple viewing angles, which not only improves the spatial consistency of the image under different viewing angles, but also enhances the effect of adversarial samples, making it more difficult for the camera pose estimation model to correctly identify the true pose of the image.
[0079] In a possible implementation, S5 is specifically:
[0080] According to the following formula, the projection direction consistency of the initial fragment image is optimized to obtain the optimized fragment image:
[0081]
[0082] in, represents the initial fragment image after the t+1th optimization, clip represents the cropping operation, represents the initial fragment image after the t-th optimization, α represents the learning rate, sign represents the sign function, and ▽ represents the gradient operator.
[0083] Reference Manual Attached Figure 4 , showing a schematic diagram of a disk background image provided by an embodiment of the present invention.
[0084] like Figure 4 First, we perform grid sampling and perspective transformation operations on the image to transform the fragment image I s The different parts of the image are mapped to the disk area, and then the fragment image is processed using perspective projection to obtain a multi-layer image structure. Finally, the final disk background image I is generated by element-by-element addition. d , the image presents a repeated pattern, forming a visual effect with a certain symmetry. This method enables the background image to effectively interfere with the camera pose estimation model in adversarial attacks.
[0085] S6: constructing a disk background image according to the optimized segment image.
[0086] It should be noted that the image is mapped to the disk area using projection transformation, and the interference effect on the camera pose estimation model is enhanced through gradual image fusion, which effectively improves the robustness against attacks and makes the optimized background image remain consistent under different viewing angles.
[0087] In a possible implementation, the size of the segment image is optimized as follows:
[0088]
[0089] Wherein, θ represents the angle of the optimized segment image, N represents the total number of optimized segment images, h represents the height of the optimized segment image, ρ represents the radius of the disk background image, w represents the width of the optimized segment image, and π represents the ratio of pi.
[0090] In a possible implementation, S6 specifically includes:
[0091] S601: Determine the vertex coordinates of the optimized fragment image:
[0092]
[0093] Wherein, (A′, B′, C′, D′) represents the vertex coordinates of the optimized fragment image, and h and w represent the height and width of the optimized fragment image, respectively.
[0094] S602: Projecting the optimized segment image to the corresponding disk position through perspective projection to obtain a disk projection image.
[0095] S603: Determine the vertex coordinates of the disk projection image.
[0096] S604: Determine a perspective transformation matrix according to the vertex coordinates of the optimized segment image and the disk projection image.
[0097] S605: Determine the disk projection image according to the perspective transformation matrix:
[0098]
[0099] in, represents the nth disk projection image, G represents the image transformation function, P n represents the projected image, I d Represents the disk background image, -1 Represents the reverse operation.
[0100] S606: Superimpose each segment of the projection image to generate a disk background image:
[0101]
[0102] Among them, I d Represents a kaleidoscope background image, Represents the nth segment of the disk projection image, n = 0, 1, ..., N-1, and N represents the total number of segments of the projection image.
[0103] In a possible implementation manner, the calculation formula for the vertex coordinates of the disk projection image is specifically:
[0104]
[0105]
[0106] Among them, (A n ,B n ,C n ,D n ) represents the vertex coordinates of the projected image, ρ 1 and ρ 2 Represents the rectangular area A after the kaleidoscope fragment image is mappedn B n C n D n The two radii of vertex A n and B n Located at a radius of ρ 1 On the circle, and C n and D n Located at a radius of ρ 2 On the circle, nθ represents the angle of rotation around the center of the kaleidoscope disk when the kaleidoscope segment image is converted into a projection image, and β represents the offset for adjusting the rotation angle.
[0107] S7: Use the disk background image as an adversarial sample to perform adversarial robust evaluation on the camera pose estimation model.
[0108] It should be noted that by using complex and optimized adversarial samples (disc background images) to test the robustness of the camera pose estimation model, the security of the model can be enhanced, making it more reliable and robust in practical applications.
[0109] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0110] In an embodiment of the present invention, by calculating the projection direction consistency loss and optimizing the fragment images, the fragment images at different viewing angles have a high degree of consistency, and the spatial information cannot be accurately estimated, thereby causing misleading in the camera estimation process, improving the adversarial attack effect, and making the disk background image have significant interference capabilities in both digital and physical environments, providing an accurate evaluation method for camera pose estimation at different viewing angles, ensuring the reliability evaluation of the model under variable viewing angles and environmental conditions, and constructing a disk background image with a complex texture structure and using the disk background as an adversarial sample, which can effectively interfere with the camera pose estimation model, causing errors when processing images from different viewing angles, effectively interfering with the output of the camera pose estimation model, and enhancing the robustness of the adversarial attack.
[0111] Reference Manual Attached Figure 5 , showing a structural schematic diagram of an adversarial robust evaluation system based on a camera posture estimation model provided by the present invention.
[0112] The present invention also provides an adversarial robust evaluation system 20 based on a camera pose estimation model, which is applied to the above-mentioned adversarial robust evaluation method based on a camera pose estimation model, comprising:
[0113] Processor 201.
[0114] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the adversarial robust evaluation method based on the camera pose estimation model in the method embodiment is implemented.
[0115] The camera pose estimation model-based adversarial robust evaluation system 20 provided in the present invention can execute the above-mentioned camera pose estimation model-based adversarial robust evaluation method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.
[0116] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0117] In an embodiment of the present invention, by calculating the projection direction consistency loss and optimizing the fragment images, the fragment images at different viewing angles have a high degree of consistency, and the spatial information cannot be accurately estimated, thereby causing misleading in the camera estimation process, improving the adversarial attack effect, and making the disk background image have significant interference capabilities in both digital and physical environments, providing an accurate evaluation method for camera pose estimation at different viewing angles, ensuring the reliability evaluation of the model under variable viewing angles and environmental conditions, and constructing a disk background image with a complex texture structure and using the disk background as an adversarial sample, which can effectively interfere with the camera pose estimation model, causing errors when processing images from different viewing angles, effectively interfering with the output of the camera pose estimation model, and enhancing the robustness of the adversarial attack.
[0118] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0119] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0120] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0121] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0122] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0123] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0124] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0125] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0126] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0127] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0128] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0129] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0130] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the adversarial robust evaluation method based on a camera pose estimation model as described in the method embodiment is implemented.
[0131] A computer-readable storage medium provided by the present invention can implement the steps and effects of the adversarial robust evaluation method based on the camera posture estimation model of the above-mentioned method embodiment. To avoid repetition, the present invention will not go into details.
[0132] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0133] In an embodiment of the present invention, by calculating the projection direction consistency loss and optimizing the fragment images, the fragment images at different viewing angles have a high degree of consistency, and the spatial information cannot be accurately estimated, thereby causing misleading in the camera estimation process, improving the adversarial attack effect, and making the disk background image have significant interference capabilities in both digital and physical environments, providing an accurate evaluation method for camera pose estimation at different viewing angles, ensuring the reliability evaluation of the model under variable viewing angles and environmental conditions, and constructing a disk background image with a complex texture structure and using the disk background as an adversarial sample, which can effectively interfere with the camera pose estimation model, causing errors when processing images from different viewing angles, effectively interfering with the output of the camera pose estimation model, and enhancing the robustness of the adversarial attack.
[0134] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
[0135] There are a few points to note:
[0136] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.
[0137] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.
[0138] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.
[0139] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A robust adversarial evaluation method based on a camera pose estimation model, characterized in that: include: S1: Initialize the fragment image to obtain the initial fragment image; S2: Randomly select 3D objects and background environment; S3: Projecting the 3D object at different projection viewing angles to obtain projection images at different viewing angles; S4: Calculate the projection direction consistency loss based on the projection images at different viewing angles; S5: Based on the projection direction consistency loss, optimizing the projection direction consistency of the initial segment image to obtain an optimized segment image; S6: constructing a disk background image according to the optimized segment image; S7: Using the disk background image as an adversarial sample, and performing adversarial robust evaluation on the camera pose estimation model.
2. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The S1 is specifically: Initialize the fragment image using uniformly distributed random noise.
3. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The S3 is specifically: Under different projection viewing angles, the 3D object is projected using a differentiable renderer to obtain projection images under different viewing angles.
4. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The S4 specifically includes: S401: extracting a point map of the projection image through a DUST3R model; S402: Mapping the pixel coordinates of the point image to camera coordinates: in, represents the camera coordinates, and All represent pixel coordinates, Φ1, Φ2 and Φ3 all represent mapping functions, i represents the i-th channel of the dot graph, H and W represent the height and width of the dot graph respectively; S403: Determine the coordinate change of the projection image in the disk area; S404: Determine the direction of the coordinate flow according to the coordinate change amount: t i =d i (l)ü+d i (l′)ü′ Among them, δ i represents the coordinate change of the ith channel, l represents the straight line dividing the disk area, l′ represents the straight line perpendicular to the straight line l, τ i represents the coordinate flow direction of the i-th channel, ü represents the normal vector of the line l, and u′ represents the normal vector of the line l′; S405: Calculate the average flow direction according to the coordinate flow direction: in, represents the average flow direction of the ith channel, u j and u j ′ respectively represents l j and l j ′, l j represents the jth straight line, l j ′ indicates perpendicular to the line l j The straight line, j = 1, 2, 3; S406: Calculate the projection direction consistency loss according to the average flow direction: Among them, L poc represents the projection direction consistency loss function, and They represent the average flow directions of view a and view b in the i-th channel, i=1,2,3 respectively.
5. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 4, characterized in that: The calculation formula of the coordinate change is specifically: Among them, δ i (l) represents the average coordinate difference between M1 and M2, δ i (l′) represents the average coordinate difference between the regions divided by the line l′, Φ i represents the i-th component of Φ, m1 and m2 represent pixels belonging to region M1 and region M2 respectively, M1 and M2 represent two regions into which the disk region M is divided by the straight line l, and l represents the straight line dividing the disk region.
6. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The S5 is specifically: According to the following formula, the projection direction consistency of the initial segment image is optimized to obtain an optimized segment image: in, represents the initial fragment image after the t+1th optimization, clip represents the cropping operation, represents the initial fragment image after the tth optimization, α represents the learning rate, sign represents the sign function, Represents the gradient operator.
7. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The size of the optimized segment image is specifically: Wherein, θ represents the angle of the optimized segment image, N represents the total number of optimized segment images, h represents the height of the optimized segment image, ρ represents the radius of the disk background image, w represents the width of the optimized segment image, and π represents the ratio of pi.
8. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 1, characterized in that: The S6 specifically includes: S601: Determine the vertex coordinates of the optimized fragment image: Where (A′, B′, C′, D′) represents the vertex coordinates of the optimized fragment image, h and w represent the height and width of the optimized fragment image respectively; S602: Projecting the optimized segment image to the corresponding disk position through perspective projection to obtain a disk projection image; S603: Determine the vertex coordinates of the disk projection image; S604: determining a perspective transformation matrix according to the vertex coordinates of the optimized segment image and the disk projection image; S605: Determine the disk projection image according to the perspective transformation matrix: in, represents the nth disk projection image, G represents the image transformation function, P n represents the projected image, I d Represents the disk background image, -1 Indicates the reverse operation; S606: Superimpose each segment of the projection image to generate a disk background image: Among them, I d Represents a kaleidoscope background image, Represents the nth segment of the disk projection image, n = 0, 1, ..., N-1, and N represents the total number of segments of the projection image.
9. The method for robust adversarial evaluation based on a camera pose estimation model according to claim 8, characterized in that: The calculation formula of the vertex coordinates of the disk projection image is specifically: Among them, (A n ,B n ,C n ,D n ) represents the vertex coordinates of the projected image, ρ1 and ρ2 represent the rectangular area A after the kaleidoscope fragment image is mapped n B n C n D n The two radii of vertex A n and B n is located on a circle with radius ρ1, and C n and D n Located on a circle with a radius of ρ2, nθ represents the angle of rotation around the center of the kaleidoscope disk when the kaleidoscope segment image is converted into a projection image, and β represents the offset for adjusting the rotation angle.
10. An adversarial robust evaluation system based on a camera pose estimation model, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the adversarial robust evaluation method based on the camera pose estimation model as described in any one of claims 1 to 9 is implemented.