Visual three-dimensional perception method and device of look-around camera suitable for complex light environment

By capturing multi-perspective images with a surround-view camera and performing illumination enhancement processing, the robot's unstable visual perception caused by light changes in a multi-floor environment is resolved, enabling high-precision three-dimensional perception and path planning, and improving its autonomous navigation capabilities.

CN120766239APending Publication Date: 2025-10-10BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510741807.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In complex multi-floor environments, robot visual perception methods are susceptible to light interference, resulting in poor perception accuracy and stability, especially in transition areas where light changes drastically, affecting autonomous navigation and safety.

Method used

The surround-view camera captures multi-view images, determines image features, generates illumination prior maps and illumination enhancement features, and uses the light environment difference constraint function between the surround-view cameras to balance and generate target image features containing illumination information. The three-dimensional information is input into the decoding network for path planning.

Benefits of technology

It improves the robot's visual perception stability and accuracy in complex lighting environments, and enhances its autonomous navigation capability and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766239A_ABST
    Figure CN120766239A_ABST
Patent Text Reader

Abstract

The invention discloses an all-around camera visual three-dimensional perception method and device suitable for a complex light environment, and belongs to the technical field of intelligent hardware, and the method comprises the steps: shooting a multi-view image through an all-around camera on a robot; according to the multi-view-angle image, determining image features of the surround-view camera; carrying out channel-dimension pixel averaging operation on the multi-view image to obtain an illumination priori image; generating illumination enhancement features according to the multi-view image and the illumination prior image; the illumination enhancement features are processed according to the overall light environment difference constraint function between the surround view cameras, and balanced illumination enhancement features are obtained; according to the illumination enhancement feature and the look-around camera image feature, generating a target image feature including illumination information; inputting the target illumination image features into a decoding network to obtain three-dimensional information of the to-be-detected target; according to the three-dimensional information of the to-be-detected target, the robot motion path is planned, and the stability and precision of visual perception of the intelligent robot can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent hardware technology, and in particular to a surround-view camera three-dimensional visual perception method and device, and electronic equipment suitable for complex lighting environments. Background Art

[0002] In the field of intelligent robotics, visual perception technology is a crucial enabler for autonomous navigation, environmental understanding, and human-machine interaction. It is widely used in tasks such as human perception and static obstacle detection. However, in complex, multi-story environments, robots face highly variable lighting conditions. Traditional visual perception methods are susceptible to interference from ambient light, impacting their perception accuracy and stability.

[0003] In a multi-floor environment, robots need to have precise visual perception capabilities to ensure safe movement and efficient operation. However, the lighting conditions on different floors may vary significantly, especially in transition areas such as corridors, hallways, and stairwells, where the light changes dramatically. For example, some corridors rely on voice-activated lighting. When the lighting equipment is not triggered, the robot still needs to have reliable visual perception capabilities to avoid collisions and operational interruptions. In addition, because the imaging quality of different cameras in the surround view camera varies in different lighting environments, the robot needs to have the ability to fuse and adaptively adjust multi-channel visual information to improve the robustness and accuracy of perception and ensure stable operation under complex lighting conditions.

[0004] Therefore, in order to address the complex and changeable lighting problems in multi-floor environments, technical personnel in this field are urgently needed to provide a three-dimensional perception method that can adapt to different lighting conditions and improve the stability of visual perception, so as to enhance the autonomous navigation capability and safety of intelligent robots in complex lighting environments. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a surround-view camera visual three-dimensional perception method and device, and electronic equipment suitable for complex lighting environments, which can solve the problem of poor perception stability under complex lighting conditions existing in the prior art.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] An embodiment of the present invention provides a surround-view camera three-dimensional perception method applicable to complex lighting environments, which is applied to a robot. The method includes:

[0008] Capturing multi-perspective images containing information about the robot's surrounding environment by using a surround-view camera on the robot;

[0009] determining the surround view camera image features based on the multi-view images;

[0010] performing a channel dimension pixel average operation on the multi-view image to obtain an illumination prior map;

[0011] generating illumination enhanced features according to the multi-view image and the illumination prior map;

[0012] processing the illumination enhanced features according to the overall light environment difference constraint function between the surround-view cameras to obtain balanced illumination enhanced features;

[0013] generating target image features containing illumination information according to the illumination enhanced features and the surround-view camera image features;

[0014] inputting the target illumination image features into a decoding network to obtain three-dimensional information of a target to be detected;

[0015] performing robot motion path planning according to the three-dimensional information of the target to be detected.

[0016] Optionally, the step of determining the surround-view camera image features according to the multi-view image comprises:

[0017] inputting the multi-view image into a three-dimensional detection network;

[0018] performing convolution processing on the multi-view image by a backbone network of the three-dimensional detection network to obtain the surround-view camera image features.

[0019] Optionally, the step of generating illumination enhanced features according to the multi-view image and the illumination prior map comprises:

[0020] performing channel-level connection fusion on the multi-view image and the illumination prior map to generate a fusion image;

[0021] inputting the fusion image into a convolution layer for fusion to obtain fusion image features;

[0022] inputting the fusion image features into a deep convolution network layer to generate illumination enhanced features.

[0023] Optionally, the step of processing the illumination enhanced features according to the overall light environment difference constraint function between the surround-view cameras to obtain balanced illumination enhanced features comprises:

[0024] determining the width of an adjacent region between any two adjacent cameras in the surround-view cameras;

[0025] performing normalization processing on the abscissas of pixel points on the adjacent region between the adjacent cameras;

[0026] constructing a loss function weight based on the normalized pixel coordinates on the adjacent region;

[0027] Based on the loss function weight and the mean square error function, constructing the overall light environment difference constraint function between the surround view cameras;

[0028] Based on the overall light environment difference constraint function between the surround view cameras, the illumination enhancement feature is processed to obtain a balanced illumination enhancement feature.

[0029] Optionally, the step of generating a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature includes:

[0030] The illumination enhancement feature and the surround view camera image feature are added and fused to generate a target image feature containing illumination information.

[0031] An embodiment of the present invention further provides a surround-view camera visual three-dimensional perception device suitable for complex lighting environments, which is applied to a robot. The device includes:

[0032] An image acquisition module is used to call the surround view camera on the robot to capture multi-view images containing information about the robot's surrounding environment;

[0033] a feature determination module, configured to determine the surround view camera image features based on the multi-view images;

[0034] An averaging module, configured to perform a pixel averaging operation in a channel dimension on the multi-view images to obtain an illumination prior map;

[0035] A first generating module, configured to generate an illumination enhancement feature based on the multi-view image and the illumination prior map;

[0036] a constraint processing module, configured to process the illumination enhancement feature according to the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature;

[0037] A second generating module is configured to generate a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature;

[0038] A three-dimensional information determination module is used to input the target illumination image features into a decoding network to obtain three-dimensional information of the target to be detected;

[0039] The planning module is used to plan the robot's motion path based on the three-dimensional information of the target to be detected.

[0040] Optionally, the feature determination module includes:

[0041] An input submodule, configured to input the multi-view images into a three-dimensional detection network;

[0042] The convolution processing submodule is used to perform convolution processing on the multi-view image through the backbone network of the three-dimensional detection network to obtain the surround camera image features.

[0043] Optionally, the first generating module includes:

[0044] The first submodule is configured to perform channel-level concatenation and fusion of the multi-view image and the illumination prior map to generate a fused image;

[0045] The second submodule is used to input the fused image into the convolution layer for fusion to obtain fused image features;

[0046] The third submodule is used to input the fused image features into the deep convolutional network layer to generate illumination enhancement features.

[0047] Optionally, the constraint processing module includes:

[0048] a width determination submodule, configured to determine, for any two adjacent cameras in the surround view cameras, a width of an adjacent area between the adjacent cameras;

[0049] A normalization processing submodule, configured to perform normalization processing on the horizontal coordinates of the pixel points in the adjacent areas between the adjacent cameras;

[0050] A weight construction submodule, configured to construct a loss function weight based on the normalized pixel coordinates of the adjacent regions;

[0051] A function construction submodule, configured to construct a constraint function for the overall light environment difference between the surround view cameras based on the loss function weight and the mean square error function;

[0052] The constraint processing submodule is used to process the illumination enhancement feature based on the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature.

[0053] Optionally, the second generating module is specifically configured to:

[0054] The illumination enhancement feature and the surround view camera image feature are added and fused to generate a target image feature containing illumination information.

[0055] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of any one of the above-mentioned surround view camera visual three-dimensional perception methods applicable to complex lighting environments are implemented.

[0056] An embodiment of the present invention provides a readable storage medium storing a program or instruction. When the program or instruction is executed by a processor, the program or instruction implements the steps of any one of the above-mentioned surround view camera three-dimensional visual perception methods applicable to complex lighting environments.

[0057] The embodiments of the present invention provide a surround-view camera visual three-dimensional perception solution suitable for complex lighting environments. The solution uses a surround-view camera on a robot to capture multi-perspective images containing information about the robot's surrounding environment. The surround-view camera determines surround-view camera image features based on the multi-perspective images. A pixel average operation in the channel dimension is performed on the multi-perspective images to obtain a lighting prior map. Lighting enhancement features are generated based on the multi-perspective images and the lighting prior map. The lighting enhancement features are processed based on a constraint function for the overall lighting environment differences between the surround-view cameras to obtain balanced lighting enhancement features. Target image features containing lighting information are generated based on the lighting enhancement features and surround-view camera image features. The target lighting image features are input into a decoding network to obtain three-dimensional information of the target to be detected. The robot's motion path is planned based on the three-dimensional information of the target to be detected. The solution provided by the embodiments of the present application, on the one hand, can effectively balance the inconsistent lighting representations from different perspectives of the surround-view camera, improving the stability and accuracy of overall perception. On the other hand, the solution is applicable to three-dimensional perception in complex lighting environments using surround-view cameras. Because the three-dimensional information of the target can be determined stably and accurately, it can improve the visual perception capabilities of autonomous robots in complex lighting environments with multiple floors. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart showing the steps of a surround view camera visual three-dimensional perception method applicable to complex lighting environments according to an embodiment of the present application;

[0059] Figure 2 is a flowchart showing the steps of a surround view camera visual three-dimensional perception method applicable to complex lighting environments according to an embodiment of the present application;

[0060] Figure 3 It is a structural block diagram of a surround-view camera visual three-dimensional perception device suitable for complex lighting environments according to an embodiment of the present application. DETAILED DESCRIPTION

[0061] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0062] The following, in conjunction with the accompanying drawings, describes in detail the surround view camera visual three-dimensional perception solution suitable for complex lighting environments provided by the embodiment of the present application through specific embodiments and their application scenarios.

[0063] As attached Figure 1As shown, the surround view camera visual 3D perception method applicable to complex light environments in an embodiment of the present application includes the following steps:

[0064] Step 101: Use the surround view camera on the robot to capture multi-perspective images containing information about the robot's surrounding environment.

[0065] The embodiments of this application provide a surround-view camera visual 3D perception method suitable for complex lighting environments. Applications include, but are not limited to, guide robots, warehouse robots, and hotel service robots. For example, a guide robot (also referred to as a guide service robot, hereinafter referred to as a robot) is an omnidirectional mobile robot capable of free movement in the x- and y-planes and z-axis rotation. The guide robot is equipped with a control module (also referred to as a robot processor) to provide transfer route planning guidance for users.

[0066] The robot uses sensory sensors to perceive its surroundings and state. These sensors may include, but are not limited to, lidar, cameras, and inertial measurement units. In this embodiment, a surround-view camera is installed on the robot to perceive its surroundings using a pre-set 3D visual perception algorithm.

[0067] Step 102: Determine surround view camera image features based on the multi-view images.

[0068] The surround-view camera can capture multiple images from different perspectives at the same time. These images are called multi-perspective images.

[0069] In an optional embodiment, the method for determining the surround-view camera image features based on the multi-view images can be: inputting the multi-view images into a three-dimensional detection network; and performing convolution processing on the multi-view images through the backbone network of the three-dimensional detection network to obtain the surround-view camera image features.

[0070] Step 103: Perform a pixel averaging operation in the channel dimension on the multi-view images to obtain a priori illumination map.

[0071] Step 104: Generate illumination enhancement features based on the multi-view images and the illumination prior map.

[0072] In an optional embodiment, the method of generating the illumination enhancement feature based on the multi-view images and the illumination prior map may include the following sub-steps:

[0073] Sub-step 1: perform channel-level concatenation and fusion of the multi-view images and the illumination prior map to generate a fused image;

[0074] Sub-step 2: Input the fused image into the convolution layer for fusion to obtain the fused image features;

[0075] In the specific implementation process, the fused image can be input into 1×1 convolution for fusion.

[0076] Sub-step 3: Input the fused image features into the deep convolutional network layer to generate illumination enhancement features.

[0077] Preferably, the deep convolutional network layer can be a 5×5 convolution, but is not limited thereto and can also be a 4×4 convolution or a 3×3 convolution.

[0078] Step 105: Process the illumination enhancement feature according to the overall light environment difference constraint function between surround view cameras to obtain a balanced illumination enhancement feature.

[0079] In an optional embodiment, the lighting enhancement feature is processed according to the overall lighting environment difference constraint function between surround view cameras to obtain a balanced lighting enhancement feature, which may include the following sub-steps:

[0080] Sub-step 1: for any two adjacent cameras in the surround view camera, determine the width of the adjacent area between the adjacent cameras;

[0081] Sub-step 2: normalize the horizontal coordinates of the pixels in the adjacent areas between adjacent cameras;

[0082] Sub-step 3: Construct the loss function weights based on the normalized pixel coordinates of the adjacent regions;

[0083] Sub-step 4: Based on the loss function weight and the mean square error function, construct the overall light environment difference constraint function between the surround view cameras;

[0084] Sub-step 5: Based on the overall light environment difference constraint function between surround view cameras, the illumination enhancement feature is processed to obtain a balanced illumination enhancement feature.

[0085] This optional way of generating illumination enhancement features makes the generated illumination enhancement features more reliable and accurate.

[0086] Step 106: Generate target image features containing lighting information based on the lighting enhancement features and the surround view camera image features.

[0087] In the actual implementation process, the illumination enhancement features and the surround camera image features can be added and fused to generate the target image features containing illumination information.

[0088] It should be noted that the illumination enhancement features and the surround camera image features may also be assigned corresponding weights and then added and fused. The specific values ​​of the weights are not limited in the embodiments of the present application.

[0089] Step 107: Input the target illumination image features into the decoding network to obtain the three-dimensional information of the target to be detected.

[0090] Step 108: According to the three-dimensional information of the to-be-detected target, the motion path of the robot is planned.

[0091] The to-be-detected target can be one or more. The visual three-dimensional perception method shown in the embodiment of the present application is executed in real time during the operation of the robot to determine the to-be-detected target. According to the three-dimensional information of the to-be-detected target recognized, obstacle avoidance and motion path planning are performed. The robot moves according to the planned motion path.

[0092] The visual three-dimensional perception method for surround-view cameras suitable for complex light environments provided by the embodiment of the present application comprises the following steps:

[0093] The visual three-dimensional perception method for surround-view cameras suitable for complex light environments provided by the embodiment of the present application will be described below with reference to a specific example. Figure 2 The visual three-dimensional perception method for surround-view cameras suitable for complex light environments provided by the embodiment of the present application will be described below with reference to a specific example.

[0094] The visual three-dimensional perception method provided by the embodiment of the present application is suitable for multi-floor complex light environments, can effectively solve the visual perception problem of intelligent robots under different light conditions, improve the robustness and accuracy of perception, and ensure the stable operation of the robot in a complex light environment.

[0095] The visual three-dimensional perception method for surround-view cameras suitable for complex light environments in the specific example comprises the following steps:

[0096] S1, input the surround-view camera image.

[0097] Then, S2 and S3 are executed in parallel.

[0098] S2, the surround-view camera image enters the backbone network to extract features.

[0099] The surround-view camera is used to obtain real-time images of the surrounding environment, and the multi-view images are input into the 3D detection network. The surround-view camera image features are obtained through the backbone network convolution. .

[0100] S3. Average the pixels in the channel dimension of the surround camera image to obtain the illumination prior map.

[0101] Specifically, for the input multi-view image Perform pixel averaging in the channel dimension to obtain the illumination prior map .

[0102] S4, input multi-view image and lighting prior map Perform channel-level connection fusion.

[0103] S5, the fused image is passed through The convolution is fused and the illumination enhancement features are obtained through a deep convolutional network.

[0104] In this step, the connected fused image is passed through Convolution is further performed to obtain image features. The fused image features are further input into a Deep convolutional network layers to obtain illumination enhancement features .

[0105] S6. Perform illumination balance on illumination enhancement features from different perspectives through loss function.

[0106] Lighting enhancement features It includes illumination enhancement features from multiple perspectives, and designs a loss function for the illumination enhancement features from multiple perspectives to balance the differences in lighting environments between cameras.

[0107] Specifically, for any two adjacent cameras and , define the width of the adjacent area between them as ,in is the width of the input image. The horizontal coordinate of the pixel point on , coordinate normalization is performed, , For cameras Adjacent areas on Pixel coordinates on The weights applied to the loss function can be obtained ;For cameras Adjacent areas on Pixel coordinates on The weights applied to the loss function can be obtained .in is a small positive value to avoid . Mean squared error ( ) function is used to measure the difference between different pixels, using Function can obtain lighting enhancement features Differences in lighting characteristics between different viewing angles. The function is defined as , and Represents two different images. For adjacent images and The illumination difference between adjacent areas can be expressed as , and finally we can get the overall light environment difference constraint function between the surround cameras ,in Indicates the number of surround view cameras.

[0108] S7. Fuse the balanced illumination enhancement features with the surround camera image features.

[0109] The balanced illumination enhancement features are obtained through the above constraint function , the illumination enhancement feature Surround view camera image features with input image Perform additive fusion to obtain the final image features containing illumination information + .

[0110] S8, image features containing illumination information Input into the decoding network to obtain the final three-dimensional information output of the detected target.

[0111] S9. Input the three-dimensional information of the detected target into the planning and control module to further support the robot's decision-making.

[0112] The three-dimensional perception method in complex lighting environments based on a surround-view camera provided in this specific example can, on the one hand, enhance the visual perception capability of autonomous robots in complex lighting environments on multiple floors; on the other hand, the self-balancing mechanism of illumination information based on the surround-view camera can effectively balance inconsistent illumination representations under different perspectives, thereby improving the stability and accuracy of overall perception.

[0113] Figure 3 A block diagram of a surround-view camera visual three-dimensional perception device suitable for complex lighting environments is provided to implement an embodiment of the present application.

[0114] The embodiment of the present application provides a surround view camera visual three-dimensional perception device suitable for complex lighting environments, which is applied to a robot and includes the following functional modules:

[0115] An image acquisition module 301 is configured to call a surround view camera on the robot to capture multi-view images containing information about the robot's surrounding environment;

[0116] A feature determination module 302 is configured to determine the surround view camera image features based on the multi-view image;

[0117] An averaging module 303 is configured to perform a channel-dimensional pixel averaging operation on the multi-view images to obtain an illumination prior map;

[0118] A first generating module 304 is configured to generate an illumination enhancement feature based on the multi-view image and the illumination prior map;

[0119] A constraint processing module 305 is configured to process the illumination enhancement feature according to the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature;

[0120] A second generating module 306 is configured to generate a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature;

[0121] The three-dimensional information determination module 307 is used to input the target illumination image features into a decoding network to obtain the three-dimensional information of the target to be detected;

[0122] The planning module 308 is used to plan the robot's motion path based on the three-dimensional information of the target to be detected.

[0123] Optionally, the feature determination module includes:

[0124] An input submodule, configured to input the multi-view images into a three-dimensional detection network;

[0125] The convolution processing submodule is used to perform convolution processing on the multi-view image through the backbone network of the three-dimensional detection network to obtain the surround camera image features.

[0126] Optionally, the first generating module includes:

[0127] The first submodule is configured to perform channel-level concatenation and fusion of the multi-view image and the illumination prior map to generate a fused image;

[0128] The second submodule is used to input the fused image into the convolution layer for fusion to obtain fused image features;

[0129] The third submodule is used to input the fused image features into the deep convolutional network layer to generate illumination enhancement features.

[0130] Optionally, the constraint processing module includes:

[0131] a width determination submodule, configured to determine, for any two adjacent cameras in the surround view cameras, a width of an adjacent area between the adjacent cameras;

[0132] A normalization processing submodule, configured to perform normalization processing on the horizontal coordinates of the pixel points in the adjacent areas between the adjacent cameras;

[0133] A weight construction submodule, configured to construct a loss function weight based on the normalized pixel coordinates of the adjacent regions;

[0134] A function construction submodule, configured to construct a constraint function for the overall light environment difference between the surround view cameras based on the loss function weight and the mean square error function;

[0135] The constraint processing submodule is used to process the illumination enhancement feature based on the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature.

[0136] Optionally, the second generating module is specifically configured to:

[0137] The illumination enhancement feature and the surround view camera image feature are added and fused to generate a target image feature containing illumination information.

[0138] The embodiment of the present invention provides a surround-view camera visual three-dimensional perception device suitable for complex lighting environments. On the one hand, it can effectively balance the inconsistent lighting representations under different viewing angles of the surround-view camera, thereby improving the stability and accuracy of the overall perception. On the other hand, it can be applied to the three-dimensional perception of the surround-view camera in complex lighting environments. Since it can stably and highly accurately determine the three-dimensional information of the target, it can enhance the visual perception capability of autonomous robots in complex lighting environments on multiple floors.

[0139] In the embodiment of the present application Figure 3 The illustrated surround-view camera visual 3D perception device, suitable for complex lighting environments, is installed in the robot's control system. The operating system on which this control system is installed can be Android, iOS, or other possible operating systems, and this embodiment of the application does not specifically limit this.

[0140] The embodiments of the present application provide Figure 3 The visual three-dimensional perception device shown can achieve Figure 1 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0141] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the various processes performed by the above-mentioned surround-view camera visual three-dimensional perception device suitable for complex lighting environments are implemented, and the same technical effect can be achieved. To avoid repetition, they will not be described here.

[0142] It should be noted that the electronic device in the embodiment of the present application includes the server described above.

[0143] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk.

[0144] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0145] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A surround view camera three-dimensional perception method suitable for complex light environments, applied to robots, characterized by: The method comprises: Capturing multi-perspective images containing information about the robot's surrounding environment by using a surround-view camera on the robot; determining the surround view camera image features based on the multi-view images; Performing a pixel averaging operation in a channel dimension on the multi-view images to obtain a lighting prior map; generating an illumination enhancement feature based on the multi-view image and the illumination prior map; Processing the illumination enhancement feature according to the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature; generating a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature; Inputting the target illumination image features into a decoding network to obtain three-dimensional information of the target to be detected; The robot motion path is planned based on the three-dimensional information of the target to be detected.

2. The method according to claim 1, characterized in that The step of determining the surround view camera image features based on the multi-view images comprises: Inputting the multi-view images into a three-dimensional detection network; The multi-view image is convolved by the backbone network of the three-dimensional detection network to obtain the surround camera image features.

3. The method according to claim 1, characterized in that The step of generating illumination enhancement features based on the multi-view images and the illumination prior map comprises: Perform channel-level concatenation and fusion of the multi-view image and the illumination prior map to generate a fused image; Inputting the fused image into the convolution layer for fusion to obtain fused image features; The fused image features are input into a deep convolutional network layer to generate illumination enhancement features.

4. The method according to claim 1, wherein The step of processing the illumination enhancement feature according to the overall light environment difference constraint function between surround view cameras to obtain a balanced illumination enhancement feature includes: For any two adjacent cameras among the surround-view cameras, determining a width of an adjacent area between the adjacent cameras; Normalizing the horizontal coordinates of the pixels in the adjacent areas between the adjacent cameras; Constructing a loss function weight based on the normalized pixel coordinates of the adjacent region; Based on the loss function weight and the mean square error function, constructing the overall light environment difference constraint function between the surround view cameras; Based on the overall light environment difference constraint function between the surround view cameras, the illumination enhancement feature is processed to obtain a balanced illumination enhancement feature.

5. The method according to claim 1, wherein The step of generating a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature comprises: The illumination enhancement feature and the surround view camera image feature are added and fused to generate a target image feature containing illumination information.

6. A surround-view camera visual three-dimensional perception device suitable for complex light environments, applied to robots, characterized by: The device comprises: An image acquisition module is used to call the surround view camera on the robot to capture multi-view images containing information about the robot's surrounding environment; a feature determination module, configured to determine the surround view camera image features based on the multi-view images; An averaging module, configured to perform a pixel averaging operation in a channel dimension on the multi-view images to obtain an illumination prior map; A first generating module, configured to generate an illumination enhancement feature based on the multi-view image and the illumination prior map; a constraint processing module, configured to process the illumination enhancement feature according to the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature; A second generating module is configured to generate a target image feature containing illumination information based on the illumination enhancement feature and the surround view camera image feature; A three-dimensional information determination module is used to input the target illumination image features into a decoding network to obtain three-dimensional information of the target to be detected; The planning module is used to plan the robot's motion path based on the three-dimensional information of the target to be detected.

7. The device according to claim 6, characterized in that The feature determination module includes: An input submodule, configured to input the multi-view images into a three-dimensional detection network; The convolution processing submodule is used to perform convolution processing on the multi-view image through the backbone network of the three-dimensional detection network to obtain the surround view camera image features.

8. The device according to claim 6, characterized in that The first generation module includes: The first submodule is configured to perform channel-level concatenation and fusion of the multi-view image and the illumination prior map to generate a fused image; The second submodule is used to input the fused image into the convolution layer for fusion to obtain fused image features; The third submodule is used to input the fused image features into the deep convolutional network layer to generate illumination enhancement features.

9. The device according to claim 6, characterized in that The constraint processing module includes: a width determination submodule, configured to determine, for any two adjacent cameras in the surround view cameras, a width of an adjacent area between the adjacent cameras; A normalization processing submodule, configured to perform normalization processing on the horizontal coordinates of the pixel points in the adjacent areas between the adjacent cameras; A weight construction submodule, configured to construct a loss function weight based on the normalized pixel coordinates of the adjacent regions; A function construction submodule, configured to construct a constraint function for the overall light environment difference between the surround view cameras based on the loss function weight and the mean square error function; The constraint processing submodule is used to process the illumination enhancement feature based on the overall light environment difference constraint function between the surround view cameras to obtain a balanced illumination enhancement feature.

10. An electronic device, characterized in that: The electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction is executed by the processor to execute the steps of any one of claims 1-5 of the surround view camera visual three-dimensional perception method applicable to complex lighting environments.