Environment sensing method and system and photovoltaic robot

By installing multiple wide-angle cameras on photovoltaic robots, combining stereo matching algorithms and deep learning algorithms, the problem of redundancy in photovoltaic robots obtaining environmental information is solved, and information processing efficiency and working stability are improved.

CN119992435APending Publication Date: 2025-05-13LEAPTING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510066628.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing photovoltaic robots are equipped with multiple sensors, which leads to redundant environmental information obtained, which increases the information processing load and reduces the processing efficiency, and is not conducive to continuous inspection in high-temperature environments.

Method used

Multiple wide-angle cameras are used to obtain 360-degree environmental information, calculate the parallax through a stereo matching algorithm, convert it into a depth map, and process the depth map and environmental images using improved FlashOCC and YOLO algorithms to detect the environmental occupation information and the working status of the photovoltaic module.

Benefits of technology

It effectively reduces the redundant information obtained by the sensor, improves the processing efficiency of perceived information, reduces the production cost of photovoltaic robots, improves the efficiency of perceived information, and enhances the working stability in high-temperature environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992435A_ABST
    Figure CN119992435A_ABST
Patent Text Reader

Abstract

The invention discloses an environment sensing method and system and a photovoltaic robot, and the method comprises the steps: collecting a surrounding environment image based on a plurality of wide-angle cameras installed on the photovoltaic robot, and enabling the visual angles of the plurality of wide-angle cameras to be different; the parallax of the surrounding environment images collected by any two adjacent wide-angle cameras is calculated through a stereo matching algorithm; the parallax is converted into a depth map of the surrounding environment in combination with the camera internal reference and the baseline distance of the wide-angle camera, environment occupation information is obtained through detection, and the environment occupation information is point cloud information of environment occupation voxel particles; and processing the surrounding environment image by using a photovoltaic detection model, and detecting to obtain photovoltaic inspection information. According to the invention, redundant information acquired by the sensor can be effectively reduced. Under the condition that few sensors are used, the processing efficiency of the sensing information is improved, the production cost of the photovoltaic robot is reduced, and the use efficiency of the sensing information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of photovoltaic robots, and more specifically, to an environmental perception method, system and photovoltaic robot. Background Art

[0002] The photovoltaic robots currently used in the market are equipped with a large number of sensors of various types, and generally rely on multiple sensors such as laser radar and cameras to perceive the environment. Although the accuracy of patrol inspection can be improved, the production cost is increased, and there is a lot of duplication in the information provided by multiple sensors, which increases the workload of information processing during data processing, causing the robot to process slower and less efficient during work. This is not conducive to the photovoltaic robot's efficient processing of effective perception information, nor is it conducive to continuous patrol inspections in high temperature environments. Summary of the invention

[0003] In order to solve the technical problem that the photovoltaic robot obtains redundant information due to the above-mentioned multiple sensors, the present application provides an environmental perception method, system, and photovoltaic robot. Through the design of the present application, the redundant information obtained by the sensor can be effectively reduced. When using fewer sensors, the processing efficiency of the perceived information is improved, the production cost of the photovoltaic robot is reduced, and the efficiency of using the perceived information is improved. Specifically, the technical solution of the present application is as follows:

[0004] In a first aspect, the present application discloses an environment perception method, comprising:

[0005] Based on multiple wide-angle cameras installed on the photovoltaic robot, the surrounding environment image is collected, and the multiple wide-angle cameras have different viewing angles;

[0006] Calculate the disparity of the surrounding environment images captured by any two adjacent wide-angle cameras through a stereo matching algorithm; convert the disparity into a depth map of the surrounding environment by combining the camera intrinsic parameters and baseline distance of the wide-angle camera;

[0007] Using an improved occupancy detection model to process the depth map, and detect and obtain environmental occupancy information; the environmental occupancy information is point cloud information of environmental occupancy voxels, which is used for path planning and safe obstacle avoidance of the photovoltaic robot during movement;

[0008] The surrounding environment image is processed using a photovoltaic detection model to detect and obtain photovoltaic inspection information; the photovoltaic inspection information is used to indicate the working status of the photovoltaic component so as to detect faults in time.

[0009] In some embodiments, the step of processing the depth map using an improved occupancy detection model to detect and obtain environmental occupancy information comprises the following steps:

[0010] Each pixel value in the depth map represents the distance from the corresponding physical point to the camera, and based on the depth map and the camera intrinsic parameters, the three-dimensional space coordinates corresponding to each pixel value are calculated;

[0011] The calculated three-dimensional spatial coordinates of the corresponding physical point are stored in a point cloud data structure to obtain point cloud data;

[0012] A global scene is constructed in combination with the surrounding environment image captured by the wide-angle camera, and any target in the global scene is identified; point cloud data of the arbitrary target is obtained; and the point cloud data is voxelized to obtain the environmental occupancy information.

[0013] In some implementations, voxelizing the point cloud data to obtain the environmental occupancy information comprises the following steps:

[0014] According to the boundary of the point cloud data and the size of the selected voxel, create a voxel grid covering the entire point cloud data; traverse each point in the point cloud data, and assign each point to a corresponding voxel according to the three-dimensional space coordinates of the physical point corresponding to each point and the boundary of the voxel grid;

[0015] Each voxel is marked for occupation, and all the voxels marked with occupation marks in the point cloud data are output. Based on the correspondence between the voxels and the corresponding physical points, the occupancy status of the corresponding physical points is determined to obtain the environmental occupancy information.

[0016] In some implementations, the environment perception method and the improved occupancy detection model specifically include:

[0017] The occupancy detection model is a FlashOCC network model;

[0018] In the FlashOCC network model, the number of input and output channels is set; in the backbone network feature extraction stage of the FlashOCC network model, an LSTM+Transformer architecture model is inserted to enhance the algorithm's ability to process contextual connections between images of different perspectives and the model's ability to process global feature information of the surrounding environment image;

[0019] In the neck network feature fusion stage of the FlashOCC network model, the FPN structure is used to provide feature representations of different scales and reduce data loss in the calculation of the FlashOCC network model;

[0020] A feature branch is added after the FPN structure, and multiple feature detection heads of different sizes are used to perform occupancy detection and recognition on objects of different sizes at multiple scales.

[0021] In some implementations, the step of calculating the parallax of the surrounding environment images captured by any adjacent wide-angle cameras by using a stereo matching algorithm comprises the following steps:

[0022] Performing feature matching on the surrounding environment images captured by any adjacent wide-angle cameras by using the stereo matching algorithm;

[0023] Based on the repeated feature points, a repeated area in the two surrounding environment images is determined, and a disparity is calculated based on the repeated feature points in the repeated area; the disparity is the horizontal displacement distance of the repeated feature points between the two surrounding environment images.

[0024] In some embodiments, the step of converting the disparity into a depth map of the surrounding environment by combining the camera intrinsic parameters and the baseline distance of the wide-angle camera comprises the following steps:

[0025] Obtaining the camera internal parameters;

[0026] Acquire the actual physical distance between the optical centers of two adjacent wide-angle cameras to obtain the baseline distance between the two adjacent wide-angle cameras;

[0027] The depth value of each pixel corresponding to the surrounding environment image is obtained by the following formula, so as to obtain a depth map corresponding to the surrounding environment image;

[0028] Z = (f*B) / d;

[0029] Among them, Z is the pixel depth value, that is, the distance from the corresponding object in the surrounding environment image to the camera; f is the focal length of the wide-angle camera; B is the baseline distance; and d is the parallax.

[0030] In a second aspect, the present application further discloses an environment perception system, which is used to implement an environment perception method described in any one of the above embodiments, including:

[0031] A plurality of wide-angle cameras are installed on the photovoltaic robot and are used to collect images of the surrounding environment, and the plurality of wide-angle cameras have different viewing angles;

[0032] A disparity calculation module, used to calculate the disparity of the surrounding environment images captured by any two adjacent wide-angle cameras through a stereo matching algorithm; combining the camera intrinsic parameters and baseline distance of the wide-angle camera, converting the disparity into a depth map of the surrounding environment;

[0033] An environment occupancy information acquisition module uses an improved occupancy detection model to process the depth map and detect the environment occupancy information; the environment occupancy information is point cloud information of environment occupancy voxels, which is used for path planning and safe obstacle avoidance of the photovoltaic robot during movement;

[0034] The photovoltaic information acquisition module uses a photovoltaic detection model to process the surrounding environment image and detects photovoltaic inspection information; the photovoltaic inspection information is used to indicate the working status of the photovoltaic component so as to detect faults in time.

[0035] In some embodiments, a plurality of wide-angle cameras are dispersedly arranged at the same height around the photovoltaic robot to collect 360° images around the photovoltaic robot;

[0036] The images captured by any wide-angle camera and the images captured by its adjacent camera generate a certain amount of repetition, so as to calculate the parallax.

[0037] In some embodiments, the improved occupancy detection model is a FlashOCC network model;

[0038] In the improved FlashOCC network model, the LSTM+Transformer architecture model is inserted into the backbone network feature extraction stage of the FlashOCC network model to enhance the contextual connection between images of different perspectives and the model's ability to process global feature information of the surrounding environment image;

[0039] The neck network feature fusion stage of the FlashOCC network model uses an FPN structure to provide feature representations of different scales and reduce data loss in the calculation of the FlashOCC network model;

[0040] A feature branch is added after the FPN structure, and multiple feature detection heads of different sizes are used to perform occupancy detection and recognition on objects of different sizes at multiple scales.

[0041] In a third aspect, the present application also discloses a photovoltaic robot, comprising the environmental perception system described in any one of the above embodiments.

[0042] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0043] 1. This application adopts a pure visual solution in the photovoltaic robot environmental perception solution, and does not include various types of sensors such as laser radar. Only multiple wide-angle cameras are installed on the photovoltaic robot to obtain 360-degree environmental information. By processing the collected multi-view images through the occupancy detection model and the photovoltaic detection model, the redundancy of environmental information acquisition can be effectively reduced, the information processing load of the photovoltaic robot can be reduced, and the working stability of the robot can be enhanced.

[0044] 2. Process the images acquired by multiple wide-angle cameras by using the deep learning-based occupancy detection model and photovoltaic detection model. This application uses FlashOCC in combination with the YOLO algorithm to perceive environmental information by processing images collected by multiple wide-angle cameras, and obtains environmental occupancy information and photovoltaic module working status information respectively to complete the inspection task. FlashOcc network detection obtains environmental occupancy information, which is used for path planning and safe obstacle avoidance of photovoltaic robots during movement. YOLO network detection obtains photovoltaic inspection information, which is used to indicate the working status of photovoltaic modules so that faults can be discovered in a timely manner.

[0045] 3. In order to improve the detection capability of environmental occupancy information and maintain a good detection speed, the present application improves the FlashOCC algorithm by adding bidirectional LSTM and Transformer modules to the backbone network feature extraction stage of the FlashOCC network model to enhance the algorithm's connection between images of different perspectives; improve the algorithm's ability to process contextual connections of images of different perspectives, and at the same time, the size of the feature map in the last layer is small to reduce the amount of calculation.

[0046] 4. This application uses the FPN structure in the feature fusion stage of the FlashOCC algorithm, and uses three feature detection heads of different sizes for prediction, which improves the model's ability to handle targets of different sizes. This improves the algorithm's ability to detect the environment without taking up too much computing power. By obtaining the real occupancy information of all objects in the three-dimensional environment, based on the identification targeting of different feature detection heads for targets of different sizes, it is easier for the photovoltaic robot to respond to the working environment and improve the algorithm's ability to handle small target objects.

[0047] 5. In the model training phase, this application collects image information collected by the photovoltaic robot at the work site for annotation processing, and produces FlashOCC and YOLO data sets suitable for photovoltaic station work scenes, and enriches the data sets through data enhancement methods. Finally, after deep learning training, FlashOCC and YOLO models suitable for photovoltaic station work scenes are obtained. Through the self-made photovoltaic station environment data set for detection, data annotation is performed for the photovoltaic station environment to enhance the model's detection capability in the photovoltaic station. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The preferred implementation scheme will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of the present application.

[0049] Figure 1 A flowchart of an embodiment of an environment perception method provided by the present application;

[0050] Figure 2A flowchart of another embodiment of an environment perception method provided by the present application;

[0051] Figure 3 It is a structural diagram of the improved FlashOcc network model in the embodiment of the present application;

[0052] Figure 4 It is a structural diagram of the FlashOcc network model in the prior art;

[0053] Figure 5 Schematic diagram of the installation distribution positions of multiple wide-angle cameras in an embodiment of the present application;

[0054] Figure 6 A structural block diagram of an embodiment of an environment perception system provided in the present application. DETAILED DESCRIPTION

[0055] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0056] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections.

[0057] In order to simplify the drawings, only the parts related to the invention are schematically shown in each figure, and they do not represent the actual structure of the product. In addition, in order to simplify the drawings and facilitate understanding, in some figures, only one of the parts with the same structure or function is schematically drawn or marked. In this article, "one" not only means "only one", but also means "more than one".

[0058] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0059] In this document, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0060] In a specific implementation, the terminal device described in the embodiments of the present application includes, but is not limited to, other portable devices such as mobile phones, laptop computers, tutoring machines, or tablet computers with touch-sensitive surfaces (e.g., touch screen displays and / or touch pads). It should also be understood that in some embodiments, the terminal device is not a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., touch screen displays and / or touch pads).

[0061] In addition, in the description of the present application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0062] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the specific implementation methods of the present application will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings and other implementation methods can be obtained based on these drawings without creative work.

[0063] In recent years, the global demand for energy supply has continued to increase, while the global climate environment has gradually deteriorated. As one of the clean, accessible and renewable energy sources, solar energy has become the first choice for solving energy problems. Among them, photovoltaic power generation is one of the ways to use solar energy. As an important component for converting solar energy into electrical energy, the reliability and life of photovoltaic modules affect the efficiency of photovoltaic power generation, so the efficient operation of photovoltaic power stations is inseparable from timely and regular inspections. Large photovoltaic power stations are usually installed in areas with high temperatures and strong light. At the same time, affected by unstable factors in the external environment, the demand for 24-hour uninterrupted inspections has increased. Traditional manual inspections in large photovoltaic power station inspection tasks are high in intensity and dangerous, so photovoltaic robots can replace manual inspections to not only improve inspection efficiency and achieve 24-hour uninterrupted inspections, but also reduce the safety of staff. When inspecting photovoltaic modules, photovoltaic robots need to safely pass through the photovoltaic array, avoid photovoltaic station working equipment and staff, and reduce the damage to their own equipment caused by severe weather such as sandstorms. Therefore, efficient and accurate environmental perception technology and hardware equipment that takes into account both economy and reliability are required.

[0064] The photovoltaic robots currently used in the market are equipped with a large number of sensors of various types. The existing environmental perception technology generally relies on multiple sensors such as laser radar and cameras to perceive the environment. Including but not limited to high-definition visible light cameras, infrared thermal imaging cameras, RTK positioning sensors, IMU (inertial measurement units), 3D vision sensors, etc. Each sensor has its own focus in obtaining information. The use of more sensors not only increases the production cost of the robot, but also requires stronger computing hardware support to process the information of multiple sensors at the same time. For example, the A200 photovoltaic robot launched by ClearPath Robot and the photovoltaic robot of Hangzhou Guochen Robot Company in my country are equipped with thermal infrared imagers, laser radars, cameras and other sensors. In order to process the information output by each sensor in a timely manner, the equipment needs to be equipped with a higher-performance processor to complete the processing of environmental information in a timely manner and obtain environmental perception information.

[0065] Although the combination of multiple sensors can improve the accuracy of patrol measurements, the environmental information provided by multiple sensors contains a lot of duplication and redundancy. The repeated and redundant information is of no help in obtaining photovoltaic site inspection data and increases the workload of information processing during data processing.

[0066] Considering that the working environment of photovoltaic robots is generally at high temperatures, high temperatures will cause the temperature inside the photovoltaic robot processor to rise, exceeding the operating temperature of the components, which may cause the components to be damaged or burned. In order to protect itself from damage, the processor may start protection measures, which will slow down the processor and reduce processing efficiency.

[0067] At present, photovoltaic robots have redundant information acquisition in terms of perception. The greater the information processing demand, the greater the operating load of the robot. Especially in high-temperature environments, this is not conducive to the photovoltaic robots’ efficient processing of effective perception information, nor is it conducive to continuous inspection work in high-temperature environments.

[0068] In order to solve the technical problem of redundant information acquisition by photovoltaic robots caused by the above-mentioned multiple sensors, the present application provides an environmental perception method based on a pure visual algorithm, which adopts multiple wide-angle cameras to acquire environmental information, and acquires complete environmental perception information by combining an improved FlashOcc with the YOLO algorithm.

[0069] The purpose is to reduce the redundant information obtained by the sensor, enhance the efficient processing of environmental information, and enhance the continuous working ability of the photovoltaic robot in a high temperature environment. This application can enable the robot to reduce the processing of redundant information in high temperature perception scenarios, reduce the working pressure of the processor, and improve the working reliability and efficiency of the photovoltaic robot. It can improve the processing efficiency of perception information, reduce the production cost of photovoltaic robots, and improve the use efficiency of perception information when using fewer sensors, so as to better carry out 24-hour continuous inspection work.

[0070] Reference Manual Attached Figure 1 An embodiment of an environment perception method provided by the present application includes the following steps:

[0071] S100, collecting surrounding environment images based on multiple wide-angle cameras installed on the photovoltaic robot, wherein the multiple wide-angle cameras have different viewing angles.

[0072] Specifically, multiple wide-angle cameras (also called multi-view RGB wide-angle cameras, or multi-view cameras) are devices that can capture RGB (red, green, and blue) images from multiple angles. These cameras can be physically arranged at different viewing angles at the same height of the photovoltaic robot to cover a wider viewing angle range.

[0073] Multi-camera systems often require synchronization mechanisms to ensure that images captured from different cameras are temporally consistent, facilitating the matching of subsequent images.

[0074] S200, calculating the disparity of the surrounding environment images captured by any adjacent wide-angle cameras through a stereo matching algorithm, and converting the disparity into a depth map of the surrounding environment in combination with the camera intrinsic parameters and baseline distance of the wide-angle camera.

[0075] Specifically, the process of acquiring depth maps by multiple wide-angle cameras mainly involves two steps: stereo matching and disparity map conversion: The first is stereo matching: This embodiment is illustrated by taking any two adjacent perspectives as an example. Use two cameras to capture two images of the same scene. Based on these two images, a disparity map is calculated by a stereo matching algorithm. The disparity map represents the horizontal displacement distance of each pixel between the two images. Then the disparity map is converted into a depth map: the unit of the disparity map is pixels, and the unit of the depth map is usually millimeters. According to the geometric relationship of parallel binocular vision, the conversion between disparity and depth can be completed using existing known conversion formulas.

[0076] S300, using an improved occupancy detection model to process the depth map, and detect and obtain environmental occupancy information. The environmental occupancy information is point cloud information of environmental occupancy voxels, which is used for path planning and safe obstacle avoidance of the photovoltaic robot during movement.

[0077] Specifically, the occupancy detection model is a FlashOcc network model, which can obtain 360-degree occupancy information of the photovoltaic robot by merging and processing multi-view RGB image information. The FlashOCC occupancy network can process the occupancy information of unlabeled objects in the data set, such as photovoltaic components, obstacles, and staff. It has strong robustness and can handle the situation where unlabeled objects enter the working scene of the photovoltaic robot in emergencies, thereby increasing work safety.

[0078] "Point cloud information of environmental occupancy voxel particles" usually refers to environmental occupancy information represented by point cloud data in three-dimensional space, where each point represents a specific location in the environment, and these points are combined to form a three-dimensional description of the environment. Voxel is a pixel concept in three-dimensional space, that is, a point in three-dimensional space, which has three dimensions: length, width and height. In point cloud data, voxel is usually used to represent a cubic unit in space, which can be used to represent obstacles, ground or other objects in the environment.

[0079] By analyzing the characteristics of voxels, it is determined which voxels in the environment are occupied and which are idle. Then it is determined whether the actual object corresponding to the voxel is occupied, which can be used by photovoltaic robots to identify and avoid obstacles during autonomous driving.

[0080] S400, using a photovoltaic detection model to process the surrounding environment image, and detect and obtain photovoltaic inspection information. The photovoltaic inspection information is used to indicate the working state of the photovoltaic component, so as to find faults in time.

[0081] Specifically, the photovoltaic detection model is a YOLO model, which is based on artificial intelligence technology such as mature YOLOv8 or YOLOv6 deep learning algorithms to intelligently identify and analyze the equipment in the photovoltaic field, accurately identify the equipment status and detect abnormal conditions. Of course, other existing photovoltaic detection models can also be used to detect the photovoltaic working status, and this application does not make specific limitations.

[0082] Photovoltaic inspection information is determined based on the specific tasks performed by the photovoltaic robot. For example, it detects whether the photovoltaic panel is blocked by foreign objects; whether the photovoltaic module bridge is normal; whether the photovoltaic module is working normally; the cleaning status of the photovoltaic panel surface, etc., so as to timely discover and deal with problems, improve the operation and maintenance efficiency and safety of photovoltaic power stations, reduce labor costs, and improve power generation efficiency.

[0083] Another embodiment of the environment perception method of the present application, based on an embodiment of the above method, step S200 includes the following steps:

[0084] S210: Perform feature matching on the surrounding environment images captured by any adjacent wide-angle cameras by using a stereo matching algorithm.

[0085] Specifically, an existing feature matching algorithm may be used, and this application does not make any specific limitation.

[0086] S220, based on the repeated feature points, determine a repeated area in the two surrounding environment images, and calculate a disparity based on the repeated feature points in the repeated area. The disparity is the horizontal displacement distance of the repeated feature points between the two surrounding environment images.

[0087] Specifically, the parallax can be expressed by the following formula: d = (x _ L)-(x _ R);

[0088] Where d is the disparity, (x_L) and (x_R) are the horizontal coordinates of the corresponding points in the left and right cameras respectively.

[0089] S230, combining the camera intrinsic parameters and the baseline distance of the wide-angle camera, converting the disparity into a depth map of the surrounding environment. Specifically comprising the following steps:

[0090] S231, obtaining the camera internal parameters.

[0091] S232: Acquire an actual physical distance between optical centers of two adjacent wide-angle cameras to obtain a baseline distance between the two adjacent wide-angle cameras.

[0092] S233, obtaining a depth value of each pixel point corresponding to the surrounding environment image by using the following formula, thereby obtaining a depth map corresponding to the surrounding environment image.

[0093] Z = (f*B) / d;

[0094] Among them, Z is the pixel depth value, that is, the distance from the corresponding object in the surrounding environment image to the camera; f is the focal length of the wide-angle camera; B is the baseline distance; and d is the parallax.

[0095] Specifically, by obtaining the camera intrinsic parameters and calibrating the extrinsic parameters of multiple wide-angle cameras relative to the center point. The baseline distance refers to the horizontal distance between the imaging planes of two adjacent cameras (or sensors) in a multiple wide-angle camera system. By calibrating the camera, the internal and external parameters of the two adjacent cameras can be obtained, where the intrinsic parameter matrix contains the focal length and optical center position information, and the extrinsic parameter matrix contains the relative position and posture information between the cameras. Through these parameters, the baseline distance between two adjacent cameras can be calculated.

[0096] There are repeated areas between the images collected by multiple wide-angle cameras. First, feature matching is performed on the image information to find the corresponding repeated feature points and calculate the disparity. The horizontal displacement distance of the same point in adjacent images is called disparity. The larger the disparity, the closer the object is to the camera. The smaller the disparity, the farther the object is from the camera. By calculating the disparity values ​​of all pixels, a depth map can be generated, and the value of each pixel represents the distance from the point to the camera. Similarly, by calculating the images from the remaining camera perspectives, the distance information in the global scene can be obtained, thereby completing occupancy recognition in the three-dimensional scene.

[0097] Another embodiment of an environment perception method of the present application, based on an embodiment of the above method, step S300 is shown, using an improved occupancy detection model to process the depth map, and detect and obtain environment occupancy information. The steps include:

[0098] S310, each pixel value in the depth map represents the distance from the corresponding physical point to the camera, and based on the depth map and the camera intrinsic parameters, the three-dimensional space coordinates corresponding to each pixel are calculated.

[0099] Specifically, load the depth image data. The depth map is a two-dimensional image in which each pixel value represents the distance from the corresponding physical point to the camera. In order to convert the depth map into a point cloud, you need to know the intrinsic parameters of the camera, including the focal length and the coordinates of the principal point. These parameters define the imaging model of the camera and participate in the calculation of the three-dimensional space coordinate conversion. Using the depth map and the camera intrinsic parameters, the three-dimensional space coordinates corresponding to each pixel can be calculated. Specifically, for each pixel in the depth map, its depth value is obtained through the depth map. Then the three-dimensional coordinates in the world coordinate system can be obtained through a fixed conversion formula.

[0100] S320, storing the calculated three-dimensional spatial coordinates of the corresponding physical point in a point cloud data structure to obtain point cloud data.

[0101] Specifically, the three-dimensional coordinates of each point calculated above are stored in a point cloud data structure. In actual use, the spatial point cloud data consists of a series of points, each of which contains its coordinates in three-dimensional space. The generated point cloud may require further processing, such as sampling, filtering, etc., to improve the quality of the point cloud or adapt to specific application requirements, which is not limited in this application.

[0102] S330, constructing a global scene in combination with the surrounding environment image captured by the wide-angle camera, identifying any target in the global scene, obtaining point cloud data of the any target, voxelizing the point cloud data, and thereby obtaining the environment occupancy information.

[0103] In one implementation of this embodiment, in step S330: voxelizing the point cloud data to obtain the environment occupancy information includes the following steps:

[0104] S331, creating a voxel grid covering the entire point cloud data according to the boundary of the point cloud data and the selected voxel size, traversing each point in the point cloud data, and assigning each point to a corresponding voxel according to the three-dimensional space coordinates of the physical point corresponding to each point and the boundary of the voxel grid.

[0105] S332, marking the occupancy of each voxel, and outputting all the voxels marked with the occupancy mark in the point cloud data, combining the correspondence between the voxels and the corresponding physical points, determining the occupancy status of the corresponding physical points, and then obtaining the environmental occupancy information.

[0106] Specifically, point cloud voxelization is a process of converting unordered point cloud data into an ordered three-dimensional voxel grid. First, the size of each voxel (i.e., the side length of the voxel) needs to be determined. The size of the voxel depends on specific requirements, such as processing speed, data accuracy, etc. Smaller voxels can provide higher accuracy, but increase the amount of calculation and memory consumption; larger voxels are the opposite. According to the boundaries of the point cloud data and the selected voxel size, a voxel grid covering the entire point cloud data is created. This usually involves determining the dimension of the grid (i.e., the number of voxels in each dimension) and calculating the position and boundaries of each voxel in three-dimensional space. Traverse each point in the point cloud and assign the point to the corresponding voxel according to the coordinates of the point and the boundaries of the voxel grid. To determine the voxel index to which the point belongs. For each voxel, different processing can be performed as needed. For example, the voxel can be marked as "occupied". After completing the above steps, the results of voxelization, i.e., the voxel grid and the occupancy status of the voxel, are output.

[0107] Another embodiment of an environment perception method of the present application, based on an embodiment of any of the above methods, before step S300, further includes improving an occupancy detection model, where the occupancy detection model is a FlashOCC network model; specifically includes the following steps:

[0108] S011, in the FlashOCC network, set the number of input and output channels. Insert the LSTM+Transformer architecture model in the backbone network feature extraction stage of the FlashOCC network model to enhance the algorithm's ability to process contextual connections of images from different perspectives and the model's ability to process global feature information of the surrounding environment image.

[0109] S012, using the FPN structure in the neck network feature fusion stage of the FlashOCC network model to provide feature representations of different scales and reduce data loss in the convolution calculation of the FlashOCC network model.

[0110] S013, adding a feature branch after the FPN structure, using multiple feature detection heads of different sizes, to perform occupancy detection and recognition on objects of different sizes at multiple scales.

[0111] This application improves the FlashOCC algorithm by adding bidirectional LSTM and Transformer modules to the backbone network feature extraction stage of the FlashOCC network model to enhance the algorithm's connection between images from different perspectives. In the feature fusion stage, multiple (for example, three) feature detection heads of different sizes are used for prediction to improve the model's ability to handle targets of different sizes. This improves the algorithm's ability to detect the environment without taking up too much computing power. By obtaining the real occupancy information of all objects in the three-dimensional environment, it is easier for the photovoltaic robot to respond to the working environment.

[0112] Another embodiment of the environment perception method of the present application summarizes the improvement, training and use process of the model, and refers to the attached manual. Figure 2 As shown, the following steps are included:

[0113] In the first step, the image information collected from the photovoltaic robot work site is annotated and processed, and FlashOCC and YOLO datasets suitable for photovoltaic station work scenes are produced respectively, and the datasets are enriched by data enhancement methods. Finally, FlashOCC and YOLO models suitable for photovoltaic station work scenes are obtained through deep learning training. This application uses a self-made photovoltaic station environment dataset for detection, annotates data for the photovoltaic station environment, and enhances the model's detection capabilities in photovoltaic stations.

[0114] In the second step, the improved FlashOCC algorithm of the present application is to enhance the perception of environmental information of photovoltaic power stations under the conditions of multiple wide-angle cameras. A bidirectional LSTM method is added to the backbone network feature extraction stage of the FlashOCC network model to enhance the algorithm's ability to process contextual connections of images from different perspectives. The Transformer model is added to the last layer of the backbone network feature extraction stage of the FlashOCC network model to improve the model's ability to process global feature information of the image and strengthen the connection between image information from different perspectives. At the same time, during the convolution operation of the backbone network of the FlashOCC network model, the size of the feature map continues to decrease. After a series of convolution operations are performed on the last layer of the backbone network of the FlashOCC network model through the ResNet network, the feature map reaches the minimum. Adding bidirectional LSTM and Transformer modules here can reduce the amount of calculation. At the same time, enhancing the connection between multi-perspective images is beneficial to improving the accuracy of the model in establishing 360-degree global photovoltaic robot working environment occupancy information. The model structure diagram is shown in the figure below. Figure 3 shown. Figure 3 The schematic diagram of the structure of the improved FlashOcc network model in the embodiment of the present application is shown in FIG. The bidirectional LSTM and Transformer modules can realize forward reasoning by setting the same number of channels.

[0115] In the third step, in order to enhance FlashOCC's ability to identify the occupancy information of objects of different sizes in the environment, the FPN structure is used in the neck network feature fusion stage of the FlashOCC network model. The FPN structure performs ConCat operations on two feature maps of different sizes to fuse the information of feature maps of different sizes, so that the model reduces the information loss in the convolution process. At the same time, two feature branches are added, such as Figure 3 Compared to Figure 4 , Figure 4 This is a structural diagram of the FlashOcc network model in the prior art. This application uses three feature detection heads of different sizes to enhance the algorithm's ability to handle occupancy of targets of different sizes, and can also enhance the algorithm's detection accuracy when the photovoltaic robot encounters small targets. At the same time, it can also improve the photovoltaic robot's ability to accurately detect obstacles when it is close to the obstacles, prevent the phenomenon of missing small targets or treating large target objects as backgrounds, and improve the photovoltaic robot's perception accuracy of the environment.

[0116] Step 4: Install Figure 5 Multiple wide-angle cameras shown, Figure 5This is a schematic diagram of the installation distribution positions of multiple wide-angle cameras in the embodiment of this application. This embodiment takes a four-view camera as an example to obtain the environmental RGB image in real time. The point cloud information of the occupied voxel particles can be calculated through the improved FlashOCC model. It is composed of a large number of four-dimensional arrays, which are X, Y, Z and color information, and the environmental occupancy information can be obtained. At the same time, the YOLO algorithm detects the four-view image again to detect the working conditions of photovoltaic components such as photovoltaic panels. Improve the efficiency of using sensor information.

[0117] The fifth step is to calculate the environmental occupancy information obtained by the Flashocc algorithm, which can be used for path planning and safe obstacle avoidance during the movement of the photovoltaic robot. It can detect whether the photovoltaic panels and other photovoltaic components are working normally. Finally, the photovoltaic robot obtains complete environmental perception information through FlashOCC and YOLO algorithms for inspection.

[0118] This application uses a pure vision solution to reduce the robot hardware cost and processor computing cost, and enhance stability in high-temperature working scenarios. By collecting multi-view RGB wide-angle images of the photovoltaic robot, the improved FlashOCC and YOLO algorithms are used to obtain environmental occupancy information and monitor the working status of photovoltaic components. This improves the current high production cost of photovoltaic robots and redundant environmental information processing, and enhances the working stability of photovoltaic robots.

[0119] Based on the same technical concept, the present application also discloses an environment perception system, which can be used to implement any of the above-mentioned environment perception methods. Specifically, an embodiment of the environment perception system of the present application is shown in the attached specification. Figure 6 As shown, including:

[0120] Multiple wide-angle cameras are installed on the photovoltaic robot to collect images of the surrounding environment.

[0121] Specifically, multiple wide-angle cameras (also known as multi-view RGB wide-angle cameras) are devices that can capture RGB (red, green, and blue) images from multiple angles. These cameras can be physically arranged at different viewing angles at the same height of the photovoltaic robot to cover a wider viewing angle range. This embodiment takes the distribution of four viewing angle cameras as an example. The four viewing angle cameras are distributed in the reference manual. Figure 5 shown.

[0122] Preferably, a multi-camera system usually requires a synchronization mechanism to ensure that images captured from different cameras are consistent in time, which facilitates the matching of subsequent images.

[0123] More preferably, a plurality of wide-angle cameras are dispersedly arranged at the same height around the photovoltaic robot to collect 360° images around the photovoltaic robot.

[0124] The images captured by any wide-angle camera and the images captured by its adjacent camera generate a certain amount of repetition, so as to calculate the parallax.

[0125] The disparity calculation module is used to calculate the disparity of the surrounding environment images collected by any adjacent wide-angle cameras through a stereo matching algorithm, and convert the disparity into a depth map of the surrounding environment in combination with the camera intrinsic parameters and baseline distance of the wide-angle camera.

[0126] Specifically, the process of acquiring depth maps by multiple wide-angle cameras mainly involves two steps: stereo matching and disparity map conversion: The first is stereo matching: This embodiment is illustrated by taking any two adjacent perspectives as an example. Use two cameras to capture two images of the same scene. Based on these two images, a disparity map is calculated by a stereo matching algorithm. The disparity map represents the horizontal displacement distance of each pixel between the two images. Then the disparity map is converted into a depth map: the unit of the disparity map is pixels, and the unit of the depth map is usually millimeters. According to the geometric relationship of parallel binocular vision, the conversion between disparity and depth can be completed using existing known conversion formulas.

[0127] The environment occupancy information acquisition module uses an improved occupancy detection model to process the depth map and detect the environment occupancy information. The occupancy detection model in this application is a FlashOCC network model. The environment occupancy information is the point cloud information of the environment occupancy voxels, such as photovoltaic components, obstacles, staff, etc. It is used for the photovoltaic robot to perform path planning and safe obstacle avoidance during movement.

[0128] Specifically, "point cloud information of environmental occupancy voxel particles" usually refers to environmental occupancy information represented by point cloud data in three-dimensional space, where each point represents a specific location in the environment, and these points are combined to form a three-dimensional description of the environment. Voxel is a pixel concept in three-dimensional space, that is, a point in three-dimensional space, which has three dimensions of length, width and height. In point cloud data, voxel is usually used to represent a cubic unit in space, which can be used to represent obstacles, ground or other objects in the environment.

[0129] By analyzing the characteristics of voxels, it is determined which voxels in the environment are occupied and which are idle. This can be used by photovoltaic robots to identify and avoid obstacles during autonomous driving.

[0130] The photovoltaic information acquisition module uses a photovoltaic detection model to process the surrounding environment image and detect photovoltaic inspection information. The photovoltaic inspection information is used to indicate the working status of the photovoltaic component so as to find faults in time.

[0131] Specifically, the photovoltaic detection model in this application is a YOLO network model. Based on the artificial intelligence technology of mature deep learning algorithms such as YOLOv8 or YOLOv6, the equipment in the photovoltaic field is intelligently identified and analyzed, and the equipment status and abnormal conditions are accurately identified. Of course, other existing photovoltaic detection models can also be used to detect the photovoltaic working status, and this application does not make specific limitations.

[0132] Photovoltaic inspection information is determined based on the specific tasks performed by the photovoltaic robot. For example, it detects whether the photovoltaic panel is blocked by foreign objects; whether the photovoltaic module bridge is normal; whether the photovoltaic module is working normally; the cleaning status of the photovoltaic panel surface, etc., so as to timely discover and deal with problems, improve the operation and maintenance efficiency and safety of photovoltaic power stations, reduce labor costs, and improve power generation efficiency.

[0133] Based on the above embodiment, the present application provides another embodiment of an environment perception system, wherein the improved FlashOCC model inserts an LSTM+Transformer architecture model into the backbone network feature extraction stage of the FlashOCC network model to enhance the processing of contextual connections between images of different perspectives and the model's ability to process global feature information of the surrounding environment image.

[0134] The neck network feature fusion stage of the FlashOCC network model uses an FPN structure to provide feature representations of different scales and reduce data loss in the convolution calculation of the FlashOCC network model.

[0135] A feature branch is added after the FPN structure, and multiple feature detection heads of different sizes are used to perform occupancy detection and recognition on objects of different sizes at multiple scales.

[0136] Preferably, the improvements of the FlashOCC model of the present application include: 1. The LSTM (Long Short-Term Memory) plug-in (or LSTM layer) is a special recurrent neural network (RNN) structure that can capture long-term dependencies in time series data. It helps to enhance the contextual connection of the algorithm to process images from different perspectives.

[0137] 2. Transformer network structure is usually composed of multiple identical encoder and decoder layers stacked together. These stacked layers help the model learn complex feature representations. It helps to improve the model's ability to process global feature information of images and strengthen the connection between image information from different perspectives.

[0138] 3. In the neck network feature fusion stage of the FlashOCC network model, the FPN (Feature Pyramid Networks) structure is used to provide feature representations of different scales, which is very critical for detecting objects of different sizes.

[0139] 4. After the FPN structure, a detection head is added independently at each scale for prediction. Multiple detection heads can make independent predictions on feature maps of different scales, which can more effectively detect objects of different sizes. Each detection head focuses on the feature map of a specific scale and can capture the detailed information of the object at that scale, thereby improving the accuracy and robustness of detection.

[0140] Preferably, in the photovoltaic module detection data set prepared in this application, images of the photovoltaic robot working scene are collected as the data set, and these images depict various situations of the photovoltaic robot during work, providing sufficient training samples for FlashOCC and YOLO network. Various data set forms are used for annotation.

[0141] This application improves the FlashOCC algorithm by adding bidirectional LSTM and Transformer modules to the backbone network of the FlashOCC network model to improve the algorithm's feature extraction capabilities, while enhancing the algorithm's contextual connections between images from different perspectives and improving occupancy detection capabilities. At the same time, in order to enhance the algorithm's ability to detect targets of different sizes, an FPN structure is used in the feature fusion stage, and three feature detection heads of different sizes are used for prediction to improve small target detection capabilities.

[0142] This application uses the FlashOCC algorithm to perform occupancy detection on the 4-view images obtained by the camera, obtains information about occupied objects in the environment, and provides walkable route information for navigation. At the same time, the 4-view information can also be used by the YOLO algorithm to detect the working status of photovoltaic modules. Make full use of existing information to perceive the environment.

[0143] Based on the same concept, the present application also discloses a photovoltaic robot, comprising the environmental perception system described in any one of the above embodiments.

[0144] Specifically, in the inspection task of photovoltaic robots, this application adopts a pure vision solution to collect 4-view RGB images, and uses improved FlashOCC and YOLO algorithms to analyze the environment. This visual perception solution has the advantages of reducing the hardware cost of the robot, using only RGB cameras to obtain environmental information, and using FlashOCC and YOLO algorithms with low processor load requirements, which can reduce the workload of the processor in high temperature environments, improve the efficiency of sensor information use, improve the working stability of photovoltaic robots, and reduce the production cost of photovoltaic robots.

[0145] An environmental perception method, system, and photovoltaic robot of the present application have the same technical concept, and the technical details of the embodiments of the two are applicable to each other. In order to reduce repetition, they will not be repeated here.

[0146] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the division of the above-mentioned program modules is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program units or modules to complete all or part of the functions described above. The program modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into a processing unit, and the above-mentioned integrated unit can be implemented in the form of hardware or in the form of software program units. In addition, the specific names of the program modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application.

[0147] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0148] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for environmental perception, characterized in that: The steps include: Based on multiple wide-angle cameras installed on the photovoltaic robot, the surrounding environment image is collected, and the multiple wide-angle cameras have different viewing angles; Calculating the parallax of the surrounding environment images captured by any two adjacent wide-angle cameras by using a stereo matching algorithm; Combining the camera intrinsic parameters and the baseline distance of the wide-angle camera, converting the disparity into a depth map of the surrounding environment; Using an improved occupancy detection model to process the depth map, and detect and obtain environmental occupancy information; the environmental occupancy information is point cloud information of environmental occupancy voxels, which is used for path planning and safe obstacle avoidance of the photovoltaic robot during movement; The surrounding environment image is processed using a photovoltaic detection model to detect and obtain photovoltaic inspection information; the photovoltaic inspection information is used to indicate the working status of the photovoltaic component so as to detect faults in time.

2. The environment perception method according to claim 1, characterized in that: The method of using the improved occupancy detection model to process the depth map and detect and obtain environmental occupancy information comprises the following steps: Each pixel value in the depth map represents the distance from the corresponding physical point to the camera, and based on the depth map and the camera intrinsic parameters, the three-dimensional space coordinates corresponding to each pixel value are calculated; The calculated three-dimensional spatial coordinates of the corresponding physical point are stored in a point cloud data structure to obtain point cloud data; Constructing a global scene in combination with the surrounding environment image captured by the wide-angle camera, and identifying any target in the global scene; Obtaining point cloud data of the arbitrary target; The point cloud data is voxelized to obtain the environment occupancy information.

3. The environment perception method according to claim 2, characterized in that: The step of voxelizing the point cloud data to obtain the environmental occupancy information comprises the following steps: According to the boundary of the point cloud data and the selected voxel size, create a voxel grid covering the entire point cloud data; traverse each point in the point cloud data, and assign each point to a corresponding voxel according to the three-dimensional space coordinates of the physical point corresponding to each point and the boundary of the voxel grid; Each voxel is marked for occupation, and all the voxels marked with occupation marks in the point cloud data are output. Based on the correspondence between the voxels and the corresponding physical points, the occupancy status of the corresponding physical points is determined to obtain the environmental occupancy information.

4. An environment perception method according to any one of claims 1 to 3, characterized in that: The improved occupancy detection model specifically includes: The occupancy detection model is a FlashOCC network model; In the FlashOCC network model, the number of input and output channels is set; in the backbone network feature extraction stage of the FlashOCC network model, an LSTM+Transformer architecture model is inserted to enhance the processing of contextual connections between images of different perspectives and the processing capability of the model for global feature information of the surrounding environment image; In the neck network feature fusion stage of the FlashOCC network model, the FPN structure is used to provide feature representations of different scales and reduce data loss in the calculation of the FlashOCC network model; A feature branch is added after the FPN structure, and multiple feature detection heads of different sizes are used to perform occupancy detection and recognition on objects of different sizes at multiple scales.

5. The environment perception method according to claim 1, characterized in that: The method of calculating the parallax of the surrounding environment images captured by any adjacent wide-angle cameras by using a stereo matching algorithm comprises the following steps: Performing feature matching on the surrounding environment images captured by any adjacent wide-angle cameras by using the stereo matching algorithm; Based on the repeated feature points, a repeated area in the two surrounding environment images is determined, and a disparity is calculated based on the repeated feature points in the repeated area; the disparity is the horizontal displacement distance of the repeated feature points between the two surrounding environment images.

6. The environment perception method according to claim 5, characterized in that: The method of converting the parallax into a depth map of the surrounding environment by combining the camera intrinsic parameters and the baseline distance of the wide-angle camera comprises the following steps: Obtaining the camera internal parameters; Acquire the actual physical distance between the optical centers of two adjacent wide-angle cameras to obtain the baseline distance between the two adjacent wide-angle cameras; The depth value of each pixel corresponding to the surrounding environment image is obtained by the following formula, so as to obtain a depth map corresponding to the surrounding environment image; Z = (f*B) / d; Among them, Z is the pixel depth value, that is, the distance from the corresponding object in the surrounding environment image to the camera; f is the focal length of the wide-angle camera; B is the baseline distance; and d is the parallax.

7. An environment perception system, characterized in that: The system is used to implement an environment perception method according to any one of claims 1 to 6, comprising: A plurality of wide-angle cameras are installed on the photovoltaic robot and are used to collect images of the surrounding environment, and the plurality of wide-angle cameras have different viewing angles; A disparity calculation module, used to calculate the disparity of the surrounding environment images captured by any two adjacent wide-angle cameras through a stereo matching algorithm; combining the camera intrinsic parameters and baseline distance of the wide-angle camera, converting the disparity into a depth map of the surrounding environment; An environment occupancy information acquisition module uses an improved occupancy detection model to process the depth map and detect the environment occupancy information; the environment occupancy information is point cloud information of environment occupancy voxels, which is used for path planning and safe obstacle avoidance of the photovoltaic robot during movement; The photovoltaic information acquisition module uses a photovoltaic detection model to process the surrounding environment image and detects photovoltaic inspection information; the photovoltaic inspection information is used to indicate the working status of the photovoltaic component so as to detect faults in time.

8. An environment sensing system as claimed in claim 7, characterized in that: A plurality of wide-angle cameras are dispersedly arranged at the same height around the photovoltaic robot to collect 360° images around the photovoltaic robot; The images captured by any wide-angle camera and the images captured by its adjacent camera generate a certain amount of repetition, so as to calculate the parallax.

9. An environment sensing system as claimed in claim 7, characterized in that: include: The improved occupancy detection model is a FlashOCC network model; In the improved FlashOCC network model, the LSTM+Transformer architecture model is inserted into the backbone network feature extraction stage of the FlashOCC network model to enhance the processing of contextual connections between images of different perspectives and the model's processing capabilities for global feature information of the surrounding environment image; The neck network feature fusion stage of the FlashOCC network model uses an FPN structure to provide feature representations of different scales and reduce data loss in the calculation of the FlashOCC network model; A feature branch is added after the FPN structure, and multiple feature detection heads of different sizes are used to perform occupancy detection and recognition on objects of different sizes at multiple scales.

10. A photovoltaic robot, characterized in that: An environment perception system comprising any one of claims 7-9.