Lawn trimming robot environment sensing method and device based on embedded vision
Through the environmental perception method of lawn pruning robot based on embedded vision, the problem of insufficient visual recognition performance of lawn pruning robots in the prior art is solved. Through data augmentation and model optimization, the intelligent perception and target recognition of the environment of the lawn pruning robot are realized.
Patent Information
- Application Number
- CN202510149013.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
AI Technical Summary
Existing lawn pruning robots face insufficient high-quality data sets in visual recognition and insufficient performance of image processing algorithms in speed and accuracy, which affects their intelligent development and large-scale application.
A lawn pruning robot environment perception method based on embedded vision is proposed. By acquiring lawn image data, fine labeling and data enhancement, a lightweight neural network model is built, the output structure of YOLOv11 instance segmentation model is optimized, and the intermediate model is quantized, and the model is finally deployed to the embedded chip to achieve intelligent perception.
Through data augmentation and model optimization, the robustness and speed of visual recognition are enhanced, the problem of difficult identification of complex boundaries and obstacles is solved, and the intelligent perception and target recognition of the environment by lawn mowing robots is realized.
Smart Images

Figure CN120147850A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robots, and in particular, to an environmental perception method and device for a lawn mowing robot based on embedded vision. Background Art
[0002] With the continuous progress of machine vision and artificial intelligence technologies, the development of lawn mowing robots has been rapid, and the requirements for their automation, convenience, and intelligence levels have also been continuously improving. Especially in the field of visual recognition technology, how to accurately and quickly identify the lawn boundary and prevent potential obstacles has become the focus of research. Lawn mowing robots integrate sensors, cameras, and efficient image processing algorithms to achieve intelligent perception of the lawn operation environment and optimize the lawn mowing effect and improve work efficiency.
[0003] However, current lawn mowing robots still face some challenges in visual recognition, including the lack of applicable high-quality data sets and the insufficient performance of existing image processing algorithms in terms of speed and accuracy. If these problems are not solved in a timely manner, they may restrict the intelligent development of lawn mowing robots and affect their popularization and promotion in large-scale applications. Summary of the Invention
[0004] The purpose of the present invention is to at least solve one of the deficiencies of the prior art, and provide an environmental perception method and device for a lawn mowing robot based on embedded vision.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] Specifically, an environmental perception method for a lawn mowing robot based on embedded vision is proposed, including the following:
[0007] Obtain lawn image data, and perform refined annotation on the lawn image data to obtain an annotated data set;
[0008] Expand the annotated data set through a data augmentation strategy to obtain a training data set;
[0009] Construct a neural network model based on model lightweight processing, the neural network model is constructed by replacing the activation function of YOLOv11, and the neural network model is trained through the training data set to obtain a feature extraction model;
[0010] Optimize the output structure of the YOLOv11 instance segmentation model, and export the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model;
[0011] Perform quantization processing on the intermediate model to obtain a quantized model;
[0012] By writing an inference program for the quantized model and deploying it to the target embedded chip, intelligent perception and target recognition of the lawn environment of the robot associated with the target embedded chip are realized.
[0013] Furthermore, specifically, lawn image data is obtained, and the lawn image data is finely annotated to obtain an annotation dataset, including,
[0014] Under different weather conditions, the weather conditions including sunny days, rainy days and nights, lawn image data including lawns, boundaries and obstacles are collected. Then, the collected lawn image data is uploaded to the Roboflow platform for fine annotation processing to obtain an annotation dataset.
[0015] Furthermore, specifically, the annotation dataset is augmented through a data augmentation strategy to obtain a training dataset, including,
[0016] The annotation dataset is augmented through image enhancement techniques to obtain a training dataset. The image enhancement techniques include rotation, cropping, changing the hue, saturation, brightness of the picture, and adding noise.
[0017] Furthermore, specifically, the construction process of the neural network model includes,
[0018] Replace the SiLU activation function of YOLOv11 with the ReLU function. The expression of the ReLU function is,
[0019] f(x) = max(0, x),
[0020] That is, when the input is greater than zero, the output is the input value itself, otherwise the output is zero.
[0021] Furthermore, specifically, optimizing the output structure of the YOLOv11 instance segmentation model includes,
[0022] Adjust the output structure of the Segment module of the output structure of the YOLOv11 instance segmentation model to 10 independent output heads, where the Bounding Box information, Masks information, and Classify information of the three feature layers are output independently.
[0023] Furthermore, specifically, quantizing the intermediate model to obtain a quantized model includes,
[0024] Determine the PTQ scheme quantization conversion process based on the model of the selected embedded chip, write a quantization program for the YOLOv11 instance segmentation model based on the PTQ scheme quantization conversion process, and convert the intermediate model format to a model format adapted to this embedded chip. Specifically,
[0025] First, determine whether the intermediate model can be used properly on the selected embedded chip;
[0026] Then, calibrate the model;
[0027] Next, write a configuration file for quantizing the YOLOv11 instance segmentation model, compress the intermediate model from FP32 to INT8 format, and obtain the quantized model.
[0028] Furthermore, specifically,
[0029] Select Horizon RDK_X3_Module as the embedded chip. The intermediate model format is ONNX format, and the model format adapted to the Horizon RDK_X3_Module chip is BIN format. Therefore, according to the model conversion toolchain provided by Horizon official, convert the ONNX intermediate format to BIN format on the development machine.
[0030] Furthermore, specifically, write an inference program for the quantized model, including,
[0031] Using the communication interface of the embedded development platform, the model is transmitted to the specified location of the device and loaded into the inference engine. Memory resources and hardware acceleration parameters are configured during initialization. The inference program preprocesses the input image, then calls the quantized model for real-time inference, outputs information such as bounding boxes, masks, class labels, and confidence levels, and generates a visualization result through post-processing.
[0032] Furthermore, specifically, the preprocessing includes size adjustment and format conversion.
[0033] The present invention proposes a device for environmental perception of a lawn mowing robot based on embedded vision, including the following:
[0034] A labeled dataset construction module, used to obtain lawn image data and perform refined labeling on the lawn image data to obtain a labeled dataset;
[0035] A training dataset construction module, used to expand the labeled dataset through a data augmentation strategy to obtain a training dataset;
[0036] A feature extraction model construction module, used to construct a neural network model based on model lightweight processing. The neural network model is constructed by replacing the activation function of YOLOv11, and the neural network model is trained through the training dataset to obtain a feature extraction model;
[0037] An intermediate model construction module, configured to optimize the output structure of the YOLOv11 instance segmentation model, and export the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model;
[0038] A quantization module, configured to perform quantization processing on the intermediate model to obtain a quantized model;
[0039] A program deployment module, configured to implement intelligent perception and target recognition of the lawn environment of a robot associated with a target embedded chip by writing an inference program for the quantized model and deploying it to the target embedded chip.
[0040] The beneficial effects of the present invention are as follows:
[0041] The present invention provides a method and device for lawn mowing robot environment perception based on embedded vision. By using a vision sensor to collect lawn image data and performing data augmentation strategies such as noise addition, hue and saturation adjustment on the image data to comprehensively simulate various changes in the real environment, the problem of difficult recognition due to environmental changes in the past is solved, and the robustness of visual recognition is enhanced. For the problem of difficult positioning and accurate recognition of complex boundaries and obstacles, the present invention selects the YOLOv11 instance segmentation model, which combines the advantages of object detection and semantic segmentation, can accurately distinguish and segment the contours and categories of each independent object in the image, has strong environmental understanding ability, and effectively reduces the number of model parameters and floating-point operation times by using the ReLU activation function to replace the original SiLU activation function of the YOLOv11 instance segmentation model, improving the recognition speed of the algorithm. For problems such as feature coupling and redundant calculations existing in the model, the present invention optimizes the output structure of the YOLOv11 instance segmentation model, improving the inference speed of the model. For the problem of limited computing resources of embedded devices, the present invention compresses the weights and activation values of the model from FP32 to INT8 format through quantization, effectively reducing the size and storage space requirements of the model, and improving the inference speed of the model on the embedded side. Finally, by deploying the model to an embedded chip and writing its inference program, intelligent perception of the surrounding environment by the lawn mowing robot using a vision sensor is realized. Description of the Drawings
[0042] By elaborating on the embodiments shown in conjunction with the drawings, the above and other features of the present disclosure will become more apparent. The same reference numerals in the drawings of the present disclosure denote the same or similar elements. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0043] Figure 1 The figure shows a flowchart of the environmental perception method of the lawn mowing robot based on embedded vision according to the present invention;
[0044] Figure 2 The figure shows a structural schematic diagram of the environmental perception device of the lawn mowing robot based on embedded vision according to the present invention;
[0045] Figure 3 The figure shows a logical schematic diagram of the environmental perception method of the lawn mowing robot based on embedded vision according to the present invention;
[0046] Figure 4 The figure shows a BoxPR curve graph of instance segmentation training after optimizing the YOLOv11 instance segmentation model;
[0047] Figure 5 The figure shows a MaskPR curve graph of instance segmentation training after optimizing the YOLOv11 instance segmentation model;
[0048] Figure 6 The figure shows a schematic diagram of the inference time and inference results of the lawn on the Horizon RDK_X3_Module chip;
[0049] Figure 7 The figure shows a schematic diagram of the lawn mowing robot in an embodiment applied by the present invention;
[0050] Figure 8 The figure shows a working schematic diagram of the environmental perception method of the lawn mowing robot based on embedded vision according to the present invention. Detailed implementation manners
[0051] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with the embodiments and the drawings, so as to fully understand the purpose, solution and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The same reference numerals used in the drawings indicate the same or similar parts everywhere.
[0052] Embodiment 1, referring to Figure 1 、 Figure 3 and Figure 8 The present invention provides an environmental perception method for a lawn mowing robot based on embedded vision, including the following:
[0053] Step 110: Obtain lawn image data, and perform refined annotation on the lawn image data to obtain an annotation data set;
[0054] Step 120: Expand the annotation data set through a data augmentation strategy to obtain a training data set;
[0055] Step 130: Construct a neural network model based on model lightweight processing. The neural network model is constructed by replacing the activation function of YOLOv11, and the feature extraction model is obtained by training the neural network model with the training dataset.
[0056] Step 140: Optimize the output structure of the YOLOv11 instance segmentation model, and export the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model.
[0057] Step 150: Perform quantization processing on the intermediate model to obtain a quantized model.
[0058] Step 160: Implement intelligent perception and target recognition of the lawn environment of the robot associated with the target embedded chip by writing an inference program for the quantized model and deploying it to the target embedded chip.
[0059] In this Embodiment 1, first, by setting up a camera acquisition device and writing a camera control program, image data of the lawn, boundaries, and common obstacles are collected and uploaded to the Roboflow platform for fine annotation processing to obtain an annotated dataset. Then, data augmentation techniques such as rotating, cropping, changing the hue, saturation, brightness of the image, and adding noise are used to expand the annotated dataset to simulate various change conditions in the real scene and obtain a training dataset. Immediately afterwards, by comparing the performance of different models in terms of recognition accuracy and frame rate (FPS), etc., and aiming at the problem of difficult positioning and accurate recognition of complex boundaries, the YOLOv11 instance segmentation model is selected for target recognition. To achieve model lightweight processing, the activation function of the original model is replaced, and the neural network model based on the lightweight processing of YOLOv11 instance segmentation is trained with the training dataset to obtain a feature extraction model. Then, the output structure of the YOLOv11 instance segmentation model is optimized, and the feature extraction model is exported in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model. Subsequently, based on the PTQ scheme quantization conversion process provided by the official of the purchased embedded chip, a model quantization program is written to compress the intermediate model from FP32 to INT8 format to obtain a quantized model. Finally, by writing an inference program for the quantized model and deploying it to the target embedded chip, intelligent perception and target recognition of the lawn environment by the robot are realized.
[0060] The embedded vision development platform applied in this Embodiment 1 includes: a host computer, an embedded chip and its related accessories, a vision sensor, a power supply, etc. Specifically, as Figure 7 shown.
[0061] As a preferred embodiment of the present invention, specifically, lawn image data is obtained, and the lawn image data is finely annotated to obtain an annotation dataset, including,
[0062] Under different weather conditions, the weather conditions include sunny days, rainy days and nights, lawn image data including lawns, boundaries and obstacles is collected, and then the collected lawn image data is uploaded to the Roboflow platform for fine annotation processing to obtain an annotation dataset.
[0063] In this preferred real-time mode, first, a camera acquisition device is built, and a camera control program is written to achieve automatic shooting. Under different weather conditions, including different conditions such as sunny days, rainy days and nights, image data including lawns, boundaries and obstacles is collected. In order to ensure the diversity of the data, various lighting and environmental changes are deliberately covered during the shooting process. Then, the collected pictures are uploaded to the Roboflow platform for fine annotation processing to obtain an annotation dataset.
[0064] As a preferred embodiment of the present invention, specifically, the annotation dataset is expanded through a data augmentation strategy to obtain a training dataset, including,
[0065] The annotation dataset is expanded through image enhancement technology to obtain a training dataset, and the image enhancement technology includes rotation, cropping, changing the hue, saturation, brightness of the picture, and adding noise.
[0066] In this preferred embodiment, after the annotation is completed, the annotation dataset is expanded through data augmentation technologies such as rotating, cropping, changing the hue, saturation, brightness of the picture, and adding noise to simulate various change conditions in the real scene to obtain a training dataset. The specific method of data augmentation is as shown in Table 1 below:
[0067] Table 1 Lawn dataset enhancement processing method
[0068]
[0069] As a preferred embodiment of the present invention, specifically, the construction process of the neural network model includes,
[0070] Replace the SiLU activation function of YOLOv11 with the ReLU function, and the expression of the ReLU function is,
[0071] f(x) = max(0, x),
[0072] That is, when the input is greater than zero, the output is the input value itself, otherwise the output is zero.
[0073] In this preferred real-time mode, by comparing the accuracy and frames per second (FPS) of different models, and aiming at the problem that it is difficult to locate and accurately identify complex boundaries, the present invention selects the YOLOv11 instance segmentation model to achieve real-time and accurate identification of lawns, boundaries, and obstacles. This model combines the advantages of object detection and semantic segmentation and has strong environmental understanding ability.
[0074] In addition, to better meet the requirements of the embedded scenario, we replace the SiLU activation function of the model with the ReLU function to achieve lightweight processing of the model. The expression of the ReLU (Rectified Linear Unit) function is:
[0075] f(x) = max(0, x),
[0076] that is, when the input is greater than zero, the output is the input value itself, otherwise the output is zero.
[0077] The ReLU activation function is simple and efficient in calculation, alleviating the problem of gradient disappearance; its sparse activation characteristic accelerates training, reduces power consumption and resource occupation, and at the same time prolongs the battery life of the device and optimizes the application performance. After replacing the activation function, the training data set is used to train a neural network model with lightweight processing based on YOLOv11 instance segmentation to obtain a feature extraction model. Its training results are as Figure 4 、 Figure 5 shown. And the number of parameters and the amount of calculation of the model are tested by writing a program. Table 2 shows the specific numerical changes before and after replacing the activation function:
[0078] Table 2 Changes in the number of parameters and the amount of calculation of the model after replacing the activation function
[0079]
[0080]
[0081] By comparing the number of parameters (Parameters) and the number of floating-point operations (FLOPs) of the SiLU and ReLU activation functions, it is revealed that the SiLU activation function is much higher than the ReLU activation function in terms of computational complexity and the number of parameters. In the process of model training and inference, using the SiLU activation function will consume more computational resources and time, and at the same time it is difficult to meet the requirement of achieving real-time performance in embedded devices. The ReLU activation function has lower computational complexity and the number of parameters and is more suitable for embedded devices with limited computing power. The comparison further verifies the rationality of our replacement of the activation function.
[0082] As a preferred implementation manner of the present invention, specifically, optimizing the output structure of the YOLOv11 instance segmentation model includes,
[0083] Adjust the output structure of the Segment module in the output structure of the YOLOv11 instance segmentation model to 10 independent output heads, where the Bounding Box information, Masks information, and Classify information of the three feature layers are output independently.
[0084] In this preferred real-time mode, after the model training is completed, optimize the output structure of the YOLOv11 instance segmentation model, adjust the output structure of the Segment module to 10 independent output heads, where the Bounding Box information, Masks information, and Classify information of the three feature layers are output independently, so as to improve the accuracy and efficiency of the model, and at the same time enhance the adaptability and flexibility. Independent output can refine the processing of each type of information, improve the accuracy of target localization, boundary extraction, and classification, while reducing feature coupling and redundant calculations. Then, export the feature extraction model in the optimized YOLOv11 instance segmentation model to an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model.
[0085] As a preferred embodiment of the present invention, specifically, perform quantization processing on the intermediate model to obtain a quantized model, including,
[0086] Determine the PTQ scheme quantization conversion process based on the selected embedded chip model, write a quantization program for the YOLOv11 instance segmentation model based on the PTQ scheme quantization conversion process, and convert the intermediate model format to a model format adapted to the embedded chip. Specifically,
[0087] First, determine whether the intermediate model can be used normally on the selected embedded chip;
[0088] Then calibrate the model;
[0089] Next, write a configuration file for quantizing the YOLOv11 instance segmentation model, and compress the intermediate model from FP32 to INT8 format to obtain a quantized model.
[0090] In this preferred real-time mode, based on the PTQ scheme quantization conversion process provided by the official of the purchased embedded chip, a quantization program for the YOLOv11 instance segmentation model is written to convert the intermediate model format into a model format adapted to this embedded chip. In the present invention, Horizon RDK_X3_Module is selected as the embedded chip, the intermediate model format is ONNX format, and the model format adapted to the Horizon RDK_X3_Module chip is BIN format. Therefore, according to the model conversion tool chain provided by Horizon official, the ONNX intermediate format is converted into BIN format on the development machine, and other embedded chips can be analogously used with basically the same method. During the quantization process, first, the model is checked to determine whether the model can be used normally on the embedded chip. Then, the model is calibrated. Since the BPU is INT8 calculation, there will be a loss of precision in the model, so calibration is necessary. Next, a configuration file for quantizing the YOLOv11 instance segmentation model is written to compress the intermediate model from FP32 to INT8 format to obtain the quantized model. This process can reduce the storage requirements and computing costs while maintaining a high level of accuracy.
[0091] As a preferred embodiment of the present invention, specifically, an inference program for the quantized model is written, including,
[0092] Using the communication interface of the embedded development platform, the model is transmitted to the specified location of the device and loaded into the inference engine. Memory resources and hardware acceleration parameters are configured during the initialization process. The inference program preprocesses the input image, and then calls the quantized model for real-time inference, outputs information such as bounding boxes, masks, class labels, and confidence levels, and generates a visual result through post-processing.
[0093] In this preferred real-time mode, the quantized model is uploaded to the embedded chip on the embedded vision development platform through means such as VNC remote control and network transmission, and a YOLOv11 instance segmentation inference program is written to implement the loading, inference, and post-processing of the model. Using the communication interface of the embedded development platform, the model is transmitted to the specified location of the device and loaded into the inference engine. Memory resources and hardware acceleration parameters are configured during the initialization process. The inference program preprocesses the input image (such as resizing and format conversion), and then calls the quantized model for real-time inference, outputs information such as bounding boxes, masks, class labels, and confidence levels, and generates a visual result through post-processing, such as Figure 6 shown Figure 6In the figure, the lawn is shown in red and the obstacles are shown in blue, and the inference time is also presented. The total inference time (including post-processing) for each image is 156.91 milliseconds. Compared with the non-lightweight YOLOv11 instance segmentation model, the performance is nearly tripled. The following table shows the comparison of the Forward Time between the unprocessed YOLOv11 public model and the optimized YOLOv11:
[0094] Table 3 Comparison of Forward Time between the unprocessed YOLOv11 public model and the optimized YOLOv11 model
[0095]
[0096] Through serial communication, the inference results recognized by the vision sensor are sent to the control module of the lawn mowing robot in the form of signals. After receiving the inference signals, the lawn mowing robot utilizes the environmental perception information provided by the embedded chip (such as the lawn area, boundary positions, and obstacle distribution), and combines its internal decision-making algorithm to optimize path planning and action execution. The working principle is as Figure 7 shown.
[0097] Referring to Figure 2 , Example 2, the present invention proposes a device for environmental perception of a lawn mowing robot based on embedded vision, including the following:
[0098] The labeled dataset construction module 100 is used to obtain lawn image data and perform refined labeling on the lawn image data to obtain a labeled dataset;
[0099] The training dataset construction module 200 is used to expand the labeled dataset through data augmentation strategies to obtain a training dataset;
[0100] The feature extraction model construction module 300 is used to construct a neural network model based on model lightweight processing. The neural network model is constructed by replacing the activation function of YOLOv11, and the feature extraction model is obtained by training the neural network model with the training dataset;
[0101] The intermediate model construction module 400 is used to optimize the output structure of the YOLOv11 instance segmentation model, and export the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model;
[0102] The quantization module 500 is used to perform quantization processing on the intermediate model to obtain a quantized model;
[0103] The program deployment module 600 is used to write the inference program of the quantized model and deploy it to the target embedded chip, so as to realize the intelligent perception and target recognition of the lawn environment of the robot associated with the target embedded chip.
[0104] In summary, a lawn mowing robot environment perception method and device based on embedded vision proposed by the present invention. By using a vision sensor to collect lawn image data, and adopting data enhancement strategies such as noise addition, hue and saturation adjustment for the image data to comprehensively simulate various changes in the real environment, the problem of difficult recognition due to environmental changes in the past is solved, and the robustness of visual recognition is enhanced. For the problems of difficult positioning and accurate recognition of complex boundaries and obstacles, the present invention selects the YOLOv11 instance segmentation model. This model combines the advantages of object detection and semantic segmentation, can accurately distinguish and segment the contours and categories of each independent object in the image, has strong environmental understanding ability, and by using the ReLU activation function to replace the original SiLU activation function of the YOLOv11 instance segmentation model, effectively reduces the number of model parameters and floating-point operation times, and improves the recognition speed of the algorithm. For problems such as feature coupling and redundant calculations existing in the model, the present invention optimizes the output structure of the YOLOv11 instance segmentation model, improving the inference speed of the model. For the problem of limited computing resources of embedded devices, the present invention compresses the weights and activation values of the model from FP32 to INT8 format through quantization, effectively reducing the size and storage space requirements of the model, and improving the inference speed of the model on the embedded side. Finally, by deploying the model to the embedded chip and writing its inference program, the lawn mowing robot realizes the intelligent perception of the surrounding environment using the vision sensor.
[0105] The present invention simplifies the structure and design complexity of the product to a certain extent. By integrating an advanced vision sensor and an efficient image processing algorithm, it realizes the intelligent perception of the lawn operation environment without relying on a complex mechanical structure or an additional sensor array. In addition, by optimizing and quantizing the model, the storage requirements and computing resources of the model are reduced to a certain extent, realizing the reasonable deployment of the model on the embedded device. Finally, the present invention improves the accuracy and speed of visual recognition to a certain extent, enabling the lawn mowing robot to complete the mowing task more efficiently.
[0106] In addition, in each embodiment of the present invention, each functional module can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0107] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.
[0108] Although the description of the present invention has been quite detailed and several of the described embodiments have been described in particular, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be regarded as providing a broad interpretation of these claims in light of the prior art by reference to the appended claims, thereby effectively covering the intended scope of the present invention. In addition, the present invention has been described above in terms of embodiments foreseeable by the inventors for the purpose of providing a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.
[0109] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as it achieves the technical effects of the present invention by the same means, it should fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and variations can be made to its technical solutions and / or embodiments.
Claims
1. A lawn mowing robot environment perception method based on embedded vision, characterized in that: Includes the following: Acquire lawn image data, and perform fine annotation on the lawn image data to obtain an annotation data set; Expanding the labeled data set through a data enhancement strategy to obtain a training data set; Constructing a neural network model based on model lightweight processing, wherein the neural network model is constructed by replacing the activation function of YOLOv11, and training the neural network model with the training data set to obtain a feature extraction model; Optimizing the output structure of the YOLOv11 instance segmentation model, and exporting the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model; Quantizing the intermediate model to obtain a quantized model; By writing the reasoning program of the quantized model and deploying it on the target embedded chip, the intelligent perception and target recognition of the lawn environment of the target embedded chip-associated robot can be realized.
2. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, lawn image data is obtained, and the lawn image data is finely annotated to obtain an annotated data set, including: Under different weather conditions, including sunny days, rainy days and nights, lawn image data including lawns, boundaries and obstacles are collected, and then the collected lawn image data are uploaded to the Roboflow platform for fine annotation processing to obtain annotated data sets.
3. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, the labeled data set is expanded through a data enhancement strategy to obtain a training data set, including: The labeled data set is expanded by image enhancement technology to obtain a training data set, and the image enhancement technology includes rotating, cropping, changing the hue, saturation, brightness of the image, and adding noise.
4. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, the construction process of the neural network model is as follows: include, Replace the SiLU activation function of YOLOv11 with the ReLU function. The expression of the ReLU function is: f(x)=max(0,x), That is, when the input is greater than zero, the output is the input value itself, otherwise the output is zero.
5. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, we optimize the output structure of the YOLOv11 instance segmentation model, including: The output structure of the Segment module of the YOLOv11 instance segmentation model output structure is adjusted to 10 independent output heads, in which the Bounding Box information, Masks information and Classify information of the three feature layers are output independently.
6. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, the intermediate model is quantized to obtain a quantized model, including: Determine the PTQ scheme quantization conversion process based on the model of the selected embedded chip, write the quantization program of the YOLOv11 instance segmentation model based on the PTQ scheme quantization conversion process, and convert the intermediate model format into a model format suitable for the embedded chip. Specifically, First, determine whether the intermediate model can be used normally on the selected embedded chip; The model is then calibrated; Next, a configuration file for quantizing the YOLOv11 instance segmentation model is written, and the intermediate model is compressed from FP32 to INT8 format to obtain a quantized model.
7. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 6, characterized in that: Specifically, Select Horizon RDK_X3_Module as the embedded chip, the intermediate model format is ONNX format, and the model format adapted to the Horizon RDK_X3_Module chip is BIN format. Therefore, according to the model conversion tool chain officially provided by Horizon, the ONNX intermediate format is converted to BIN format on the development machine.
8. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 1, characterized in that: Specifically, writing the inference program of the quantized model includes: Using the communication interface of the embedded development platform, the model is transferred to the specified location of the device and loaded into the inference engine. During the initialization process, memory resources and hardware acceleration parameters are configured. The inference program preprocesses the input image and then calls the quantized model for real-time inference. It outputs information such as bounding boxes, masks, category labels, and confidence levels, and generates visualization results through post-processing.
9. The method for environmental perception of a lawn mowing robot based on embedded vision according to claim 8, characterized in that: Specifically, the preprocessing includes size adjustment and format conversion.
10. A device for environmental perception of a lawn mowing robot based on embedded vision, characterized in that: Includes the following: A labeling data set construction module is used to obtain lawn image data, and to perform fine labeling on the lawn image data to obtain a labeling data set; A training data set construction module is used to expand the labeled data set through a data enhancement strategy to obtain a training data set; A feature extraction model construction module is used to construct a neural network model based on model lightweight processing, wherein the neural network model is constructed by replacing the activation function of YOLOv11, and the neural network model is trained by the training data set to obtain a feature extraction model; An intermediate model construction module is used to optimize the output structure of the YOLOv11 instance segmentation model, and export the feature extraction model in the optimized YOLOv11 instance segmentation model into an intermediate model format compatible with a specific embedded deployment platform to obtain an intermediate model; A quantization module, used for quantizing the intermediate model to obtain a quantized model; The program deployment module is used to write the reasoning program of the quantized model and deploy it to the target embedded chip to realize the intelligent perception and target recognition of the lawn environment of the target embedded chip-associated robot.