Network model and method suitable for lawn scene environment perception
By adopting a network model of shared encoder in a lawn environment, combining CBAM attention mechanism and pyramid pooling, lawn area prediction and obstacle detection are realized, and the problem of relying on GPS signals or manual layout of markers in the existing technology is solved, and the autonomous operation ability and environmental perception accuracy of the mowing robot are improved.
Patent Information
- Application Number
- CN202510485253.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the prior art automatically mowing grass in a lawn environment, it is necessary to rely on GPS signals or a large number of manually deployed markers, which has problems such as signal loss, marker offset or loss, making it difficult for the mowing robot to accurately identify the lawn range and its own position.
The network model of shared encoder is adopted, and the feature extraction module, feature enhancement module, feature fusion module and detection module are combined with the CBAM attention mechanism and pyramid pooling to realize lawn area prediction and obstacle detection, providing semantic segmentation and object detection functions.
Without relying on GPS signals or manual layout of markers, it can accurately identify lawn areas and obstacles in the lawn environment, improving the autonomous operation ability and environmental perception accuracy of the mowing robot.
Smart Images

Figure CN119992353A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a network model and method suitable for environmental perception of a lawn scene, belonging to the technical field of image processing. Background Art
[0002] As lawn planting areas grow, the demand for mowing tasks also increases. Therefore, how to achieve automatic mowing has become a hot research topic for intelligent mowing robots. One of the computational challenges is how to enable mowing robots to accurately and quickly understand the environmental information in the lawn.
[0003] The existing technologies mainly include the following: The lawn mowing robot is taught by manually controlling the lawn mowing area. During this process, the position is recorded by satellite and a map of the work area is drawn in proportion in the background. The smart mowing robot then ensures that mowing operations are carried out within the designated work area based on the drawn map and its own position.
[0004] One method of special boundary object recognition technology is to place special markers on the boundary and use sensors set on the mowing robot to recognize the markers to determine the working area; Image recognition technology is used to predict the working area of the lawn, and the single-category features in the lawn are identified through the algorithm, providing environmental information for the autonomous operation of the intelligent mowing robot.
[0005] The first method mentioned above needs to be used when the GPS signal is good. Once the GPS signal is lost, the lawn mower robot cannot confirm the lawn range and its own position and will lose its direction.
[0006] The second method requires a large number of special markers to be set up in the lawn boundary area, which requires a lot of manpower and materials. In addition, the special markers may be offset and lost outdoors, causing the lawn map to change and providing wrong information to the mowing robot.
[0007] In the third method, in the image-based recognition method, the single model provides limited environmental information, but it is easy to deploy on embedded devices; when multiple task models are required to complement each other, there is redundancy in model design, which brings huge pressure to the deployment of embedded devices.
[0008] In summary, the existing technology has the following disadvantages: ① requires a specific environment; ② high labor cost; ③ high material usage cost; ④ the environmental information provided is single; ⑤ difficult to deploy on embedded devices. Summary of the invention
[0009] The purpose of the present invention is to overcome the problems existing in the prior art and provide an environment perception model and environment perception method suitable for lawn scenes. By sharing the encoder method, detection heads for different tasks are introduced at the end of the decoder, including two tasks: target detection and semantic segmentation, which are respectively aimed at lawn area prediction and obstacle detection in the lawn.
[0010] To solve the above technical problems, the present invention proposes a network model suitable for lawn scene environment perception, including a feature extraction module, a feature enhancement module, a feature fusion module and a detection module, wherein the feature extraction module is improved by using a lightweight network: a CBAM attention mechanism is embedded in the sampling process with a ghost net backbone network step size of one; the feature fusion module includes extraction network output layers B2, B3 and B4 and a pyramid output layer B5.
[0011] The present invention also proposes an environment perception method suitable for lawn scenes, comprising the following steps: Step 1: Preprocess the lawn environment image and label the collected lawn environment data; Step 2: extracting features from the lawn image preprocessed in step 1 through a feature extraction module; Step 3: Enhance the features extracted in step 2 through the feature enhancement module; Step 4: Improve the inference accuracy of the model through the feature fusion module; Step 5: Match the extracted features with the tasks through the detection module, and use the specific detection module of the corresponding task to implement semantic segmentation and target detection for the final fusion result output in step 4.
[0012] Furthermore, in step 1, semantic segmentation labels and target detection labels are obtained by annotation, wherein the semantic segmentation labels are static features, and the target detection labels are dynamic features.
[0013] Furthermore, the features extracted in step 2 include static features and dynamic features, the static features include lawns, trees, bushes, and stones, and the dynamic features include people, cats, and dogs.
[0014] Furthermore, the image preprocessing in step 1 includes size adjustment, noise removal, color space conversion, and normalization.
[0015] Furthermore, step 3 specifically includes the following steps: Step 3.1: Use convolution to create the spatial features required for the pyramid pooling layer on the features of the last layer of the feature extraction module; Step 3.2: Three 5×5 pooling kernels are used in series to extract feature representations of different scales, and feature maps of three scales are obtained in sequence; Step 3.3: Concatenate the three pooling kernel extraction results with the previously created spatial features for output; Step 3.4: Use convolution operation to smooth the feature conflicts between the four branches and obtain feature fusion representations of different scales, thereby extracting multiple levels of features and realizing feature extraction of objects of different sizes.
[0016] Furthermore, the feature fusion module includes a feature alignment module FAM, an information fusion module IFM and an information injection module IJM, and step 4 specifically includes the following steps: Step 4.1: Use the feature alignment module FAM to unify the features of different levels; Step 4.2: Use the information fusion module IFM to further couple the aligned features; Step 4.3: The enhanced global guidance information is obtained through the feature alignment module FAM and the information fusion module IFM, and the effective fusion of the global guidance information and the original information is achieved through the information injection module IJM.
[0017] Furthermore, step 4.1 specifically includes the following steps: bilinearly upsampling B5 and average pooling B2 and B3 to the same size as B4, and then stacking the three results with B4 for output.
[0018] Furthermore, step 4.3 specifically includes the following steps: Step 4.3.1: Use convolution to perform nonlinear transformation on the original information to obtain branch 1; Step 4.3.2: The global guidance information is obtained through two convolutions to obtain different branches. One branch is normalized by the sigmoid function, then upsampled and multiplied with branch one to obtain branch two. The other branch is added to branch two and then convolved to obtain the final result. Step 4.3.3: Inject the final result obtained in step 4.3.2 into the information injection module IJM module to obtain the first stage fusion information P3, P4, and P5; then obtain N5 after 3×3 convolution of P5. After downsampling, N5 is stacked with P4 and then undergoes 3×3 convolution to obtain N4; After downsampling, N4 is first stacked with P3 and then undergoes 3×3 convolution to obtain N3, from which the final fusion results N3, N4, and N5 are obtained.
[0019] Compared with the prior art, the present invention achieves the following beneficial effects: the ghost net with CBAM attention mechanism is added to improve the feature extraction capability; the pyramid pooling is adapted to targets of different sizes; the feature fusion module introduces global information to enhance the effect of feature fusion; the target detection head and the language segmentation head are used simultaneously to complete two tasks in lawn environment perception at one time. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The accompanying drawings are only provided for reference and explanation and are not intended to limit the present invention.
[0021] Figure 1 The structural diagram of the network model for lawn scene environment perception; Figure 2 Schematic diagram of CBAM attention mechanism; Figure 3 This is the structural diagram of the feature enhancement module; Figure 4 It is the structural diagram of the feature fusion module; Figure 5 This is the structure diagram of the feature alignment module FAM; Figure 6 This is the structure diagram of the information injection module IJM; Figure 7 is a general flow chart of the environment perception method of the present invention; Figure 8 This is a flow chart of step 3 in the present invention; Fig. 9 is a flow chart of step 4 in the present invention; Fig.10 This is a flow chart of step 4.3 in the present invention. DETAILED DESCRIPTION
[0022] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below with reference to specific diagrams.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0024] like Figure 1 As shown, the invention proposes a network model suitable for lawn scene environment perception, including a feature extraction module, a feature enhancement module, a feature fusion module and a detection module. The feature enhancement module is a spatial pyramid module, and the feature fusion module includes a feature alignment module FAM, an information fusion module IFM and an information injection module IJM. Figure 2 As shown in the figure, the feature extraction module is improved by using a lightweight network: an efficient CBAM attention mechanism is embedded in the sampling process of the ghost net backbone network with a step size of one; the feature fusion module includes the extraction network output layers B2, B3 and B4 and the pyramid output layer B5.
[0025] based on Figure 1The network model of the present invention is an environment perception method suitable for lawn scenes, such as Figure 7 As shown, the following steps are included: Step 1: Perform image preprocessing on the lawn environment and annotate the collected lawn environment data; Step 1 image preprocessing includes resizing, noise removal, color space conversion, and normalization. In step 1, semantic segmentation labels and target detection labels are obtained through annotation. The semantic segmentation label is a static feature, and the target detection label is a dynamic feature.
[0026] Step 2: Extract features from the lawn image preprocessed in step 1 through a feature extraction module; the extracted static features include lawn, trees, shrubs, stones, etc., and the dynamic features include pedestrians, cats, dogs, etc. The present invention effectively reduces the amount of calculation during feature extraction by using ghost net, and adds CBAM to further improve the network feature extraction capability.
[0027] like Figure 8 As shown, step 3: the features extracted in step 2 are enhanced by the feature enhancement module; the feature enhancement module is the spatial pyramid module. Figure 3 As shown, Step 3 specifically includes the following steps: Step 3.1: Use convolution to create the spatial features required for the pyramid pooling layer for the features of the last layer of the feature extraction module; pyramid pooling can effectively take care of objects of different sizes; Step 3.2: Three 5×5 pooling kernels are used in series to extract feature representations of different scales, and feature maps of three scales are obtained in turn, which is the maximum pooling kernel in the figure; the scale information includes: k is the pooling window size, s is the step size; p is the padding size; Step 3.3: Concatenate the three pooling kernel extraction results with the previously created spatial features for output; Step 3.4: Use convolution operation to smooth the feature conflicts between the four branches and obtain feature fusion representations of different scales, thereby extracting multiple levels of features and realizing feature extraction of objects of different sizes.
[0028] like Fig. 9 As shown, step 4: improve the reasoning accuracy of the model through the feature fusion module, the feature fusion module includes the feature alignment module FAM, the information fusion module IFM and the information injection module IJM, as shown Figure 4 As shown, step 4 specifically includes the following steps Step 4.1: Use the feature alignment module FAM to make the features of different levels the same size; Step 4.1 specifically includes the following steps: bilinear upsampling B5 and averaging pooling B2 and B3 to the same size as B4, and then stacking the three results with B4 for output. The structure diagram of the feature alignment module FAM is shown in the figure. Figure 5 As shown; Step 4.2: Use the information fusion module IFM to further couple the aligned features; Step 4.3: The enhanced global guidance information is obtained through the feature alignment module FAM and the information fusion module IFM, and the effective fusion of the global guidance information and the original information is achieved through the information injection module IJM. The structure diagram of the information injection module IJM is shown in Figure 6 The present invention effectively solves the problem of missing global guidance information by using multi-scale information and multi-head attention mechanism through feature fusion layer.
[0029] like Fig.10 As shown, step 4.3 specifically includes the following steps: Step 4.3.1: Use convolution to perform nonlinear transformation on the original information to obtain branch 1; Step 4.3.2: The global information is obtained through two convolutions to obtain different branches. One branch is normalized by the sigmoid function, then upsampled and multiplied with branch one to obtain branch two. The other branch is added to branch two and then convolved to obtain the final result. Step 4.3.3: Inject the final result obtained in step 4.3.2 into the information injection module IJM module to obtain the first stage fusion information P3, P4, and P5; then obtain N5 after 3×3 convolution of P5. After downsampling, N5 is stacked with P4 and then undergoes 3×3 convolution to obtain N4; After downsampling, N4 is stacked with P3 and then undergoes 3×3 convolution to obtain N3. The final fusion results N3, N4, and N5 are obtained. The fusion results include the fusion results in the target detection task and the fusion results in the semantic segmentation task. The fusion results in the target detection task output multiple forms such as tuple of is the coordinate of the upper left corner of the detection box, and are the width and height of the detection box respectively, is the category to which the target in the detection box belongs, that is, the category index, is the confidence score that the detection box contains the target. These results represent the location, category and credibility of the target detected by the network in the image. The fusion result in the semantic segmentation task is a segmentation map with the same size as the input image. Each pixel in the segmentation map has a category label or a category probability distribution, which is used to indicate the semantic category to which the pixel belongs. In the present invention, each pixel of the output segmentation map will be marked as a corresponding category such as lawn, stone, etc.
[0030] Step 5: Match the extracted features with the tasks through the detection module, and use the specific detection module of the corresponding task to implement semantic segmentation and target detection for the final fusion result output in step 4.
[0031] The present invention no longer requires manual deployment of any special signs, does not rely on GPS signals, can be run completely offline, and provides the intelligent lawn mowing robot with lawn workable area information and obstacle avoidance information at the same time, solving the problem that the lawn information currently obtained by the intelligent lawn mowing robot is single and lacks coordination.
[0032] The above are only the preferred feasible embodiments of the present invention, which show and describe the basic principles and main features of the present invention and the advantages of the present invention, but do not limit the patent protection scope of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. In addition to the above embodiments, the present invention may also have other implementation modes without departing from the spirit and scope of the present invention. The present invention may also have various changes and improvements, and all technical solutions formed by equivalent substitution or equivalent transformation fall within the protection scope required by the present invention. The protection scope required by the present invention is defined by the attached claims and their equivalents. The technical features not described in the present invention can be realized by or using the existing technology, which will not be repeated here.
Claims
1. A network model suitable for lawn scene environment perception, characterized by: It includes a feature extraction module, a feature enhancement module, a feature fusion module and a detection module, wherein the feature extraction module is improved by using a lightweight network: a CBAM attention mechanism is embedded in the sampling process of the ghost net backbone network with a step length of one; The feature fusion module includes extraction network output layers B2, B3 and B4 and a pyramid output layer B5.
2. An environment perception method suitable for lawn scenes based on the network model described in claim 1, comprising the following steps: Step 1: Preprocess the lawn environment image and label the collected lawn environment data; Step 2: extracting features from the lawn image preprocessed in step 1 through a feature extraction module; Step 3: The feature enhancement module enhances the features extracted in step 2; Step 4: Improve the inference accuracy of the model through the feature fusion module; Step 5: Match the extracted features with the tasks through the detection module, and use the specific detection module of the corresponding task to implement semantic segmentation and target detection for the final fusion result output in step 4.
3. The environment perception method applicable to lawn scenes according to claim 2 is characterized in that: In step 1, semantic segmentation labels and target detection labels are obtained by annotation, wherein the semantic segmentation labels are static features and the target detection labels are dynamic features.
4. The environment perception method applicable to lawn scenes according to claim 3 is characterized in that: The features extracted in step 2 include static features and dynamic features. The static features include lawns, trees, bushes, and stones, and the dynamic features include people, cats, and dogs.
5. The environment perception method applicable to lawn scenes according to claim 2, characterized in that: The image preprocessing in step 1 includes size adjustment, noise removal, color space conversion, and normalization.
6. The environment perception method applicable to lawn scenes according to claim 4, characterized in that: Step 3 specifically includes the following steps: Step 3.1: Use convolution to create the spatial features required for the pyramid pooling layer on the features of the last layer of the feature extraction module; Step 3.2: Three 5×5 pooling kernels are used in series to extract feature representations of different scales, and feature maps of three scales are obtained in sequence; Step 3.3: Concatenate the three pooling kernel extraction results with the previously created spatial features for output; Step 3.4: Use convolution operation to smooth the feature conflicts between the four branches and obtain the fusion representation of features of different scales, so as to extract multiple levels of features and realize feature extraction of objects of different sizes.
7. The environment perception method applicable to lawn scenes according to claim 6, characterized in that: The feature fusion module includes a feature alignment module FAM, an information fusion module IFM and an information injection module IJM. Step 4 specifically includes the following steps: Step 4.1: Use the feature alignment module FAM to unify the features of different levels; Step 4.2: Use the information fusion module IFM to further couple the aligned features; Step 4.3: The enhanced global guidance information is obtained through the feature alignment module FAM and the information fusion module IFM, and the effective fusion of the global guidance information and the original information is achieved through the information injection module IJM.
8. The environment perception method applicable to lawn scenes according to claim 7 is characterized in that: Step 4.1 specifically includes the following steps: bilinearly upsampling B5 and average pooling B2 and B3 to the same size as B4, and then stacking the three results with B4 for output.
9. The environment perception method applicable to lawn scenes according to claim 7, characterized in that: Step 4.3 specifically includes the following steps: Step 4.3.1: Use convolution to perform nonlinear transformation on the original information to obtain branch 1; Step 4.3.2: The global guidance information is obtained through two convolutions to obtain different branches. One branch is normalized by the sigmoid function, then upsampled and multiplied with branch one to obtain branch two. The other branch is added to branch two and then convolved to obtain the final result. Step 4.3.3: Inject the final result obtained in step 4.3.2 into the information injection module IJM module to obtain the first stage fusion information P3, P4, and P5; then obtain N5 after 3×3 convolution of P5. After downsampling, N5 is stacked with P4 and then undergoes 3×3 convolution to obtain N4; After downsampling, N4 is first stacked with P3 and then undergoes 3×3 convolution to obtain N3, from which the final fusion results N3, N4, and N5 are obtained.
Citation Information
Patent Citations
Unmanned aerial vehicle target detection model based on YOLOv5 network
CN114612835A
Automatic driving multi-task visual perception method
CN116824537A
Intelligent driving multi-task visual environment perception method
CN119693920A