High-speed sweeper target identification method and system based on machine vision
By integrating speed information in the intelligent sweeper for adaptive chunking and image enhancement processing, the problem of image blurring at high speed movement is solved, and the target recognition accuracy and processing efficiency are improved.
Patent Information
- Application Number
- CN202510737791.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
When the intelligent sweeper moves at high speed, image acquisition is prone to blur, resulting in a decrease in garbage recognition accuracy. The existing recognition model increases the data processing volume, affecting the continuous performance of the sweeper.
By acquiring real-time image data and integrating vehicle speed information, adaptive chunking processing and image enhancement are performed, and target recognition is performed using the MobileNetv3 network.
It improves the clarity of the detailed features in the image and the target recognition accuracy, reduces the data processing volume, and adapts to the target recognition task of high-speed sweepers.
Smart Images

Figure CN120259950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine vision high-speed cleaning, in particular to a method and system for target recognition of a high-speed cleaning vehicle based on machine vision. Background Art
[0002] Currently, machine vision technology has been widely applied to intelligent cleaning vehicles, enabling the intelligent cleaning vehicles to perform targeted cleaning on the garbage in front of the operation route and improving the intelligent level of automatic cleaning of the cleaning vehicles.
[0003] However, during the current intelligent cleaning process, when the moving speed of the cleaning vehicle exceeds a certain level, it is likely to cause the collected images to be blurred, thereby reducing the accuracy of subsequent garbage recognition. And if a recognition model with corresponding performance is used, it will greatly increase the data processing volume or data transmission volume during the task execution of the intelligent cleaning vehicle, affecting the continuous performance of the cleaning vehicle. Therefore, it is highly necessary to propose a machine vision target recognition method that can adapt to high-speed cleaning vehicles. Summary of the Invention
[0004] In view of the above problems, the present invention aims to provide a method and system for target recognition of a high-speed cleaning vehicle based on machine vision.
[0005] The object of the present invention is achieved by the following technical solutions: In a first aspect, the present invention proposes a method for target recognition of a high-speed cleaning vehicle based on machine vision, including the following steps: S1 Obtain real-time image data of the vehicle operation area, where the obtained real-time image data carries real-time vehicle speed information; S2 Perform adaptive block processing on the obtained real-time image data, divide the real-time image data into multiple sub-image blocks; perform image enhancement processing on the obtained sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks; S3 Based on the enhanced image data, use the trained target recognition model to process the enhanced image data to obtain the target recognition result of the operation area.
[0006] Preferably, step S1 includes: Obtain the real-time speed information v(t) of the cleaning vehicle based on the on-vehicle IMU; Collect real-time image data X(t) of the cleaning vehicle operation area based on the camera; Integrate the real-time speed information into the real-time image data.
[0007] Preferably, step S2 specifically includes: S21 Extract real-time image frames and corresponding real-time speed information from the obtained real-time image data; S22 performs adaptive block processing on the acquired real-time image frame, dividing the real-time image frame into multiple sub-image blocks; S23 performs image enhancement processing on each sub-image block respectively to obtain enhanced sub-image blocks; S24 reconstructs based on the enhanced sub-image blocks to obtain an enhanced real-time image frame; Preferably, in step S21, real-time image frames are extracted from the real-time image data according to a set time interval.
[0008] Preferably, in step S22, the adaptive block processing based on the acquired real-time image frame specifically includes: Dividing the real-time image frame into N sub-image blocks with equal shapes and areas; 1) For each sub-image block, calculate the complexity factor of the current sub-image block; 2) Calculate the division parameter value of the sub-image block according to the complexity factor of the current sub-image block and the current speed information; 3) Compare the division parameter value of the current sub-image block with a set division threshold. When the division parameter value of the sub-image block is [condition not specified in the original], further divide the current sub-image block into M sub-image blocks with equal shapes and areas; Repeat the above steps 1)-3) until there are no sub-image blocks that can be further divided or the size of the sub-image block is smaller than a preset minimum size, and complete the division of the sub-image blocks of the real-time image frame.
[0009] Preferably, in step S22, the complexity factor calculation function adopted is: ; In the formula, C i represents the complexity factor of the i th sub-image block φ i , N i represents the total number of pixel points of the i th sub-image block, (x, y) ∈ φ i represents the pixel point (x, y) which is the pixel point in the sub-image block φ i ; Sobel(x, y) represents the Sobel eigenvalue extracted with the pixel point (x, y) as the center through the Sobel operator; λ represents a set balance factor, where λ ∈ [0.4, 0.6] ; LBP(x, y) represents the LBP eigenvalue extracted with the pixel point (x, y)The local binary feature values extracted centered; The partitioning parameter value calculation function used is: ; In the formula, S i represents the i th sub-image block φ i 's partitioning parameter value, v(t) represents the real-time speed, v max represents the set maximum speed, C i represents the i th sub-image block φ i 's complexity factor, C TH represents the set complexity factor threshold.
[0010] Preferably, in step S23, image enhancement processing is performed on each sub-image block respectively, specifically including: Statistical analysis is performed on the gray values of each pixel point in the sub-image block to obtain the gray statistical histogram φ i of the sub-image block T i = {p(n)} , where p(n) represents the statistical probability value of the pixel points with gray value i in the n th sub-image block; For the obtained gray statistical histogram T i cropping processing is performed, and the cropping function used is: ; Among them, p cro (n) represents the statistical probability value of the pixel points with gray value n after cropping processing, represents the cropping threshold, , , S i represents the i th sub-image block's total number of pixel points, γ represents the localization factor, where γ ∈ [0.01, 0.05] , α i represents the dynamic limit factor, v(t) represents the real-time speed, v max represents the set maximum speed, Ci represents the i sub - image block φ i complexity factor, C TH represents the set complexity factor threshold; d i represents the pixel distance from the center point of the i sub - image block to the center point of the image; D represents the diagonal pixel distance of the image; μ represents the set limit adjustment parameter, where μ ∈ [0.5, 2] ; Perform compensation processing on the cropped grayscale statistical histogram, and the compensation processing function used is: ; In the formula, p cps (n) represents the statistical probability value of the pixel point with the grayscale value of n after compensation processing; p cro (n) represents the statistical probability value of the pixel point with the grayscale value of n after cropping processing; j represents a variable, p(j) represents the statistical probability value of the pixel point with the grayscale value of j , β represents the cropping threshold; Perform grayscale equalization processing on the sub - image block according to the compensated grayscale statistical histogram to obtain the enhanced sub - image block, and the equalization processing function used is: ; In the formula, represents the updated grayscale value of the pixel point with the grayscale value of v in the i sub - image block after grayscale equalization processing; round represents the rounding function, S i The i total number of pixel points in the sub - image block, CDF(v) represents the cumulative distribution statistical value of the grayscale value of v in the compensated grayscale statistical histogram.
[0011] Preferably, in step S24, after completing the enhancement processing of each sub - image block, merge the enhanced sub - image blocks to obtain the complete enhanced real - time image frame.
[0012] Preferably, in step S3, the enhanced real-time image frame is input into the trained target recognition model to process the enhanced image data, and the corresponding garbage recognition result is output by the target recognition model; Among them, the target recognition model is built based on the MobileNetv3 network, including an input layer, a first convolutional layer, a bottleneck layer, and an output layer; the input layer is used to input the enhanced real-time image frame, the first convolutional layer uses a common convolutional kernel, where the convolutional kernel size is 3×3, the stride is 2, and the activation function is ReLU; the bottleneck layer contains multiple (for example, 15) stacked convolutional modules, where the depth convolutional kernel used is 3×3 or 5×5, the stride is 1 or 2, the activation function is ReLU, and the expansion ratio is 1-6; an attention module is also set in the convolutional module; the output layer contains a fully connected module, where the fully connected module uses global average pooling operation, and the classification function uses the Softmax function to output the garbage recognition result.
[0013] In a second aspect, the present invention proposes a target recognition system for a high-speed sweeper based on machine vision, including: An acquisition module, configured to acquire real-time image data of the vehicle operation area, where the acquired real-time image data carries real-time vehicle speed information; An enhancement module, configured to perform adaptive block processing according to the acquired real-time image data, divide the real-time image data into multiple sub-image blocks; perform image enhancement processing on the acquired sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks; A recognition module, configured to process the enhanced image data based on the enhanced image data by using a trained target recognition model to obtain the target recognition result of the operation area.
[0014] The beneficial effects of the present invention are: A target recognition method and system for a high-speed sweeper based on machine vision are proposed. By integrating the real-time moving speed information of the sweeper into the corresponding real-time image data during the process of collecting real-time image data of the operation area. Adaptive block processing is performed according to the acquired real-time speed information and the feature area of the image, and targeted enhancement processing is performed based on the sub-image blocks, thereby improving the clarity of the detailed feature parts in the image. Finally, the target recognition processing of the operation area is completed based on the enhanced image data through the image processing model.
[0015] In view of the problem that the images collected by the intelligent cleaning vehicle at high moving speeds are prone to blurring, the above-described embodiments of the present invention can, through the proposed enhancement processing method, first perform adaptive block processing on the parts of the image where features are concentrated. During the block processing, speed information is further added to intelligently adjust the size of the blocks, so as to perform targeted enhancement processing on the feature regions. At the same time, an image enhancement processing method based on sub-image blocks is proposed, which can perform adaptive enhancement processing on the sub-image blocks based on speed features and complexity features, improving the performance of local detail features in the image. Thereby, the accuracy of subsequent target recognition based on image data is improved, and the data processing efficiency of the enhancement processing is also effectively improved, enabling it to adapt to the target recognition tasks during the operation of high-speed cleaning vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention will be further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the following drawings without creative efforts.
[0017] Figure 1 FIG. is a flowchart of the steps of a method for target recognition of a high-speed cleaning vehicle based on machine vision according to an embodiment of the present invention; Figure 2 is Figure 1 a detailed flowchart of step S2 in the embodiment; Figure 3 FIG. is a framework structure diagram of a target recognition system for a high-speed cleaning vehicle based on machine vision according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present invention will be further described in combination with the following application scenarios.
[0019] See Figure 1 , which shows a method for target recognition of a high-speed cleaning vehicle based on machine vision, including the following steps: S1 Obtain real-time image data of the vehicle operation area, where the obtained real-time image data carries the vehicle's real-time speed information; S2 Perform adaptive block processing on the obtained real-time image data, dividing the real-time image data into multiple sub-image blocks; perform image enhancement processing on the obtained sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks; S3 Based on the enhanced image data, use the trained target recognition model to process the enhanced image data to obtain the target recognition result of the operation area.
[0020] In the above embodiment of the present invention, a method for target recognition of a high-speed sweeper based on machine vision is proposed. During the process of collecting real-time image data of the operation area, the real-time moving speed information of the sweeper is integrated into the corresponding real-time image data. Adaptive block processing is performed according to the obtained real-time speed information and the feature regions of the image, and targeted enhancement processing is performed based on the sub-image blocks, thereby improving the clarity of the detailed feature parts in the image. Finally, based on the enhanced image data, the target recognition processing of the operation area is completed through an image processing model.
[0021] In an exemplary scenario, the above-mentioned method and system for target recognition of a high-speed sweeper based on machine vision proposed by the present invention can be built based on the built-in data processing device of the high-speed sweeper or a cloud server, etc.
[0022] Preferably, step S1 includes: Obtaining the real-time speed information v(t) of the sweeper based on the vehicle-mounted IMU; Collecting the real-time image data X(t) of the operation area of the sweeper based on the camera; Integrating the real-time speed information into the real-time image data.
[0023] In the above embodiment of the present invention, a camera is set in front of the sweeper to align with the operation area of the sweeper, and the image data of the operation area of the sweeper is collected in real time through the camera. At the same time, the real-time speed data of the sweeper can be synchronously collected based on the IMU built in the sweeper.
[0024] Preferably, referring to Figure 2 , step S2 includes: S21 Extracting the real-time image frame X(t) and the corresponding real-time speed information v(t) from the obtained real-time image data; S22 Performing adaptive block processing on the obtained real-time image frame, and dividing the real-time image frame into multiple sub-image blocks; S23 Performing image enhancement processing on each sub-image block respectively to obtain the enhanced sub-image blocks; S24 Reconstructing based on the enhanced sub-image blocks to obtain the enhanced real-time image frame X ’ (t).
[0025] In the above embodiments of the present invention, aiming at the situation that the images collected at high moving speeds of the intelligent cleaning vehicle are prone to blurred images, the above embodiments of the present invention can first perform adaptive block processing on the parts where features are aggregated in the image. During the block division process, speed information is further added to intelligently adjust the size of the blocks, so as to perform targeted enhancement processing on the feature regions. At the same time, an image enhancement processing method based on sub-image blocks is proposed, which can perform adaptive enhancement processing on sub-image blocks based on speed features and complexity features, improving the performance of local detail features in the image. Thus, the accuracy of subsequent target recognition based on image data is improved, and the data processing efficiency of the enhancement processing is also effectively improved, being able to adapt to the target recognition tasks during the operation of the high-speed cleaning vehicle.
[0026] Preferably, in step S21, real-time image frames are extracted from the real-time image data according to a set time interval.
[0027] In an exemplary scenario, the time interval can be set to 16 - 200 ms in combination with the comprehensive computing power and the reaction speed of the cleaning vehicle.
[0028] Preferably, in step S22, adaptive block processing is performed based on the obtained real-time image frames, specifically including: Initialize the number of division layers T = 1, and divide the real-time image frame into N sub-image blocks with equal shapes and areas; 1) For each sub-image block, calculate the complexity factor of the current sub-image block, where the complexity factor calculation function used is:
[0029] In the formula, C i represents the i th sub-image block φ i 's complexity factor, N i represents the total number of pixel points of the i th sub-image block, (x, y) ∈ φ i represents the pixel point (x, y) being the pixel point in the sub-image block φ i ; Sobel(x, y) represents the Sobel eigenvalue extracted with the pixel point (x, y) as the center through the Sobel operator; λ represents a set balance factor, where λ ∈ [0.4, 0.6] ; LBP(x, y) represents the local binary eigenvalue extracted with the pixel point (x, y) as the center through the LBP operator; 2) Calculate the partitioning parameter value of the sub-image block according to the complexity factor and current speed information of the current sub-image block, where the partitioning parameter value calculation function used is:
[0030] In the formula, S i represents the i th sub-image block φ i 's partitioning parameter value, v(t) represents the real-time speed, v max represents the set maximum speed, C i represents the i th sub-image block φ i 's complexity factor, C TH represents the set complexity factor threshold; 3) Compare according to the partitioning parameter value of the current sub-image block and the set partitioning threshold S TH . When the partitioning parameter value of the sub-image block S i >S TH , further divide the current sub-image block into M sub-image blocks with equal shape and area; Update the decomposition layer number T = T + 1, and repeat the above steps 1)-3) until there are no sub-image blocks that can be further divided or the size of the sub-image block is smaller than the preset minimum size, and complete the sub-image block partitioning of the real-time image frame.
[0031] Among them, when performing the first image partitioning, the value of N can be set to 1-9, and in subsequent partitionings, the value of M can also be set to 1-9 according to actual needs.
[0032] Among them, the specific calculation method of the Sobel eigenvalue is to calculate the vertical Sobel eigenvalue (x, y) and the horizontal Sobel eigenvalue G x (x, y) and G y (x, y) respectively by aligning the pixel point with the Sobel operator in the vertical and horizontal directions, and then further comprehensively calculating to obtain the Sobel eigenvalue ; Among them, the calculation method of the local binary eigenvalue is , where s(h(x, y) - h(i)) represents a threshold function. When (h(x, y) - h(i)) > 0 , s(h(x, y) - h(i)) = 1 , otherwise s(h(x, y) - h(i)) = 0 , h(x, y) represents the grayscale feature value of the pixel point (x, y) . h(i) represents the grayscale feature value of the (x, y) th neighborhood pixel point centered on the pixel point i .
[0033] Preferably, according to experience, the complexity factor threshold C TH can be set in the range of 20-70 , and the optimal value can be selected as C TH = 50 ; The maximum speed v max can be set in the range of 15 - 20 m / s , that is 18 - 72 km / h .
[0034] In the above embodiments of the present invention, when performing adaptive block processing on the extracted image frames, the feature complexity of the image is calculated for the edge features and texture features of each sub-image block, and the dispersion and aggregation levels of the feature information in the image are represented by the feature complexity; further, the sub-image blocks are further divided based on the complexity factor and the real-time speed feature. Among them, when the speed is faster or the complexity is higher, the image block is divided into a finer size for targeted enhancement processing (which is more helpful for restoring the detailed features damaged by motion blur and improving the characterization level of feature details) in the subsequent enhancement processing, improving the fineness of the enhancement processing. At the same time, for general areas, subsequent enhancement processing is based on larger image blocks as the basis, thereby avoiding the occurrence of redundant enhancement and ensuring the efficiency of the enhancement processing. Through the above adaptive sub-image block division method, the effect of subsequent enhancement processing and the data processing volume can be balanced.
[0035] Preferably, in step S23, image enhancement processing is performed on each sub-image block respectively, specifically including: Statistical analysis is performed on the grayscale values of the pixel points in the sub-image block to obtain the grayscale statistical histogram φ i of the sub-image block T i = {p(n)} , where p(n) represents the statistical probability value of the pixel points with the grayscale value of i in the n th sub-image block; For the obtained grayscale statistical histogram T i Perform cropping processing, where the cropping function used is:
[0036] Among them, p cro (n) represents the statistical probability value of the pixel points with grayscale value n after cropping processing, represents the cropping threshold, , , S i represents the i total number of pixel points in the th sub-image block, γ represents the localization factor, where γ ∈ [0.01, 0.05] , α i represents the dynamic limit factor, v(t) represents the real-time speed, v max represents the set maximum speed, C i represents the i th sub-image block φ i complexity factor, C TH represents the set complexity factor threshold; d i represents the i pixel distance from the center point of the th sub-image block to the center point of the image; D represents the diagonal pixel distance of the image; μ represents the set limit adjustment parameter, where μ ∈ [0.5, 2] ; Perform compensation processing on the cropped grayscale statistical histogram, where the compensation processing function used is:
[0037] In the formula, p cps (n) represents the statistical probability value of the pixel points with grayscale value n after compensation processing; p cro (n) represents the statistical probability value of the pixel points with grayscale value n after cropping processing; j represents a variable, p(j) represents the statistical probability value of the pixel points with grayscale value j , β represents the cropping threshold; Perform gray-level equalization processing on the sub-image blocks according to the compensated gray-level statistical histogram to obtain the enhanced sub-image blocks, where the equalization processing function used is:
[0038] In the formula, represents the updated gray value of the pixel points with gray value v in the i th sub-image block after completing the gray-level equalization processing; round represents the rounding function, S i the total number of pixel points in the i th sub-image block, CDF(v) represents the cumulative distribution statistical value of the gray value v in the compensated gray-level statistical histogram.
[0039] Among them, the above d i can also represent the distance from the sub-image block to the center point of the operation area according to the actual situation, where the operation area or the center point of the operation area can be set in advance by artificial calibration according to the setting position or the picture of the camera.
[0040] In the above embodiment of the present invention, after completing the division of the sub-image blocks, an image enhancement processing for the sub-image blocks is also specifically proposed. First, the gray-level statistical histogram is statistically calculated for the gray-level feature information in the image, and the part where the gray levels converge in the histogram is specifically trimmed and compensated. During the process of histogram trimming, the influence of the speed feature on the gray-level feature of the image is particularly considered (the image blur caused by high-speed movement makes the gray-level feature of the image more concentrated, resulting in a decrease in the contrast of the detail part), the trimming effect is intelligently adjusted, and further gray-level equalization is performed on the image, specifically stretching the gray levels of the blurred part and the feature detail part, so as to effectively repair the blurred area in the image and highlight the feature detail part.
[0041] Preferably, in step S24, after completing the enhancement processing of each sub-image block, the enhanced sub-image blocks are merged according to the enhanced sub-image blocks to obtain a complete enhanced real-time image frame.
[0042] After completing the enhancement processing of each sub-image, the sub-image blocks are merged to obtain an enhanced real-time image frame as the basis for subsequent target recognition.
[0043] Preferably, in step S3, the enhanced real-time image frame is input into the trained target recognition model to process the enhanced image data, and the target recognition model outputs the corresponding garbage recognition result.
[0044] Preferably, the target recognition model is built based on the MobileNetv3 network, including an input layer, a first convolutional layer, a bottleneck layer, and an output layer; wherein the input layer is used to input enhanced real-time image frames, the first convolutional layer uses ordinary convolutional kernels, where the convolutional kernel size is 3×3, the stride is 2, and the activation function is ReLU; the bottleneck layer contains multiple (for example, 15) stacked convolutional modules, where the depth convolutional kernels used are 3×3 or 5×5, the stride is 1 or 2, the activation function is ReLU, and the expansion ratio is 1-6; wherein an attention module is also set in the convolutional module; the output layer contains a fully connected module, where the fully connected module uses global average pooling operation, and the classification function uses the Softmax function to output the garbage recognition result.
[0045] In an alternative embodiment, the specific structure of the MobileNetv3 network is as follows: INPUT: Input enhanced real-time image frames; Conv2D: Convolution kernel size is 3×3, stride is 2, activation function is ReLU; Bottleneck1: Expansion ratio is 1, convolution kernel is 3×3, stride is 1, activation function is ReLU, with an attention module; Bottleneck2: Expansion ratio is 4, convolution kernel is 3×3, stride is 2, activation function is ReLU; Bottleneck3: Expansion ratio is 3, convolution kernel is 3×3, stride is 1, activation function is ReLU; Bottleneck4: Expansion ratio is 3, convolution kernel is 5×5, stride is 2, activation function is ReLU, with an attention module; Bottleneck5: Expansion ratio is 3, convolution kernel is 5×5, stride is 1, activation function is ReLU, with an attention module; Bottleneck6: Expansion ratio is 3, convolution kernel is 5×5, stride is 1, activation function is ReLU, with an attention module; Bottleneck7: Expansion ratio is 6, convolution kernel is 3×3, stride is 2, activation function is ReLU; Bottleneck8: Expansion ratio is 2, convolution kernel is 3×3, stride is 1, activation function is ReLU; Bottleneck9: Expansion ratio is 2, convolution kernel is 3×3, stride is 1, activation function is ReLU; Bottleneck10: Expansion ratio is 2, convolution kernel is 3×3, stride is 1, activation function is ReLU; Bottleneck11: Expansion ratio is 6, convolution kernel is 3×3, stride is 1, activation function is ReLU, with an attention module; Bottleneck12: Expansion ratio is 6, convolution kernel is 3×3, stride is 1, activation function is ReLU, with an attention module; Bottleneck13: Expansion ratio is 6, convolution kernel is 5×5, stride is 2, activation function is ReLU, with an attention module; Bottleneck14: Expansion ratio is 6, convolution kernel is 5×5, stride is 1, activation function is ReLU, with an attention module; Bottleneck15: Expansion ratio is 6, convolution kernel is 5×5, stride is 1, activation function is ReLU, with an attention module; Conv2D: Convolution kernel is 1×1; AvgPool: Average pooling operation, pooling kernel is 7×7; Conv2D: Convolution kernel is 1×1; Conv2D: Classifier, classification function is Softmax; OUTPUT: Output classification results.
[0046] Preferably, for the training of the target recognition model, the model can be trained by constructing a standard data set (including image data and corresponding recognition results), and after the training is completed, the accuracy test can be completed through the test set to obtain the trained target recognition model.
[0047] In the above embodiments of the present invention, a target recognition model based on the MobileNetv3 network is also proposed. Based on the lightweight MobileNetv3 network, it can accurately recognize the targets (such as garbage, etc.) in the image, and improve the target recognition speed and accuracy.
[0048] Preferably, the method further includes: Controlling the cleaning system to complete targeted cleaning tasks according to the target recognition result.
[0049] After the target recognition of the operation area is completed, when garbage or other targets that need to be cleaned are recognized, the cleaning equipment carried on the cleaning vehicle is further controlled to complete the cleaning of the target, so as to complete the target cleaning task.
[0050] See Figure 3 As shown in the embodiment, it shows a target recognition system for a high-speed cleaning vehicle based on machine vision, including: An acquisition module, configured to acquire real-time image data of the vehicle operation area, where the acquired real-time image data carries vehicle real-time speed information; An enhancement module, configured to perform adaptive block processing on the acquired real-time image data, divide the real-time image data into multiple sub-image blocks; perform image enhancement processing on the acquired sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks; A recognition module, configured to process the enhanced image data by using the trained target recognition model based on the enhanced image data to obtain the target recognition result of the operation area.
[0051] Preferably, the system further includes a control module, configured to control the cleaning system to complete targeted cleaning tasks according to the target recognition result.
[0052] At the same time, each module in the above-mentioned target recognition system for a high-speed cleaning vehicle based on machine vision is also used to implement the corresponding method steps in the target recognition method for a high-speed cleaning vehicle based on machine vision as described above Figure 1 This is not repeated in the present invention.
[0053] It should be noted that in each embodiment of the present invention, each functional unit / module can be integrated into one processing unit / module, or each unit / module can exist physically alone, or two or more units / module can be integrated into one unit / module. The above integrated unit / module can be implemented in the form of hardware or in the form of a software functional unit / module.
[0054] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments described herein can be implemented by hardware, software, firmware, middleware, code, or any appropriate combination thereof. For hardware implementation, the processor can be implemented in one or more of the following units: application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, or other electronic units designed to implement the functions described herein, or a combination thereof. For software implementation, part or all of the processes of the embodiments can be completed by instructing relevant hardware through a computer program. When implemented, the above program can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a computer. The computer-readable medium can include, but is not limited to, RAM, ROM, EEPROM, CD-ROM, or other optical disc storage, magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A target recognition method for a high-speed sweeper based on machine vision, characterized in that, It includes the following steps: S1 Obtain the real-time image data of the vehicle operation area, where the obtained real-time image data carries the real-time vehicle speed information; S2 Perform adaptive block processing on the obtained real-time image data, and divide the real-time image data into multiple sub-image blocks; Perform image enhancement processing on the obtained sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks. Specifically, it includes: S21 Extract the real-time image frames and the corresponding real-time speed information from the obtained real-time image data; S22 Perform adaptive block processing based on the obtained real-time image frames, and divide the real-time image frames into multiple sub-image blocks; S23 Perform image enhancement processing on each sub-image block respectively to obtain enhanced sub-image blocks; S24 Reconstruct based on the enhanced sub-image blocks to obtain an enhanced real-time image frame; S3 Based on the enhanced image data, use the trained target recognition model to process the enhanced image data to obtain the target recognition result of the operation area.
2. The object recognition method of a high-speed sweeper based on machine vision according to claim 1, characterized in that, Step S1 includes: Obtain the real-time speed information of the sweeper based on the in-vehicle IMU; Collect the real-time image data of the sweeper operation area based on the camera; Integrate the real-time speed information into the real-time image data.
3. The object recognition method of a high-speed sweeper based on machine vision according to claim 2, characterized in that, In step S21, extract the real-time image frames from the real-time image data according to the set time interval.
4. A method for target recognition of a high-speed cleaning vehicle based on machine vision according to claim 2, characterized in that, In step S22, perform adaptive block processing based on the obtained real-time image frames. Specifically, it includes: Divide the real-time image frames into N sub-image blocks with equal shapes and areas; 1) For each sub-image block, calculate the complexity factor of the current sub-image block; 2) Calculate the division parameter value of the sub-image block according to the complexity factor of the current sub-image block and the current speed information; 3) Compare according to the partitioning parameter value of the current sub-image block and the set partitioning threshold. When the partitioning parameter value of the sub-image block is [condition not provided in the original], further partition the current sub-image block into M sub-image blocks with equal shapes and areas; Repeat the above steps 1)-3) until there are no sub-image blocks that can be further divided or the size of the sub-image block is smaller than the preset minimum size, and complete the division of the sub-image blocks of the real-time image frame.
5. A method for target recognition of a high-speed cleaning vehicle based on machine vision according to claim 4, characterized in that, In step S22, the complexity factor calculation function used is: ; In the formula, C i represents the i th sub-image block φ i 's complexity factor, N i represents the total number of pixel points of the i th sub-image block, (x,y) ∈ φ i represents the pixel point (x,y) is the pixel point in the sub-image block φ i ; Sobel(x,y) represents the Sobel eigenvalue extracted with the pixel point (x,y) as the center by the Sobel operator; λ represents the set balance factor, where λ ∈ [0.4,0.6] ; LBP(x,y) represents the local binary eigenvalue extracted with the pixel point (x,y) as the center by the LBP operator; The division parameter value calculation function used is: ; In the formula, S i represents the i sub-image block φ i partition parameter value, v(t) represents the real-time speed, v max represents the set maximum speed, C i represents the i sub-image block φ i complexity factor, C TH represents the set complexity factor threshold.
6. The object recognition method of a high-speed sweeper based on machine vision according to claim 5, characterized in that, In step S23, perform image enhancement processing on each sub-image block respectively. Specifically, it includes: Statistical analysis is performed based on the gray values of each pixel in the sub-image block to obtain the sub-image block φ i gray-scale statistical histogram T i = {p(n)} , where p(n) represents the statistical probability value of the pixel points with the gray value of i in the n th sub-image block; For the obtained grayscale statistical histogram T i Perform cropping processing, and the cropping function used is as follows: ; Among them, p cro (n) represents the statistical probability value of the pixel points with a gray value of n after the cropping process, and represents the cropping threshold, , , S i represents the total number of pixel points in the i th sub-image block, γ represents the localization factor, where γ ∈ [0.01,0.05] , α i represents the dynamic limit factor, v(t) represents the real-time speed, v max represents the set maximum speed, C i represents the i th sub-image block φ i 's complexity factor, C TH represents the set complexity factor threshold; d i represents the pixel distance from the center point of the i th sub-image block to the center point of the image; D represents the diagonal pixel distance of the image; μ represents the set limit adjustment parameter, where μ ∈ [0.5,2] ; Perform compensation processing on the cropped grayscale statistical histogram. The compensation processing function used is: ; In the formula, p cps (n) represents the statistical probability value of the pixel point with the gray value of n after the compensation process; p cro (n) represents the statistical probability value of the pixel point with the gray value of n after the clipping process; j represents a variable, p(j) represents the statistical probability value of the pixel point with the gray value of j ; β represents the clipping threshold; Perform grayscale equalization processing on the sub-image block according to the compensated grayscale statistical histogram to obtain the enhanced sub-image block. The equalization processing function used is: ; In the formula, represents the updated gray value of the pixel points with gray value v in the i th sub-image block after gray level equalization processing; round represents the rounding function, S i the i th sub-image block, the total number of pixel points, CDF(v) represents the cumulative distribution statistical value of the gray value v in the compensated gray statistical histogram.
7. A method for target recognition of a high-speed sweeper based on machine vision according to claim 6, characterized in that, In step S24, after completing the enhancement processing of each sub-image block respectively, merge according to the enhanced sub-image blocks to obtain the complete enhanced real-time image frame.
8. A method for target recognition of a high-speed sweeper based on machine vision according to claim 1, characterized in that, In step S3, input the enhanced real-time image frame into the trained target recognition model to process the enhanced image data, and the target recognition model outputs the corresponding garbage recognition result; Among them, the target recognition model is built based on the MobileNetv3 network, including an input layer, a first convolutional layer, a bottleneck layer, and an output layer. The input layer is used to input enhanced real-time image frames. The first convolutional layer uses ordinary convolutional kernels with a size of 3×3, a stride of 2, and a ReLU activation function. The bottleneck layer contains multiple stacked convolutional modules, where the depth convolutional kernels used are 3×3 or 5×5, the stride is 1 or 2, the activation function is ReLU, and the expansion ratio is 1-6. An attention module is also set in the convolutional module. The output layer contains a fully connected module, where the fully connected module uses global average pooling operation, the classification function uses the Softmax function, and the garbage recognition result is output.
9. A target recognition system for a high-speed road sweeper based on machine vision, characterized in that, Including: An acquisition module, which is used to acquire real-time image data of the vehicle operation area, and the acquired real-time image data carries vehicle real-time speed information; An enhancement module, which is used to perform adaptive block processing on the acquired real-time image data, divide the real-time image data into multiple sub-image blocks, perform image enhancement processing on the acquired sub-image blocks, and obtain enhanced image data based on the enhanced sub-image blocks; A recognition module, which is used to process the enhanced image data based on the enhanced image data by using the trained target recognition model to obtain the target recognition result of the operation area.
Citation Information
Patent Citations
Real-time enhanced processing system for foggy continuous video image
CN102611828A
Road background extraction and updating method with fusion of real-time traffic state information
CN104077757A
Ultrahigh definition video image quality objective evaluation method based on visual perception characteristic
CN104079925A
Video frame complexity measurement method based on moving object and image analysis
CN105208402A
Object vehicle video tracking system of complex scene
CN106296731A
Cited By
Image enhancement method and device for X-ray security inspection machine
CN121095080A
An image enhancement method and device for X-ray security inspection machines
CN121095080B