A method for intelligent recognition of robot targets
By using a hybrid Gaussian background and foreground model in the robot reconnaissance system, the enhancement probability of pixel points and the grayscale value frequency is corrected, the problem of poor processing of grayscale features with small frequency is solved, and the accuracy of identification of small targets is improved.
Patent Information
- Application Number
- CN202510377818.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing histogram equalization algorithm can easily lead to excessive compression or complete loss when processing grayscale features with smaller frequency, reducing the accuracy of identification of small targets in robot detection videos.
By collecting reconnaissance videos in real time and using robot motion information for inter-frame pixel matching, a mixed Gaussian background model and foreground model are established, and the enhancement probability of each pixel is calculated based on these models and the frequency of its grayscale value is corrected for histogram equalization.
It effectively improves the contrast and detailed expression of small targets in video, makes the target characteristics more prominent, and improves the accuracy of target detection and recognition.
Smart Images

Figure CN119904626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition. More specifically, the present invention relates to a method for intelligent recognition of robot targets. Background Art
[0002] With the rapid development of artificial intelligence and robot technology, robot target recognition technology has become one of the core technologies in the fields of intelligent manufacturing, autonomous driving, intelligent security, and military reconnaissance.
[0003] In the robot reconnaissance mission, due to the continuous movement of the robot, the clarity of the reconnaissance video captured by the on-board camera decreases, which has a significant impact on the accurate recognition of the target. Therefore, performing image enhancement processing on the reconnaissance video to improve the video clarity is of great significance for improving the target recognition accuracy.
[0004] Currently, the histogram equalization algorithm is widely used in the field of image enhancement due to its simple implementation method and good enhancement effect. For example, the Chinese patent documents with the authorization announcement numbers CN114882346B and CN114092793B both use the histogram equalization algorithm to enhance the images collected by the robot camera.
[0005] However, the histogram equalization algorithm has obvious limitations in the processing process: for the gray-scale features with a higher occurrence frequency, its enhancement effect is significant; but for the gray-scale features with a smaller frequency, they may be over-compressed or even completely lost. In the reconnaissance video, since the target to be recognized is usually small, the corresponding gray-scale features often have a lower occurrence frequency. If the histogram equalization processing is directly performed on each frame of the video, these key target information may be swallowed, resulting in a decrease in the accuracy of target recognition. Summary of the Invention
[0006] To solve the above technical problem that the gray-scale features with a smaller frequency may be over-compressed or even completely lost, thereby reducing the accuracy of target recognition, the present invention provides a method for intelligent recognition of robot targets, including:
[0007] During the process of robot reconnaissance, the reconnaissance video is collected in real time by a camera mounted on the robot; according to the moving speed and direction of the robot at each moment, the pixel points between different frames of the video are matched, and based on the matching result, a Gaussian mixture background model is established to classify the pixel points in the video into foreground pixel points and background pixel points, obtaining the first Gaussian mixture model of the background pixel points; corner matching is performed on the foreground regions composed of foreground pixel points between frames to obtain the corresponding relationship of foreground pixel points between different frames, and based on the corresponding relationship of foreground pixel points, Gaussian mixture model fitting is performed to obtain the second Gaussian mixture model of each foreground pixel point; the enhancement probabilities of background pixel points and foreground pixel points are determined respectively according to the first Gaussian mixture model and the second Gaussian mixture model; for any frame in the video, according to the enhancement probabilities of the pixel points in this frame, the correction frequency of each gray value in this frame is determined, and histogram equalization is performed on this frame according to the correction frequency of each gray value to obtain the enhanced image of this frame; target recognition is performed according to the enhanced images of each frame.
[0008] The present invention can effectively establish a scene background model by collecting reconnaissance video in real time and using the motion information of the robot for inter-frame pixel matching. Based on the Gaussian mixture background modeling method, the video pixel points are accurately divided into two categories: foreground and background, and the first Gaussian mixture model of the background pixel points is established, which not only realizes the accurate extraction of the scene background but also lays a foundation for subsequent target detection. By performing corner matching on the foreground region, the present invention obtains the corresponding relationship of foreground pixel points between different frames and establishes the second Gaussian mixture model of foreground pixel points based on this, which helps to accurately capture the characteristic information of moving targets and improve the accuracy of foreground detection. Based on the first Gaussian mixture model of background pixel points and the second Gaussian mixture model of foreground pixel points, the present invention calculates the enhancement probabilities of background pixel points and foreground pixel points respectively, and corrects the frequency of each gray value based on the enhancement probabilities of each pixel point, increasing the frequency of stable gray features in consecutive frames in each frame, avoiding the over-compression or even complete loss of stable gray features in consecutive frames due to too small a frequency during enhancement, effectively improving the contrast and detail expressiveness of small targets in the video, making the target features more prominent, and performing target recognition based on the enhanced image, improving the accuracy of target detection and recognition, and providing strong technical support for the robot reconnaissance task.
[0009] Preferably, the matching of pixel points between different frames of the video according to the moving speed and direction of the robot at each moment includes: for any two adjacent frames in the video, according to the moving speed, direction of the robot at the corresponding moment and the time interval between the two frames, determining the corresponding pixel points of each pixel point in the previous frame in the next frame.
[0010] Preferably, for the Gaussian mixture background modeling based on the matching result, dividing the pixel points in the video into foreground pixel points and background pixel points to obtain the first Gaussian mixture model of the background pixel points includes: forming a pixel point sequence with the corresponding pixel points in each frame, performing Gaussian mixture background modeling on each pixel point sequence respectively to obtain the Gaussian mixture model of each pixel point sequence, and dividing the pixel points in each pixel point sequence into foreground pixel points and background pixel points; using the Gaussian mixture model of each pixel point sequence as the first Gaussian mixture model corresponding to each background pixel point in the pixel point sequence.
[0011] The present invention performs Gaussian mixture background modeling on the pixel point sequence, uses the Gaussian mixture model to reflect the gray-scale statistical characteristics of the corresponding pixel points in each frame, and not only can distinguish foreground pixel points from background pixel points based on the gray-scale statistical characteristics, but also provides a mathematical basis for obtaining the enhanced probability of background pixel points subsequently.
[0012] Preferably, the method for obtaining the foreground region is: performing connected component analysis on the foreground pixel points in each frame, and using the connected component containing the number of foreground pixel points greater than a preset number threshold as the foreground region.
[0013] The present invention judges the number of foreground pixel points contained in the connected component, eliminates the interference of noise, and obtains a more accurate foreground region.
[0014] Preferably, for the Gaussian mixture model fitting based on the corresponding relationship of the foreground pixel points to obtain the second Gaussian mixture model of each foreground pixel point includes: forming a foreground pixel point sequence with the corresponding foreground pixel points in each frame; performing Gaussian mixture fitting on each foreground pixel point sequence respectively to obtain the Gaussian mixture model of each foreground pixel point sequence, and using the Gaussian mixture model of each foreground pixel point sequence as the second Gaussian mixture model corresponding to the foreground pixel points in the foreground pixel point sequence.
[0015] The present invention performs Gaussian mixture fitting on the foreground pixel point sequence, uses the Gaussian mixture model to reflect the gray-scale statistical characteristics of the foreground pixel points representing the same motion feature in each frame, and provides a mathematical basis for obtaining the enhanced probability of foreground pixel points subsequently.
[0016] Preferably, the enhanced probabilities of the background pixel points and the foreground pixel points satisfy the expression: ; where represents the enhanced probability of the th pixel point in the target frame; represents the set composed of the foreground pixel points in the target frame; represents the th pixel point in the target frame, when the th pixel point is a foreground pixel point, , Indicates the weight of the sub-Gaussian model with the maximum probability density among the sub-Gaussian models of the th pixel in the target frame in its second Gaussian mixture model, as well as the corresponding probability density; is a preset foreground parameter; when the th pixel is a background pixel, , Indicates the weight of the sub-Gaussian model with the maximum probability density among the sub-Gaussian models of the th pixel in the target frame in its first Gaussian mixture model, as well as the corresponding probability density.
[0017] The present invention uses the weights and probability densities of sub-Gaussian models to characterize the gray-scale statistical characteristics of pixel points in consecutive multiple frames, assigns a greater enhancement probability to the stable gray-scale features in consecutive multiple frames, and avoids the stable gray-scale features in consecutive multiple frames being over-compressed or even completely lost due to their small frequency in a single frame, effectively improving the reliability and stability of the enhancement result; the present invention takes into account the different feature requirements of foreground and background regions, can adjust the enhancement degree ratio of foreground and background according to specific application scenarios, can ensure the coordination of the enhancement effect, improve the overall quality and visual perception effect of the image, and provides a higher-quality data basis for subsequent target recognition.
[0018] Preferably, determining the correction frequency of each gray value in the frame according to the enhancement probability of each pixel point in the frame includes: determining the enhancement threshold of each gray value in the frame according to the enhancement probability of each pixel point; in response to the frequency of the gray value being less than the enhancement threshold of the gray value, taking the normalized result of the enhancement threshold of the gray value as the correction frequency of the gray value; in response to the frequency of the gray value being not less than the enhancement threshold of the gray value, taking the normalized result of the frequency of the gray value as the correction frequency of the gray value.
[0019] The present invention corrects the frequency of the gray value according to the enhancement threshold of the gray value, which not only enables the stable gray-scale features in consecutive multiple frames to be enhanced emphatically, avoiding the over-compression or loss of the gray-scale feature due to its small frequency in a single frame, but also ensures that the main gray-scale features in each frame can be enhanced emphatically, making the target features more prominent and providing a basis for subsequent target recognition.
[0020] Preferably, determining the enhancement threshold of each gray value in the frame according to the enhancement probability of each pixel point includes: for any gray value in any frame, taking the mean value of the enhancement probabilities of all pixel points corresponding to the gray value in the frame as the enhancement threshold of the gray value.
[0021] Preferably, the corner matching uses the SIFT feature matching algorithm.
[0022] Preferably, the object recognition based on the enhanced images of each frame includes: inputting the enhanced images of each frame into a trained neural network and outputting the recognition result.
[0023] The beneficial effects of the present invention are as follows: The present invention realizes the accurate extraction of the scene background, laying a foundation for subsequent object detection. The present invention accurately captures the feature information of moving objects, improving the accuracy of foreground detection. The present invention avoids the situation where the frequency of stable gray-scale features in consecutive multiple frames is too small, resulting in over-compression or even complete loss during enhancement, effectively improving the contrast and detail expressiveness of small objects in the video, making the object features more prominent, improving the accuracy of object detection and recognition, and providing strong technical support for the robot reconnaissance task. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart schematically showing a method for intelligent object recognition of a robot in the present invention;
[0025] Figure 2 is a schematic diagram showing a frame image in the reconnaissance video in the present invention;
[0026] Figure 3 is a flowchart schematically showing step S2 of a method for intelligent object recognition of a robot in the present invention;
[0027] Figure 4 is a schematic diagram showing the foreground area in the present invention;
[0028] Figure 5 is a flowchart schematically showing step S3 of a method for intelligent object recognition of a robot in the present invention;
[0029] Figure 6 is a schematic diagram showing, according to the correction frequency of each gray value, Figure 2 the enhanced image obtained after histogram equalization;
[0030] Figure 7 is a schematic diagram showing, according to the frequency before correction of each gray value, Figure 2 the enhanced image obtained after histogram equalization. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0032] The following will describe the specific embodiments of the present invention in detail in conjunction with the accompanying drawings.
[0033] An embodiment of the present invention discloses an intelligent method for robot target recognition. Refer to Figure 1 , which includes steps S1 to S4:
[0034] S1. During the detection process of the robot, the detection video is collected in real time through a camera mounted on the robot.
[0035] In the present invention, the collection frequency is 0.02 frames per second, and the implementer can set the collection frequency according to the actual implementation situation. For the convenience of processing, each frame of the collected video is a grayscale image, Figure 2 which is a frame image in the detection video.
[0036] S2. Match the pixel points between different frames of the video, perform Gaussian mixture background modeling according to the matching results, classify the pixel points in the video into background pixel points and foreground pixel points, and respectively obtain the first Gaussian mixture model of the background pixel points and the second Gaussian mixture model of the foreground pixel points.
[0037] The flowchart of step S2 refers to Figure 3 , which includes steps S201 to S204, specifically:
[0038] S201. Match the pixel points between different frames of the video according to the moving speed and direction of the robot at each moment.
[0039] It should be noted that since the robot is constantly moving, the objects in the video are moving relative to the lens of the camera mounted on the robot. Therefore, in the present invention, the pixel points between adjacent two frames are matched according to the time interval between adjacent two frames, the moving speed and direction of the robot at the corresponding moment.
[0040] Specifically, for any two adjacent frames in the video, according to the time interval between adjacent two frames, the moving speed and direction of the robot at the corresponding moment, a motion model of the camera mounted on the robot during this period is constructed. According to the internal parameters of the camera, the pixel coordinates of the previous frame are converted into three-dimensional world coordinates. According to the motion model of the camera, the three-dimensional world coordinates are re-projected onto the image plane of the latter frame using the internal parameters of the camera, and the corresponding pixel positions in the latter frame can be obtained. In this way, the pixel points corresponding to each pixel point in the previous frame in the latter frame can be obtained.
[0041] It should be noted that constructing the motion model of the camera, converting the pixel coordinates into three-dimensional world coordinates, and projecting the three-dimensional world coordinates onto the image plane are well-known technologies and will not be elaborated in detail here.
[0042] It should be further noted that when the robot moves forward, the distance between the object in front of the robot and the robot shortens. Based on the principle that objects appear larger when closer and smaller when farther away, the object in the subsequent frame in the video becomes larger relative to the object in the previous frame. Thus, one pixel point in the previous frame may correspond to multiple pixel points in the subsequent frame. When the robot moves backward, the distance between the object in front of the robot and the robot increases, and the object in the subsequent frame in the video becomes smaller relative to the object in the previous frame. Multiple pixel points in the previous frame may correspond to one pixel point in the subsequent frame.
[0043] S202. Perform Gaussian mixture background modeling according to the matching results, classify the pixel points in the video into foreground pixel points and background pixel points, and obtain the first Gaussian mixture model of the background pixel points.
[0044] It should be noted that the Gaussian mixture background modeling constructs a Gaussian mixture model for the pixel points at the same position in each frame of the video. Since the robot is moving, the camera carried by the robot is also moving, resulting in different image features being represented by the pixel points at the same position in each frame of the video. Therefore, according to the correspondence relationship of the pixel points in each frame of the video obtained in step S201, the present invention constructs a pixel point sequence and performs Gaussian background modeling on the pixel point sequence.
[0045] Specifically, the corresponding pixel points in each frame form a pixel point sequence. For example, pixel point A1 in the first frame corresponds to pixel points A2 and B2 in the second frame, pixel point A2 in the second frame corresponds to pixel points A3 and C3 in the third frame, and pixel point B2 in the second frame corresponds to pixel points B3 and D3 in the third frame. Then, A1, A2, B2, A3, C3, B3, and D3 are formed into a pixel point sequence.
[0046] Perform Gaussian mixture background modeling on each pixel point sequence respectively to obtain the Gaussian mixture model of each pixel point sequence, and classify the pixel points in each pixel point sequence into foreground pixel points and background pixel points. The Gaussian mixture model of each pixel point sequence is the first Gaussian mixture model corresponding to each background pixel point in the pixel point sequence.
[0047] It should be noted that in the present invention, the number of sub-Gaussian models in the Gaussian mixture background modeling is set to 5, and the implementer can set the number of sub-Gaussian models according to the actual implementation situation. Distinguishing foreground pixel points and background pixel points is a well-known technique in Gaussian mixture modeling technology and will not be elaborated here.
[0048] S203. Perform corner point matching on the foreground regions formed by the foreground pixel points between each frame to obtain the correspondence relationship of the foreground pixel points between different frames.
[0049] It should be noted that the foreground pixel points are the pixel points corresponding to moving objects. In different frames, the positions of the moving objects relative to the camera lens of the robot are different, resulting in different positions and sizes of the regions where the moving objects are located in different frames. Therefore, the present invention matches the foreground pixel points between different frames.
[0050] Specifically, perform connected component analysis on the foreground pixel points in each frame, and regard the connected component containing the number of foreground pixel points greater than the preset number threshold as the foreground region. Regard each pixel point in the connected component containing the number of foreground pixel points not greater than the preset number threshold as a noise point, and reclassify the noise points as background pixel points. Figure 4 It is a schematic diagram of the foreground region.
[0051] In the present invention, the number threshold is set by the implementer according to the actual implementation situation, such as 2.
[0052] Perform corner matching on the foreground regions between two adjacent frames to obtain the matching feature points in the foreground regions of two adjacent frames. Use the matching feature points to calculate the geometric transformation parameters between the foreground regions of two adjacent frames. Based on the geometric transformation parameters, deduce the corresponding relationship of the foreground pixels of non-feature points between two adjacent frames through interpolation operation.
[0053] It should be noted that the present invention uses the Scale Invariant Feature Transform (SIFT) algorithm to perform corner matching on the foreground regions between two adjacent frames, uses the Random Sample Consensus (RANSAC) algorithm to calculate the geometric transformation parameters between the foreground regions of two adjacent frames, and deduces the corresponding relationship of the foreground pixels of non-feature points between two adjacent frames based on bilinear interpolation. The implementer can select the corner matching algorithm, the algorithm for solving the geometric transformation parameters, and the interpolation operation algorithm according to the actual implementation situation. The SIFT algorithm, the RANSAC algorithm, and bilinear interpolation are well-known technologies and will not be elaborated in detail here.
[0054] It should be further noted that in the front and back two frames, the position of the object corresponding to the foreground pixel points changes relative to the position of the robot. When the distance from the object corresponding to the foreground pixel points to the robot becomes larger, based on the principle that objects appear smaller when farther away and larger when closer, the object corresponding to the foreground pixel points becomes smaller in the latter frame, and multiple foreground pixel points in the former frame may correspond to one foreground pixel point in the latter frame; when the distance from the object corresponding to the foreground pixel points to the robot becomes smaller, the object corresponding to the foreground pixel points becomes larger in the latter frame, and one foreground pixel point in the former frame may correspond to multiple foreground pixel points in the latter frame.
[0055] S204. Perform Gaussian mixture model fitting according to the corresponding relationship of the foreground pixel points between each frame to obtain the second Gaussian mixture model of each foreground pixel point.
[0056] The corresponding foreground pixel points in each frame form a foreground pixel point sequence. Gaussian mixture fitting is performed on each foreground pixel point sequence respectively to obtain the Gaussian mixture model of each foreground pixel point sequence, and the Gaussian mixture model of each foreground pixel point sequence is the second Gaussian mixture model corresponding to the foreground pixel points in the foreground pixel point sequence.
[0057] It should be noted that in the present invention, Gaussian mixture fitting is performed on each foreground pixel point sequence respectively, and the least squares method is adopted. Implementers can select a fitting algorithm according to the actual implementation situation, such as the maximum likelihood method. When performing Gaussian mixture fitting on each foreground pixel point sequence in the present invention, the number of sub-Gaussian models included in the Gaussian mixture model is 5, and implementers can adjust the number of sub-Gaussian models according to the actual implementation situation.
[0058] S3. Enhance each frame according to the Gaussian mixture model corresponding to the pixel points in each frame of the video and the pixel point type.
[0059] The flowchart of step S3 refers to Figure 5 , including step S301 to step S303, specifically:
[0060] S301. Take any frame as the target frame, and determine the enhancement probabilities of the background pixel points and the foreground pixel points in the target frame according to the first Gaussian mixture model and the second Gaussian mixture model respectively.
[0061] Specifically, take any frame as the target frame, and the enhancement probability of each pixel point in the target frame satisfies the expression:
[0062] ;
[0063] Among them, represents the enhancement probability of the th pixel point in the target frame; represents the set composed of the foreground pixel points in the target frame; represents the th pixel point in the target frame. When the th pixel point is a foreground pixel point, represents the weight of the sub-Gaussian model with the largest probability density among the sub-Gaussian models of the th pixel point in the target frame in its second Gaussian mixture model, represents the probability density corresponding to the sub-Gaussian model with the largest probability density among the sub-Gaussian models of the th pixel point in the target frame in its second Gaussian mixture model; is a preset foreground parameter used to adjust the importance of the foreground; when the th pixel point is a background pixel point, represents the The weight of the sub-Gaussian model with the maximum probability density among the sub-Gaussian models of the first Gaussian mixture model of a pixel denotes the th pixel in the target frame, and the corresponding probability density of the sub-Gaussian model with the maximum probability density among the sub-Gaussian models of its first Gaussian mixture model
[0064] When the th pixel in the target frame is a foreground pixel, and the probability density of the th pixel in the target frame under a sub-Gaussian model of its second Gaussian mixture model is greater, the th pixel better conforms to the foreground feature corresponding to this sub-Gaussian model. If the weight of this sub-Gaussian model is greater, it indicates that in multiple consecutive frames, the moving object has exhibited this foreground feature, and this foreground feature is a stable gray feature in multiple consecutive frames. This foreground feature is an important feature of the moving object. To make the moving object clearer, it is necessary to focus on enhancing the pixels corresponding to this foreground feature. Therefore, the present invention obtains the enhancement probability of the th pixel according to the probability density and the weight . When the probability density and the weight of the corresponding sub-Gaussian model are greater, the enhancement probability of the th pixel is greater
[0065] When the th pixel in the target frame is a background pixel, and the probability density of the th pixel in the target frame under a sub-Gaussian model of its first Gaussian mixture model is greater, the th pixel better conforms to the background feature corresponding to this sub-Gaussian model. If the weight of this sub-Gaussian model is greater, it indicates that in multiple consecutive frames, the background object has exhibited this background feature, and this background feature is a stable gray feature in multiple consecutive frames. This background feature is an important feature of the object in the background. To make the background clearer, it is necessary to focus on enhancing the pixels corresponding to this background feature. Therefore, the present invention obtains the enhancement probability of the th pixel according to the probability density and the weight . When the probability density and the weight of the corresponding sub-Gaussian model are greater, the enhancement probability of the th pixel is greater
[0066] In the formula, the foreground parameter For adjusting the importance of the foreground, in the present invention, the foreground objects in the motion state are the main detection targets of the robot, and the foreground pixel points of the target frame need to be emphasized. Therefore, in the present invention , in other scenarios, the target to be recognized may be static. Therefore, the implementer can set according to the actual implementation situation .
[0067] It should be noted that as the robot moves, some areas may disappear from the robot's field of view, and some areas may appear in the robot's field of view, resulting in some pixel points in the previous frame having no corresponding pixel points in the next frame, and some pixel points in the next frame having no corresponding pixel points in the previous frame. Therefore, the lengths of the pixel point sequences obtained in step S202 are different. Similarly, as the objects corresponding to the foreground move, some foreground areas may disappear from the robot's field of view, and some foreground areas may appear in the robot's field of view, resulting in some foreground pixel points in the previous frame having no corresponding foreground pixel points in the next frame, and some foreground pixel points in the next frame having no corresponding foreground pixel points in the previous frame. Therefore, the lengths of the foreground pixel point sequences obtained in step S204 are different. When the length of the pixel point sequence is too small, the accuracy of performing Gaussian mixture background modeling on the pixel point sequence is poor. When the length of the foreground pixel point sequence is too small, the accuracy of performing Gaussian mixture fitting on the foreground pixel point sequence is poor. Therefore, in another embodiment, when the length of the pixel point sequence or the foreground pixel point sequence is too small, Gaussian mixture background modeling or Gaussian mixture fitting is not performed on the pixel point sequence or the foreground pixel point sequence, and the enhancement probabilities of the pixel points or foreground pixel points in the pixel point sequence or the foreground pixel point sequence are directly set.
[0068] In another embodiment, in response to the length of the pixel point sequence being not greater than a preset length threshold, Gaussian mixture background modeling is not performed on the pixel point sequence. For each pixel point in the pixel point sequence, it is stipulated that the enhancement probability of each pixel point is . In response to the length of the foreground pixel point sequence being not greater than a preset length threshold, Gaussian mixture fitting is not performed on the foreground pixel point sequence. For each foreground pixel point in the foreground pixel point sequence, it is stipulated that the enhancement probability of each foreground pixel point is . In this embodiment, the length threshold is set by the implementer according to the actual implementation situation. For example, the length threshold is 5. is a preset foreground parameter, and the empirical value is , and the implementer can set according to the actual implementation situation .
[0069] S302. Determine the enhancement threshold of each gray value in the target frame according to the enhancement probability of each pixel point in the target frame.
[0070] It should be noted that when the enhancement probability of each pixel corresponding to a certain gray value in the target frame is greater, it indicates that this gray value is a stable gray feature in multiple consecutive frames, and this gray value needs to be enhanced emphatically. Therefore, the present invention sets the enhancement threshold of each gray value according to the enhancement probability of each pixel corresponding to each gray value, so that when the frequency of each gray value is corrected by using the enhancement threshold of each gray value later, the higher the enhancement threshold of the gray value, the greater the correction frequency of the gray value, thereby achieving the emphatic enhancement of this gray value.
[0071] Specifically, the enhancement threshold of each gray value satisfies the expression:
[0072] ;
[0073] Wherein, represents the enhancement threshold of the gray value in the target frame; represents the enhancement probability of the s-th pixel with the gray value in the target frame; represents the number of pixels with the gray value in the target frame. When the enhancement probability of each pixel with the gray value in the target frame is greater, the enhancement threshold of the gray value in the target frame is higher. After the frequency of the gray value is corrected according to the enhancement threshold later, the greater the correction frequency obtained, so that when the histogram equalization of the target frame is performed according to the correction frequency of each gray value, the image features corresponding to the gray value can be enhanced emphatically.
[0074] S303. Determine the correction frequency of each gray value according to the enhancement threshold of each gray value, and perform histogram equalization on the target frame according to the correction frequency of each gray value to obtain the enhanced image of the target frame.
[0075] In response to the frequency of the gray value being less than the enhancement threshold of the gray value, use the normalized result of the enhancement threshold of the gray value as the correction frequency of the gray value; in response to the frequency of the gray value being not less than the enhancement threshold of the gray value, use the normalized result of the frequency of the gray value as the correction frequency of the gray value.
[0076] Specifically, the correction frequency of each gray value satisfies the expression:
[0077] ;
[0078] Wherein, represents the correction frequency of the gray value in the target frame; represents the enhancement threshold of the gray value in the target frame; represents the gray value in the target frame Frequency; Indicates the grayscale value in the target frame Frequency process quantity. When the grayscale value Frequency is not less than the grayscale value Enhancement threshold, the frequency process quantity takes the value of the grayscale value Frequency. When the grayscale value Frequency is less than the grayscale value Enhancement threshold, the frequency process quantity takes the value of the grayscale value Enhancement threshold; when the grayscale value Frequency is larger, the grayscale value Is the main grayscale feature in the target frame. When the grayscale value Enhancement threshold is larger, the grayscale value Is the stable grayscale feature in multiple consecutive frames; Therefore, in the present invention, when the grayscale value Frequency is less than the enhancement threshold, the grayscale value Corrected frequency is positively correlated with the enhancement threshold. When the grayscale value Frequency is not less than the enhancement threshold, the grayscale value Corrected frequency is positively correlated with the grayscale value Original frequency, thus ensuring that in subsequent enhancement, the main grayscale features in the target frame and the stable grayscale features in multiple consecutive frames can be enhanced emphatically.
[0079] Taking the grayscale value as the horizontal axis and the corrected frequency of the grayscale value as the vertical axis, construct a grayscale correction histogram, and perform histogram equalization on the grayscale correction histogram to achieve the enhancement of the target frame.
[0080] Similarly, obtain the enhanced images of each frame in the video. Figure 6 Is the enhanced image obtained by performing histogram equalization on Figure 2 According to the corrected frequencies of each grayscale value in Figure 2 , Figure 7 Is the enhanced image obtained by performing histogram equalization on Figure 2 According to the frequencies before correction of each grayscale value in Figure 2 . It can be seen that Figure 7 The background in is over-enhanced, the foreground is blurred and difficult to identify, while Figure 6 The foreground in has a good enhancement effect and is easier to identify.
[0081] S4. Perform target recognition according to the enhanced images of each frame.
[0082] Specifically, the present invention uses the YOLOv5 network to recognize the targets in the enhanced image, including: drones, vehicles, etc. Implementers can select a neural network and set the targets to be recognized according to the actual situation.
Claims
1. A robot target intelligent recognition method, characterized in that: include: During the robot reconnaissance process, the camera on the robot collects reconnaissance videos in real time; According to the robot's moving speed and direction at each moment, the pixels between different frames of the video are matched, and a mixed Gaussian background model is performed based on the matching results. The pixels in the video are divided into foreground pixels and background pixels to obtain the first mixed Gaussian model of the background pixels. Corner point matching is performed on the foreground area composed of foreground pixels between frames to obtain the correspondence between foreground pixels between different frames, and a mixed Gaussian model is fitted according to the correspondence between foreground pixels to obtain a second mixed Gaussian model of each foreground pixel; Determine the enhancement probability of background pixels and foreground pixels respectively according to the first mixed Gaussian model and the second mixed Gaussian model; For any frame in the video, the correction frequency of each gray value in the frame is determined according to the enhancement probability of each pixel in the frame, and the frame is histogram equalized according to the correction frequency of each gray value to obtain an enhanced image of the frame; target recognition is performed based on the enhanced image of each frame; The enhancement probability of the background pixel and the foreground pixel satisfies the expression: ; in, Indicates the target frame The enhancement probability of each pixel; Represents the set of foreground pixels in the target frame; Indicates the target frame pixels, when the When a pixel is a foreground pixel, , Indicates the target frame The weight of the sub-Gaussian model with the largest probability density under each sub-Gaussian model in the second mixed Gaussian model for each pixel point and the corresponding probability density; is the preset foreground parameter; when When pixels are background pixels, , Indicates the target frame The weight of the sub-Gaussian model with the largest probability density for each sub-Gaussian model in its first mixed Gaussian model and the corresponding probability density.
2. A robot target intelligent recognition method according to claim 1, characterized in that: The matching of pixels between different frames of the video according to the moving speed and direction of the robot at each moment includes: For any two adjacent frames in the video, the corresponding pixel points in the next frame are determined according to the robot's moving speed and direction at the corresponding moment and the time interval between the two frames.
3. A robot target intelligent recognition method according to claim 1 or 2, characterized in that: The mixed Gaussian background modeling is performed according to the matching result, and the pixels in the video are divided into foreground pixels and background pixels to obtain a first mixed Gaussian model of the background pixels, including: The corresponding pixels in each frame constitute a pixel sequence, and a mixed Gaussian background model is performed on each pixel sequence to obtain a mixed Gaussian model of each pixel sequence, and the pixels in each pixel sequence are divided into foreground pixels and background pixels; the mixed Gaussian model of each pixel sequence is used as the first mixed Gaussian model corresponding to each background pixel in the pixel sequence.
4. A robot target intelligent recognition method according to claim 1, characterized in that: The method for obtaining the foreground area is: A connected domain analysis is performed on the foreground pixels in each frame, and a connected domain containing foreground pixels whose number is greater than a preset threshold is taken as a foreground area.
5. A robot target intelligent recognition method according to claim 1, characterized in that: The method of fitting a mixed Gaussian model according to the corresponding relationship of the foreground pixels to obtain a second mixed Gaussian model of each foreground pixel includes: The corresponding foreground pixel points in each frame constitute a foreground pixel point sequence; mixed Gaussian fitting is performed on each foreground pixel point sequence to obtain a mixed Gaussian model of each foreground pixel point sequence, and the mixed Gaussian model of each foreground pixel point sequence is used as the second mixed Gaussian model corresponding to the foreground pixel point in the foreground pixel point sequence.
6. A robot target intelligent recognition method according to claim 1, characterized in that: Determining the correction frequency of each gray value in the frame according to the enhancement probability of each pixel in the frame includes: Determine the enhancement threshold of each gray value in the frame according to the enhancement probability of each pixel; In response to the grayscale value frequency being less than the grayscale value enhancement threshold, the grayscale value enhancement threshold normalized result is used as the grayscale value correction frequency; in response to the grayscale value frequency being not less than the grayscale value enhancement threshold, the grayscale value frequency normalized result is used as the grayscale value correction frequency.
7. A robot target intelligent recognition method according to claim 6, characterized in that: Determining the enhancement threshold of each gray value in the frame according to the enhancement probability of each pixel point includes: For any gray value in any frame, the average value of the enhancement probabilities of all pixels corresponding to the gray value in the frame is used as the enhancement threshold of the gray value.
8. A robot target intelligent recognition method according to claim 1, characterized in that: The corner point matching adopts the SIFT feature matching algorithm.
9. A robot target intelligent recognition method according to claim 1, characterized in that: The target recognition is performed according to the enhanced image of each frame, including: The enhanced images of each frame are input into the trained neural network and the recognition results are output.
Citation Information
Patent Citations
End-to-end biological target detection method for complex underwater environments
CN114092793B
A vision-based autonomous target recognition method for underwater robots
CN114882346B
Moving workpiece target unsupervised segmentation method suitable for high-dynamic light condition
CN107516320A
Motion foreground detection method and device, terminal equipment and storage medium
CN113409353A
Dioscorea zingiberensis saponin fermentation process monitoring method and system based on image processing
CN118570113A