Object state detection method and device, electronic equipment and storage medium
By using dual Kalman filters and height difference correction methods in intelligent driving, the problem of difficulty in taking into account the accuracy and stability of single-eye distance measurement is solved, and the pedestrian speed measurement and distance measurement effect is achieved with high accuracy, timeliness and stability.
Patent Information
- Application Number
- CN202510043726.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, it is difficult to take into account both the accuracy and stability of the single-eye ranging, which makes it difficult to take into account both the accuracy and detection stability of the ranging.
Using the method of dual Kalman filter, stability is ensured on long-distance objects through the first Kalman filter, and timeliness and accuracy are achieved on close-distance objects through the second Kalman filter. At the same time, by correcting the height difference between the camera position and the plane where the object is located, the accuracy of the measurement results is improved.
In intelligent driving, pedestrian speed measurement and distance measurement are achieved with high accuracy and fast response speed, taking into account detection accuracy, timeliness and stability, and reducing the cost of autonomous driving vehicles.
Smart Images

Figure CN119942595A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual sensing technology, and in particular to an object state detection method, device, electronic equipment and storage medium. Background Art
[0002] In the field of intelligent driving of high-end cars, pedestrian speed and distance measurement plays a very important role in safety warning, active safety and driving safety of intelligent driving. The trajectory of pedestrians can be predicted by their speed and position, so that active avoidance can be carried out.
[0003] In the related technology, the following method is used to measure the speed and distance of pedestrians: pedestrian targets are detected through images, and then the distance between them and the vehicle is measured through the principle of monocular ranging, and the target speed is calculated using the position information. Monocular ranging has a small amount of calculation and low cost, so it is widely used. However, the monocular ranging method lacks depth information, relies on image features and assumptions, and is affected by the stability of image frame detection, making it difficult to balance ranging accuracy and detection stability. Summary of the invention
[0004] One of the purposes of the present invention is to provide an object state detection method to solve the problem in the prior art that it is difficult to balance the accuracy and stability of monocular ranging; the second purpose is to provide an object state detection device; the third purpose is to provide an electronic device.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] A method for detecting an object state, comprising:
[0007] Obtaining a first state detection result of the object based on a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information, wherein a confidence level of the first state prediction value is greater than a confidence level of the state measurement value;
[0008] When it is determined according to the state measurement value that the measured distance of the object is less than or equal to the preset distance, a second state detection result of the object is obtained according to the state measurement value of the object obtained from the image information of the object and the second state prediction value of the image information, wherein the confidence of the state measurement value is greater than the confidence of the second state prediction value;
[0009] A state output result of the object is obtained according to the first state detection result and the second state detection result, and the state output result includes a speed and a target distance.
[0010] According to the above technical means, the state measurement value of the object obtained according to the image information of the object and the first state prediction value of the image information are used to obtain the first state detection result of the object. In this step, the confidence of the first state prediction value is greater than the confidence of the state measurement value; and the state measurement value of the object obtained according to the image information of the object and the second state prediction value of the image information are used to obtain the second state detection result of the object. In this step, the confidence of the state measurement value is greater than the confidence of the second state prediction value. That is, in the first state detection result, the first state prediction value is tended to be trusted, and the first state detection result has hysteresis, but is stable, while in the second state detection result, the prediction value is tended to be trusted, and the second state detection result is obtained in time, and the detection accuracy is high, but the stability is poor. Therefore, combining the advantages of the two detection results, when the close-range object requires timeliness, stability and accuracy, the first state detection result and the second state detection are combined as the state output result of the object, the first state detection result makes up for the stability of the second state detection result, and the second state detection result makes up for the timeliness of the first state detection result, so the object state detection method in this embodiment takes into account the detection accuracy, timeliness and stability.
[0011] Further, obtaining the state output result of the object according to the first state detection result and the second state detection result includes:
[0012] Determine the information entropy of the second state detection result according to the second state detection result and multiple frames of historical second state detection results;
[0013] When the information entropy of the second state detection result is greater than or equal to a preset threshold, obtaining a state output result of the object according to the first state detection result and the second state detection result, wherein a weight of the first state detection result is less than or equal to a preset weight;
[0014] When the information entropy of the second state detection result is less than the preset threshold, the state output result of the object is obtained according to the first state detection result and the second state detection result.
[0015] According to the above technical means, since information entropy can reflect the change of the object state, when the object state changes greatly, the second state detection result is used as the main output of the object state, which can ensure the timeliness of the output of the detection result, and in the second state detection result, the confidence of the state measurement value is greater than the confidence of the second state prediction value, and the accuracy is higher. Based on the second state detection result, an accurate response can be made to the state change in a timely manner. When the object state changes slightly, the first state detection result and the second state detection result are combined as the state output result of the object to ensure the stability of the state output result.
[0016] Further, obtaining the state output result of the object according to the first state detection result and the second state detection result includes:
[0017] Determine the information entropy of the second state detection result according to the second state detection result and multiple frames of historical second state detection results;
[0018] Determine a second weight of the second state detection result according to the information entropy of the second state detection result, and determine a first weight of the first state detection result;
[0019] The state output result of the object is obtained by fusing the first state detection result and the first weight, and the second state detection result and the second weight.
[0020] According to the above technical means, since information entropy can reflect the change of the object state, the weights of the first state detection result and the second state detection result are determined by information entropy, which can improve the first state detection result and the second state detection.
[0021] Further, obtaining the state output result of the object according to the first state detection result and the second state detection result includes:
[0022] Determine a state difference corresponding to the first state detection result and the second state detection result;
[0023] If the state difference meets a preset condition, obtaining a state output result of the object according to the first state detection result and the second state detection result;
[0024] If the state difference does not satisfy the preset condition, the state output result of the object is obtained according to the first state detection result and the second state detection result, wherein the weight of the second state detection result is less than or equal to the preset weight.
[0025] According to the above technical means, due to the poor stability of the second state detection result, there may be abnormal detection results in certain time steps. When the detection result is abnormal, the abnormal situation is eliminated in time.
[0026] Furthermore, the preset conditions include:
[0027] The running direction of the object reflected by the first state detection result is the same as the running direction of the object reflected by the second state detection result;
[0028] A difference between a speed of the object reflected by the first state detection result and a speed of the object reflected by the second state detection result is less than or equal to a preset difference.
[0029] According to the above technical means, by limiting the preset conditions, the situation where the second state detection result is abnormal can be excluded, avoiding outputting abnormal state output results to the vehicle, and further ensuring the accuracy of the state output results.
[0030] Furthermore, the state measurement value of the object obtained according to the image information of the object includes:
[0031] Acquire a detection frame of the object in the image information, and determine position information of the object in the image coordinate system according to the detection frame;
[0032] Determine the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located according to the image information and the camera height;
[0033] Using a monocular ranging algorithm, based on the position information, the camera height and the height difference, the object is measured to obtain an initial distance of the object;
[0034] Obtaining an estimated height of the object according to the initial distance, the imaging height of the object in the camera imaging plane, and the focal length of the camera, and optimizing the estimated height to obtain a target height of the object;
[0035] The measured distance of the object is determined according to the target height, the imaging height of the object in the camera imaging plane and the focal length of the camera, and the speed of the object is determined according to the measured distance obtained from multiple frames of historical image information. The state measurement value includes the speed and measured distance of the object.
[0036] According to the above technical means, during the distance measurement process, the initial distance of the object is corrected based on the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located, thereby improving the accuracy of the distance measurement. In addition, during the distance measurement process, after the target height of multiple historical frames is used to optimize the estimated height of the object, the target height of the object is fixed, and then the measured distance of the object is calculated in combination with the pinhole imaging principle. After the target height of the object is fixed, the influence of height jump on the detection accuracy can be avoided, making the detection result more stable.
[0037] Further, determining the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located according to the image information and the camera height includes:
[0038] Determining a difference value based on the lane line detection width in the image information and the actual width of the lane line;
[0039] The height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located is determined according to the difference value and the camera height.
[0040] According to the above technical means, the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located is obtained through the lane line, and then in the distance measurement process, the height difference is used for correction, which can improve the accuracy of the distance measurement result.
[0041] Further, the optimizing the estimated height to obtain the target height of the object includes:
[0042] Identify object attributes, wherein the object attributes include object type and / or object posture, the object type includes adult or child, and the object posture includes standing or squatting;
[0043] The estimated height is optimized according to the object attributes to obtain a target height of the object.
[0044] According to the above technical means, it is possible to identify the attributes of an object and optimize the estimated height of the object based on the object data. On the one hand, the accuracy of distance measurement can be improved by estimating the height.
[0045] Further, the monocular ranging algorithm is used to measure the distance of the object based on the position information, the camera height and the height difference to obtain the initial distance of the object, including:
[0046] Determine the target grounding point of the object according to the position information, and according to the angle between the target grounding point and the optical center line of the camera;
[0047] An initial distance of the object is obtained according to the camera height, the height difference, the camera pitch angle and the included angle.
[0048] According to the above technical means, based on the monocular ranging algorithm, the initial distance of the object can be quickly obtained with low computing power requirements.
[0049] Further, determining the target grounding point of the object according to the position information includes:
[0050] When the detection frame is a human body detection frame, determining the intersection of the object and the horizon according to the position information as the target grounding point of the object;
[0051] When the detection frame is a head detection frame, the target grounding point is determined by one or more of the following methods:
[0052] estimating a target grounding point of the object according to a ratio of the head detection frame to the body detection frame, a ratio of the head to the body, and the position information;
[0053] The height of the object is estimated according to the head detection frame and the height of adjacent vehicles, and the target grounding point of the object is determined according to the height of the object and the position information.
[0054] According to the above-mentioned technical means, for the scene where the ghost head is blocked during the automatic driving process, the object is identified based on the head detection frame, and the head detection frame and the body detection frame are set to be displayed according to the normal proportions of the human. Therefore, the height of the human can be estimated through the head detection frame and the body detection frame, and then the target grounding point of the object can be determined; and further combined with the height of nearby vehicles, the target height of the object is corrected, thereby improving the accuracy of the target distance of the object. Combined with the setting of the second Kalman filter, the purpose of rapid convergence can be achieved. After inspection, in the scene where the ghost head is blocked, the timeliness and accuracy of the state output results meet the need to obtain the state output results in time for response in the scene where the ghost head is blocked.
[0055] Further, the state measurement value of the object obtained according to the image information of the object and the first state prediction value of the image information are used to obtain the first state detection result of the object, including:
[0056] A first state detection result of the object is obtained by using a first Kalman filter according to a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information;
[0057] Alternatively, obtaining a second state detection result of the object based on a state measurement value of the object obtained from the image information of the object and a second state prediction value of the image information includes:
[0058] A second state detection result of the object is obtained by using a second Kalman filter based on the state measurement value of the object obtained according to the image information of the object and the second state prediction value of the image information.
[0059] An object state detection device, the device comprising:
[0060] A first state detection module is used to obtain a state measurement value of the object and a first state prediction value of the image information according to the image information of the object, and obtain a first state detection result of the object, wherein the first state detection module is configured such that: a confidence level of the prediction value is greater than a confidence level of the measurement value;
[0061] a second state detection module, configured to obtain a state measurement value of the object and a second state prediction value of the image information according to the image information of the object when the measured distance of the object determined based on the state measurement value is less than or equal to a preset distance, and obtain a second state detection result of the object, wherein the second state detection module is configured such that: a confidence level of the measurement value is greater than a confidence level of the prediction value;
[0062] A fusion module is used to obtain a state output result of the object according to the first state detection result and the second state detection result, wherein the state output result includes a speed and a target distance.
[0063] An electronic device, comprising: a memory, a processor;
[0064] The memory is used to store computer programs / instructions; the processor is used to implement the above-mentioned method according to the computer programs / instructions stored in the memory.
[0065] A computer-readable storage medium stores a computer program / instruction, wherein the computer program / instruction is used to implement the method described above when executed by a processor.
[0066] A computer program product, comprising a computer program / instruction, wherein the computer program / instruction is used to implement the method as described above when executed by a processor.
[0067] Beneficial effects of the present invention:
[0068] (1) The present invention sets a dual Kalman filter to optimize the monocular ranging result, wherein the observation noise of the first Kalman filter is large and the process noise is small, and the observation noise of the second Kalman filter is small and the process noise is large. The first Kalman filter ensures the stable output of the detection result of the distant object, and the second Kalman filter can make the detection result of the close-range object output timely and accurate. The first Kalman filter and the second Kalman filter are combined, and according to the scene where the object is located, the advantages of the first Kalman filter and the second Kalman filter are combined, so that the object state output result is highly accurate, timely and stable.
[0069] (2) In the monocular ranging process of the present invention, the height difference between the plane where the camera is located and the plane where the object is located is corrected to improve the measurement result, thereby accelerating the convergence speed of the Kalman filter, so that the Kalman filter can output accurate results in a timely manner.
[0070] (3) The present invention aims at the scene where the ghost head is blocked during the automatic driving process. The object is identified based on the head detection frame, and the head detection frame and the body detection frame are set to be displayed according to the normal proportion of the person. Therefore, the height of the person can be estimated by the head detection frame and the body detection frame, and then the target grounding point of the object can be determined; and the target height of the object can be corrected in combination with the height of nearby vehicles, thereby improving the accuracy of the target distance of the object. Combined with the setting of the second Kalman filter, the purpose of rapid convergence can be achieved. After inspection, in the scene where the ghost head is blocked, the timeliness and accuracy of the state output result meet the need to obtain the state output result in time for response in the scene where the ghost head is blocked. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 A schematic diagram of a flow chart of an object state detection method provided by an embodiment of the present invention;
[0072] Figure 2 for Figure 1 A detailed flow chart of an embodiment of S104;
[0073] Figure 3 A schematic diagram of a flow chart of an object state detection method provided by another embodiment of the present invention;
[0074] Figure 4 A schematic diagram of a flow chart of a method for detecting an object state based on monocular ranging in another embodiment of the present invention;
[0075] Figure 5 A schematic diagram of a height difference between a horizontal plane where a camera is located and a horizontal plane where an object is located based on lane line correction;
[0076] Figure 6 This is a schematic diagram of the monocular ranging principle;
[0077] Figure 7 Schematic diagram of the pinhole imaging principle. DETAILED DESCRIPTION
[0078] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, not for limiting the scope of protection of the present invention.
[0079] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0080] In the field of intelligent driving of high-end cars, pedestrian speed and distance measurement plays a very important role in safety warning, active safety and driving safety of intelligent driving. The trajectory of pedestrians can be predicted by their speed and position, so that active avoidance can be carried out.
[0081] In the related technology, the following method is used to measure the speed and distance of pedestrians: pedestrian targets are detected through images, and then the distance between them and the vehicle is measured through the principle of monocular ranging, and the target speed is calculated using the position information. Monocular ranging has a small amount of calculation and low cost, so it is widely used. However, the monocular ranging method lacks depth information, relies on image features and assumptions, and is affected by the stability of image frame detection, making it difficult to balance ranging accuracy and detection stability.
[0082] Based on this, an embodiment of the present application proposes a multi-filter fusion speed and distance measurement method to achieve high accuracy, fast response speed and low computing power in object speed and distance measurement, thereby improving the perception performance of pedestrians in intelligent driving.
[0083] In this application, an electronic device is used as the execution subject to execute the object state detection method of the following embodiment. Specifically, the execution subject can be a hardware device of the electronic device, or a software application that implements the following embodiment in the electronic device, or a computer-readable storage medium installed with the software application that implements the following embodiment, or a code that implements the software application of the following embodiment. Among them, the electronic device can be a vehicle, or a visual perception module (visual sensor) in a vehicle, or the electronic device can also be other devices.
[0084] Figure 1 A flow chart of an object state detection method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes:
[0085] S101, obtaining a first state detection result of the object according to a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information, wherein a confidence level of the first state prediction value is greater than a confidence level of the state measurement value;
[0086] In the above steps, the image information is acquired by a camera. After acquiring the image information, the object in the image information is acquired through a deep learning detection model. It should be noted that the "object" mentioned in this embodiment and the following embodiments includes but is not limited to pedestrians, for example, it can also be a moving bicycle, an electric vehicle, or other vehicles other than the vehicle. The following is an example based on the "object" for pedestrians.
[0087] After the object is detected, the state measurement value of the object is obtained based on the monocular ranging and speed measurement method.
[0088] In order to further optimize the measurement results of the monocular ranging and speed measurement method and improve the detection accuracy, in some examples, the monocular ranging and speed measurement algorithm is combined with the Kalman filter to detect the state of the object. It should be noted that the processing process and principle of the Kalman filter for data is as follows: if the current moment is t1, before the moment t1 (for example, the moment t-1), the dynamic model is used to predict the state prediction value at the moment t1; based on the image information collected at the moment t1, the state measurement value of the object is obtained by the monocular ranging and speed measurement method, and the state estimation is corrected (corrected) according to the state measurement value and the state prediction value to obtain the state detection result of the object. At the same time, based on the dynamic model and the updated state estimation, the state prediction value at the moment t+1 is continued to be predicted, and after the moment t+1, the above process is repeated (iteration). Through the above prediction and correction steps, the Kalman filter can recursively update the state estimation at each time step, so that the estimated value is as close to the actual state of the system as possible, so that the accuracy of the detection result is higher.
[0089] In the scheme combining the monocular ranging and speed measurement method with the Kalman filter, the monocular ranging and speed measurement method has instability problems, as well as the characteristics of large errors in long-distance ranging and small errors in close-range ranging. If the observation noise of the Kalman filter is set to be large and it is more dependent on the predicted data, the accuracy of the detection result will be affected, and the state estimation will be updated based on the predicted data, and the Kalman filter will be difficult to converge quickly, and the detection result cannot be output in time. If the process noise of the Kalman filter is set to be large, it will be more dependent on the observed data (measured value), and the measurement value accuracy will be higher. The Kalman filter can converge quickly, but it is subject to external interference based on the measured data, resulting in unstable measurement results. Therefore, it is difficult to take into account both real-time and stability in the monocular ranging and speed measurement method combined with a Kalman filter.
[0090] Based on this, in this embodiment, at least two Kalman filters are established according to the characteristics of the visual sensor. The following description is based on two Kalman filters as examples:
[0091] The two Kalman filters are the first Kalman filter and the second Kalman filter, wherein the process noise of the first Kalman filter is small and the observation noise is large, thereby ensuring the stability of the output detection result; the process noise of the second Kalman filter is large and the observation noise is small, thereby ensuring the timeliness (small lag) of the output detection result.
[0092] Based on the observation value setting rules of the first Kalman filter and the second Kalman filter, the first Kalman filter has the advantages of stable and accurate output when detecting distant objects, and the second Kalman filter has the advantages of timely and accurate detection of close objects. Therefore, two Kalman filters are set, and the advantages of the two Kalman filters can be selected according to the specific detection scenario to check the object state.
[0093] For example, through testing, it can be known that since the monocular camera has the characteristics of large error in long-distance ranging and small error in close-range, for long-distance targets, the first Kalman filter is used for detection to ensure the stability of the detection results. Although the output results will have a certain lag, the distance is far, and the lag has little impact on safe driving, and the accuracy can be gradually improved based on multiple predictions and state updates. Therefore, for the state detection of objects at a distance, the detection results of the first Kalman filter can be used as the main one to ensure the stability of the detection results. For close-range objects, the monocular camera has a relatively high short-distance accuracy and the image deep learning model has a high accuracy. The second Kalman filter is used for detection, which has good real-time performance and high sensitivity. Although the stability is not increased, the distance is short and the lag of the detection results has a greater impact on safe driving. Therefore, for the state detection of close-range objects, the detection results of the second Kalman filter can be used as the basis to ensure real-time performance and accuracy.
[0094] After testing, the preset distance is used as the critical value to determine whether the object is far or near.
[0095] As an example, the preset distance is 30m.
[0096] S102, determining whether the measured distance of the object is less than or equal to a preset distance according to the state measurement value;
[0097] The state measurement value includes the measured distance and the measured speed of the object, and determines whether the measured distance is less than or equal to the preset distance.
[0098] In this step, the measured distance of the object refers to the physical distance of the object relative to the vehicle, and the measured distance is calculated by a monocular distance measurement algorithm. By comparing the measured distance with the preset distance, it is determined whether the object is a long-distance object. If it is a long-distance object, the first Kalman filter is used as the main filter to optimize the measured distance, and the second Kalman filter may not be constructed (in some embodiments, it may also be constructed). If it is a close-range object, a second Kalman filter is constructed, and the state detection result is optimized based on the first Kalman filter and the second Kalman filter.
[0099] That is, if the measured distance is less than or equal to the preset distance, S103 is executed to obtain a second state detection result of the object based on the state measurement value of the object obtained from the image information of the object and the second state prediction value of the image information, wherein the confidence of the state measurement value is greater than the confidence of the second state prediction value;
[0100] A second Kalman filter is constructed, and for the state measurement value and the second state prediction value, the second Kalman filter is used to obtain the second state detection result of the object. Based on the fact that the observation noise of the second Kalman filter is small and the process noise is large, the confidence of the state measurement value predicted by the second Kalman filter is greater than the confidence of the second state prediction value.
[0101] For the state measurement value and the first state prediction value, the first Kalman filter is used to obtain the first state detection result of the object. Based on the fact that the observation noise of the first Kalman filter is large and the process noise is small, therefore, for the first Kalman filter, the confidence of the first state prediction value predicted is greater than the confidence of the state measurement value.
[0102] The following describes the parameter settings of the first Kalman filter and the second Kalman filter through specific examples:
[0103] In most cases, pedestrians (objects) can be approximated as uniform linear motion, so the linear Kalman filter is used in the pedestrian speed measurement process. In practical applications, the effect of the linear Kalman filter is better than that of the nonlinear filter. Normal pedestrians are approximately uniform straight-line motion, so the CV kinematic model is used to establish the linear Kalman filter. Taking the first Kalman filter as an example, it is as follows:
[0104] The state quantity is:
[0105] Where x represents the horizontal position of the object; y represents the vertical position of the object; v x Indicates the lateral velocity of the object; v y Indicates the longitudinal velocity of the object.
[0106] The state transfer equation F is:
[0107] Measurement formula:
[0108] Prediction equation: X t =F*X t-1 +Q;
[0109] Correct the equation:
[0110] Where R represents observation noise, Q represents process noise, is a priori observable quantity.
[0111] According to the characteristics of monocular ranging of visual sensors, the target is far away and the observation noise is large, and the observation noise has a certain relationship with the distance. After system testing, the observation noise error is about 5% within 60m, so the observation noise is set as follows:
[0112] lg noise =scale*(0.05*pos) 2 ;
[0113] Among them, pos represents the target distance; scale represents the dynamic adjustment of noise according to the jump of the observed value (that is, the measured value); lg_noise represents the position noise.
[0114] The scale value in the above formula is dynamically adjusted based on the variance of the measured distance and the measured speed of historical observations. The larger the variance, the smaller the scale value. Since the main linear Kalman filter establishes a uniform straight-line model, the actual acceleration of the object will affect the accuracy of its prediction. Therefore, the relationship between process noise and acceleration is set as follows:
[0115]
[0116] V noise =(acc*Δt) 2 ;
[0117] Among them, P noise Represents distance noise; V noise represents velocity noise, acc represents the acceleration of the object, which is set to a fixed value according to the system situation, and △t is the time interval; optionally, acc can be set to 1.
[0118] Since pedestrians do not walk at an absolutely uniform speed during their walking process, by setting the relationship between process noise and acceleration as described above, the distance and acceleration can be dynamically measured to adjust the process noise corresponding to the Kalman filter, making the detection result more accurate.
[0119] The above is an example of parameter setting of the first Kalman filter.
[0120] Optionally, when the embodiment of the present application is applied to a vehicle, it is used to perceive the surrounding information of the vehicle during the intelligent driving process of the vehicle. In the scene of intelligent driving of the vehicle, there may be scenes such as pedestrians crossing, pedestrians peeking out, crossing and turning straight, etc. Based on the above scenes, the first Kalman filter will have slow speed convergence, making the final output speed have obvious lag characteristics. In addition, the above-mentioned first Kalman filter is more biased towards the predicted value, and the front observation noise is set relatively large. The overall position speed output is relatively stable, and the sudden change of the object shows a large lag.
[0121] Therefore, the parameter setting example of the second Kalman filter is similar to the parameter setting example of the first Kalman filter, except that the second Kalman filter is more biased towards the measured value. The following is an explanation of the setting of the second Kalman filter: a 30m short-distance linear Kalman filter is established based on the characteristic of small target close-range error of the monocular vision sensor. The entire noise setting method is similar to the above-mentioned first Kalman filter, except that the scale is reduced when observing noise. After a large number of tests, the second Kalman filter shows good real-time performance and high sensitivity. In the speed of the scene with sudden speed changes such as pedestrian crossing, crossing termination and two-wheeled vehicles crossing and turning straight, the speed convergence is greatly improved, meeting the system requirements. At the same time, due to its high sensitivity, the second Kalman filter is greatly affected by the stability of the model grounding point detection effect. If the filter is used directly in a short distance, the final result will also fluctuate greatly. Therefore, after the dual filter is established, the value of the corresponding filter is reasonably selected to highlight the advantages of each filter and make up for each other's shortcomings, so that the output result is better. The following is a statement of the specific way to reasonably select the value of the corresponding filter.
[0122] S104, obtaining a state output result of the object according to the first state detection result and the second state detection result, where the state output result includes a speed and a target distance.
[0123] By combining the detections output by the two Kalman filters, stability, timeliness, and accuracy can be achieved. The following describes the specific fusion methods of the first state detection result and the second state detection result in different scenarios:
[0124] As an example, the weight of the first state detection result and the weight of the second state detection result can be set, and the fusion is performed according to the weight distribution. The weight can be preset, for example, the weight of the first state detection result is 0.2, and the weight of the second state detection result is 0.8, then the state output result = first state detection result * 0.2 + second state detection result * 0.8.
[0125] As an example, the first state detection result and the second state detection result may also be fused in the following manner:
[0126] According to the second state detection result and the second state detection results of multiple frames of history, the information entropy of the second state detection result is determined; according to the information entropy of the second state detection result, the second weight of the second state detection result is determined, and the first weight of the first state detection result is determined; according to the first state detection result and the first weight, the second state detection result and the second weight, the state output result of the object is obtained by fusing.
[0127] Wherein, based on the second state detection result and the multi-frame historical second state detection results, the information entropy of the second state detection result is obtained, and the change of the object can be judged by the information entropy, thereby determining the credibility of the second state detection result. The second weight of the second state detection result is determined based on the information entropy, and the sum of the first weight and the second weight is 1.
[0128] For example, if the information entropy is relatively large, the second weight is relatively large, the first weight is relatively small, and the first difference between the second weight and the first weight is large. If the information entropy is small, the second weight is large, the first weight is small, but the second difference between the second weight and the first weight is smaller than the first difference.
[0129] In some embodiments, the second weight and the first weight are adjusted by information entropy. In each time step, when the determined information entropy is different, the corresponding first weight and the second weight are different, thereby dynamically adjusting the weight distribution of the first state detection result and the second state detection result, and adjusting the fusion result in time according to the changes to improve the accuracy of the detection result.
[0130] If the measured distance is greater than the preset distance, S105 is executed to obtain a state output result of the object based on the first state detection result.
[0131] In this step, if the measurement distance is large, the first state detection result is used as the state output result of the object, and there is no need to establish a second Kalman filter. Since the object is far away, the output result of the first Kalman filter can also meet the detection accuracy and timeliness requirements.
[0132] In this embodiment, the state measurement value of the object obtained according to the image information of the object and the first state prediction value of the image information are used to obtain the first state detection result of the object. In this step, the confidence of the first state prediction value is greater than the confidence of the state measurement value; and the state measurement value of the object obtained according to the image information of the object and the second state prediction value of the image information are used to obtain the second state detection result of the object. In this step, the confidence of the state measurement value is greater than the confidence of the second state prediction value. That is, in the first state detection result, the first state prediction value is tended to be trusted, and the first state detection result has hysteresis but is stable, while in the second state detection result, the prediction value is tended to be trusted, and the second state detection result is obtained in time, with high detection accuracy, but poor stability. Therefore, by combining the advantages of the two detection results, when timeliness, stability and accuracy are required for close-range objects, the first state detection result and the second state detection are combined as the state output result of the object. The first state detection result makes up for the stability of the second state detection result, and the second state detection result makes up for the timeliness of the first state detection result. Therefore, the object state detection method in this embodiment takes into account detection accuracy, timeliness and stability. Its application in monocular ranging has low requirements on computing power. At the same time, it takes into account detection accuracy and stability, which can reduce the cost of autonomous driving vehicles.
[0133] Figure 2 Another embodiment of a state detection method provided based on the above embodiment is as follows: Figure 2 As shown, based on the above embodiment, S104 may specifically be:
[0134] S201, determining the information entropy of the second state detection result according to the second state detection result and multiple frames of historical second state detection results;
[0135] As an example, the information entropy of the second state detection result is calculated in the following way:
[0136]
[0137]
[0138] Among them, E-information entropy of state (speed), P-weight ratio of object.
[0139] It should be noted that the above information entropy reflects the dynamics of the object. For example, when the information entropy of the second state detection result is large, it reflects that the object has stopped crossing or turned from crossing to going straight, resulting in a large change in the state of the object, so the information entropy is large. When the object crosses or goes straight, the state of the object changes little, so the information entropy is small.
[0140] After testing, a preset threshold is set as a critical value to identify the change scene of the object. For example, when the information entropy is greater than or equal to the preset threshold, the change scene of the object is likely to be: traversal termination or traversal to straight. If the information entropy is less than the preset threshold, the change scene of the object is likely to be traversal or straight.
[0141] S202, determining whether the information entropy of the second state detection result is greater than or equal to a preset threshold;
[0142] When the information entropy of the second state detection result is greater than or equal to the preset threshold, executing S203, obtaining a state output result of the object according to the first state detection result and the second state detection result, wherein the weight of the first state detection result is less than or equal to the preset weight;
[0143] When the object stops crossing or turns to straight, the vehicle is likely to collide with the object and needs to respond in time. Therefore, in this scenario, the second state detection result is used as the basis and feedback is given to the vehicle in time so that the vehicle can take timely collision prevention measures.
[0144] In this scenario, the preset weight can be set to zero, that is, the first state detection result is zero, and the first state detection result is removed based on the second state detection result as the state output result of the object. Optionally, the preset weight can also be greater than 0, but the preset weight is much smaller than the weight of the second state detection result, that is, the weight of the first state detection result is extremely small, and the second state detection result is mainly used.
[0145] When the information entropy of the second state detection result is less than a preset threshold, S204 is executed to obtain a state output result of the object according to the first state detection result and the second state detection result.
[0146] For scenes like pedestrians crossing, the information entropy value is relatively small, and the state output result of the object is obtained based on the first state detection result and the second state detection. In this step, the specific fusion method is the same as that of the first embodiment above, and the details can be referred to the first embodiment above, which will not be repeated here.
[0147] In this embodiment, the scene in which the object is located is identified by the information entropy of the second state detection result, and it is determined whether the state output result is based on the first state detection result, the second state detection result, or a combination of the two according to the scene in which the object is located. The corresponding state output results are determined differently in different scenarios. For scenarios with high timeliness requirements, the state output results are output in a timely manner, which can provide a timely response for the vehicle.
[0148] Figure 3 A state detection method according to another embodiment provided based on all the above embodiments is as follows: Figure 3As shown, this embodiment is based on all the above embodiments. In this embodiment, the object state detection method includes:
[0149] S301, obtaining a first state detection result of the object according to a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information, wherein a confidence level of the first state prediction value is greater than a confidence level of the state measurement value;
[0150] S302, determining whether the measured distance of the object is less than or equal to a preset distance according to the state measurement value;
[0151] If the measured distance is less than or equal to the preset distance, executing S303, obtaining a second state detection result of the object according to the state measurement value of the object obtained from the image information of the object and the second state prediction value of the image information, wherein the confidence of the state measurement value is greater than the confidence of the second state prediction value;
[0152] The specific implementation of steps S301 to S303 is the same as above. Figure 1 Steps S101 to S103 in the embodiment are similar, and reference may be made to the above embodiment for details.
[0153] S304, determining a state difference corresponding to the first state detection result and the second state detection result;
[0154] As in the above embodiment, if the measured distance between the object and the vehicle is less than or equal to the preset distance, under normal circumstances, the weight of the second state detection result is large and the weight of the first state detection result is small. In the final state output result of the object, the second state detection result is dominant.
[0155] However, the second Kalman filter has an instability problem, so the second state detection result may have inaccuracies such as sudden changes. Therefore, when this embodiment determines that the measured distance of the object is less than the preset distance, it compares the state difference between the first state detection result and the second state detection result, and judges whether the second state detection result is accurate based on the state difference.
[0156] S305, determining whether the state difference meets a preset condition;
[0157] As an example, the preset conditions include:
[0158] The running direction of the object reflected by the first state detection result is the same as the running direction of the object reflected by the second state detection result. The detection results include speed and measured distance, wherein the changes in speed and measured distance can reflect the moving direction of the object. If the first state detection result reflects that the running direction of the object is moving toward the vehicle, and the second state detection result reflects that the running direction of the object is moving away from the vehicle, then it is determined that the directions of movement are opposite. Since the first state detection result is more biased towards the prediction result, and the prediction result is based on the prediction of the dynamic model, the accuracy of the first state detection result is higher. Therefore, when the directions of movement are opposite, the probability of the second state detection result being wrong is greater. Therefore, in this case, it is determined that the state difference does not meet the preset conditions.
[0159] The difference between the speed of the object reflected by the first state detection result and the speed of the object reflected by the second state detection result is less than or equal to the preset difference. Similarly, if the speed difference between the object reflected by the first state detection result and the second state detection result is very large, it means that the probability of the second state detection result being wrong is high, and it is determined that the state difference does not meet the preset condition.
[0160] As an example, the preset difference may be 0.5 m / s.
[0161] On the contrary, if the above conditions are met, it is determined that the state difference meets the preset conditions.
[0162] If the state difference meets the preset condition, executing S306, obtaining the state output result of the object according to the first state detection result and the second state detection result;
[0163] When the state difference meets the preset conditions, it means that the second state detection result is normal. For the detection of close-range objects, the state output result of the object is obtained by combining the first state detection result and the second state detection result. For the specific combination method, please refer to the above Figure 1 The embodiment shown.
[0164] If the state difference does not meet the preset condition, S307 is executed to obtain the state output result of the object according to the first state detection result and the second state detection result, wherein the weight of the second state detection result is less than or equal to the preset weight.
[0165] When the state difference does not meet the preset conditions, it means that the second state detection result is an abnormal result. In this scenario, the preset weight can be zero, that is, the second state detection result is zero, then the state output result based on the first state detection result as the object is removed. Optionally, the preset weight can also be greater than 0, but the preset weight is much smaller than the weight of the first state detection result, that is, the weight of the second state detection result is extremely small, and the first state detection result is mainly used.
[0166] If the measured distance is greater than the preset distance, S308 is executed to obtain a state output result of the object based on the first state detection result.
[0167] When detecting close objects, the second state detection result is unstable, so the second state detection result may be abnormal. This embodiment sets preset conditions to determine whether the second state detection result is abnormal. If abnormal, the second state detection result is discarded. If normal, the state output result is obtained based on the second state detection result and the first state detection result, thereby improving the accuracy of the state output result.
[0168] Figure 4 A state detection method according to another embodiment provided based on all the above embodiments is as follows: Figure 4 As shown, in all the above embodiments, the specific implementation method of obtaining the state measurement value of the object according to the image information can be:
[0169] S401, obtaining a detection frame of an object in image information, and determining position information of the object in an image coordinate system according to the detection frame;
[0170] In the embodiment of the present application, the image information is collected by a camera, such as a vehicle-mounted camera. The image is recognized by a pre-trained deep neural network model to detect the object, and the position information of the object is reflected by the position of the detection frame in the image.
[0171] Optionally, before application, multiple groups of image information are collected by the camera, manually annotated, and a deep neural network model (such as Yolov5) capable of detecting objects is trained to complete multi-target detection. The output parameters of the deep neural network model include the object detection frame, the category of the object detection frame, and the location information (location coordinates) of the object.
[0172] Furthermore, in order to achieve object tracking, when the detection frame of the object is obtained, the IOU (Intersection over Union, an indicator used to measure the degree of overlap between two bounding boxes) between the detection frames is calculated, and then Hungarian matching is used to associate the objects to achieve object tracking.
[0173] S402, determining the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located according to the image information and the camera height;
[0174] It should be noted that there may or may not be a height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located. Taking a vehicle driving on a horizontal road as an example, the position of the object (pedestrian) and the position of the vehicle are on the same horizontal plane, and there is no height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located, that is, the height difference is 0; when the vehicle is driving on a slope, or the current position of the vehicle is a pit, the position of the object and the position of the vehicle are not on the same horizontal plane, so there is a height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located.
[0175] In intelligent driving scenarios, the vehicle or the object is usually not on the same horizontal plane and there is a height difference. Therefore, in order to improve the accuracy of object state detection, the height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located is considered, and the image captured by the camera is corrected according to the height difference, so that the output result of the camera is more accurate, thereby making the subsequent state detection results more accurate.
[0176] As an example, Figure 5 As shown in the figure, the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located is estimated by using the lane lines in the image information and the actual width of the lane. The specific implementation process is as follows:
[0177] Determine a difference value based on the lane line detection width in the image information and the actual width of the lane line;
[0178] According to the difference value and the camera height, the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located is determined.
[0179] As an example, the actual width of the lane line can be determined based on the standard width of the lane line, or the actual width of the current lane line can be determined based on map information.
[0180] As an example, the difference value may be reflected by a ratio of the lane line detection width to the actual width, or by a scaling ratio of the lane line detection width to the actual width.
[0181] The difference between the lane line detection width and the actual width can reflect the change in the camera's pitch angle. Therefore, the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located can be determined through the difference value and the camera height.
[0182] S403, using a monocular ranging algorithm to measure the distance of the object based on the position information and the height difference to obtain an initial distance of the object;
[0183] This step uses a monocular ranging algorithm, the position information of the object and the height difference of the camera to measure the distance. The position information of the object refers to the position information of the object in the image coordinate system, and the object ranging refers to the physical distance of the object (relative to the actual distance of the vehicle).
[0184] like Figure 6 As shown, ranging can be performed in the following ways:
[0185] First, the target ground point of the object is determined based on the position information, and the angle between the target ground point and the optical center line of the camera is used. The target ground point refers to the intersection of the object and the ground. The initial relative distance of the object refers to the longitudinal distance between the target ground point and the vehicle (camera).
[0186] If there is a height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located, there will be an angle between the camera light and the target ground point ray. The high angle can be determined based on the height difference and the focal length of the camera, for example
[0187] The initial distance of the object is obtained based on the camera height, height difference, camera pitch angle and angle.
[0188] Based on the monocular ranging algorithm, the calculation formula for the initial distance of the object is:
[0189] x = Cam_h*tan(θ1-θ2);
[0190] Among them, Cam_h is the height of the camera from the ground plane, θ1 is the pitch angle of the camera (the pitch angle is obtained through calibration); θ2 is the angle between the camera light and the target ground point ray.
[0191] If the height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located is 0, then θ2 is zero.
[0192] If there is a height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located, the initial distance of the object is calculated by the following formula:
[0193] x=(Cam h -h Δ )*tan(θ1-θ2);
[0194] In this example, the initial distance is corrected based on the height difference between the horizontal plane where the camera is located and the horizontal plane where the object is located. This can improve the accuracy of the detection result and avoid the problem of the camera pitch in the scene where the object and the car are not on the same horizontal plane, which affects the accuracy of the distance measurement.
[0195] S404, obtaining an estimated height of the object according to the initial distance, the imaging height of the object in the camera imaging plane, and the focal length of the camera, and optimizing the estimated height to obtain a target height of the object;
[0196] In this step, after the initial distance of the object is measured by the monocular ranging method, the initial distance is optimized using the pinhole imaging principle. This is specifically achieved by the following methods:
[0197] like Figure 7 As shown, based on the pinhole imaging principle, the physical height of the object (that is, the estimated height of the object) is estimated. For example, the estimated height of the object is:
[0198]
[0199] Where X is the initial distance, H is the imaging height of the object in the camera imaging plane, and f is the focal length of the camera.
[0200] It should be noted that the estimated height of the object is estimated based on the initial distance and the pinhole imaging principle, and its accuracy is not high. In addition, the estimated height of the same object may be different in different time steps, so it may fluctuate. To avoid this, the estimated height is optimized in this embodiment.
[0201] As an example, the process of optimizing the estimated height: Based on object tracking, the relevant information of the same object at different time steps can be determined, and the estimated height can be optimized based on the historical target height of the object. For example, if the historical target height of the object is 1.75m, and the estimated height of the object at the current time step is 1.71m, then 1.75 is used as the estimated height of the object. Alternatively, the estimated height of the object can be optimized by the average height of the historical target height.
[0202] As an example, the process of optimizing the estimated height is as follows: the upper and lower limits of the height are set based on the object attributes. For example, the object data includes the object type, and the object type includes adults or children. If the object is an adult, the upper limit of the height is set to 1.8m~1.9m, and the lower limit of the height is set to 1.1m~1.5m; if the object is a child, the upper limit of the height is set to 1.3m~1.5m. When the object is identified as an adult by the deep learning model, if the estimated height of the object is 1.1m or less than 1.5m, 1.1m or 1.5m is used as the target height to optimize the estimated height. If the estimated height of the object is above 1.8m or 1.9m, 1.8m or 1.9m is used as the target height. Correspondingly, if the object is identified as a child, the estimated height is optimized with the height limit of the child. For example, if the estimated height is greater than 1.3m or 1.5m, the height of 1.3m is directly used as the target height.
[0203] Alternatively, in some examples, the object attribute also includes the object posture, and the object posture includes standing or squatting, wherein standing is a normal pedestrian state, and squatting is an object in a wheelchair, etc. Based on the object posture, a corresponding height optimization method can also be set, which is similar to the height optimization method of the above object type, and can refer to the above example.
[0204] Based on this, in some embodiments, the optimization of the estimated height may be implemented in the following manner:
[0205] Identify object attributes;
[0206] The estimated height is optimized based on the object properties to obtain the target height of the object.
[0207] By optimizing the estimated height, the stability of the height can be further improved. The height is initialized and dynamically adjusted according to the object attributes (adult or child, squatting attributes), and the final target height is obtained by establishing an EKF height filter.
[0208] S405 , outputting a measured distance of the object according to the target height, the imaging height of the object in the camera imaging plane, and the focal length of the camera, wherein the state measurement value includes the measured distance of the object.
[0209] Based on the optimized target height of the object, the measurement distance of the object is calculated using the pinhole imaging principle. The specific estimation is done in the following way:
[0210]
[0211] This method is used to measure the distance of pedestrians (objects), which is less affected by changes in the grounding point. At the same time, for the same object, the target height of the object can be fixed in the above manner. In this way, the distance jump is related to the detection frame and has nothing to do with the target height, thereby improving the stability of distance detection.
[0212] Furthermore, in some ghost peeking scenes during autonomous driving, the ghost peeking object is partially blocked, and the blocked part is often in contact with the ground. In this scenario, since the height of the object cannot be detected, the error of the monocular camera in ranging and speed measurement is large, resulting in a slow convergence speed of the Kalman filter. In this scenario, it is necessary to quickly output information such as the position and speed of the object so that the vehicle can respond in time and avoid collision.
[0213] Based on this, the embodiment of the present application combines a dual filter to solve the problem of fast convergence of the speed of unobstructed objects. For obstructed objects, the main reason is the height of the object and the inaccurate distance measurement accuracy. Therefore, the embodiment of the present application sets the head detection of the object in the image information and the whole body detection of the object in the image information in the deep learning detection model, classifies the head detection and the whole body inspection, and outputs the body detection frame and the head detection frame respectively. When an object is identified and the detection frame of the object is a head detection frame, the object is determined to be an obstructed object. When the detection frame of an object is identified as a body detection frame, the object is determined to be an unobstructed object.
[0214] After using the detection frame to distinguish whether the object is blocked, the height and distance of the object are detected in the following ways:
[0215] For example, in the above steps, determining the target grounding point of the object according to the position information includes:
[0216] When the detection frame is a human body detection frame, the intersection of the object and the horizon is determined according to the position information, which is the target grounding point of the object;
[0217] When the detection frame is a head detection frame, the target ground point is determined by one or more of the following methods:
[0218] Estimate the target grounding point of the object based on the ratio of the head detection frame to the body detection frame, the ratio of the head to the body, and the position information;
[0219] The height of the object is estimated based on the head detection frame and the height of adjacent vehicles, and the target grounding point of the object is determined based on the height and position information of the object.
[0220] For the obscured ghost peeking scene, the target ground point is estimated by using the human body proportion through the head detection frame, and the target distance is measured based on the ground point. Alternatively, the distance between pedestrians and the vehicles next to them is relatively close, and the height of the ghost peeking object is obtained based on the proportion of its detection frame and the fixed height of the vehicle.
[0221] Alternatively, in some embodiments, when the detection frame is a head detection frame, the target grounding point of the object is first estimated based on the ratio of the head detection frame to the body detection frame, the ratio of the head to the body, and the position information;
[0222] Through the target grounding point, the initial distance of the object is obtained, and the estimated height of the object is obtained according to the initial distance, the imaging height of the object in the camera imaging plane and the focal length of the camera. Then, according to the ratio of the head detection frame to the height of the adjacent vehicle, the estimated height is corrected to obtain the target distance of the object. This improves the ranging accuracy of the object after the lower half is blocked, and improves the reliability of the position information of the blocked object. By using the target distance estimation speed of the object, fast convergence can be achieved in this scenario.
[0223] As an example, when two target distances are calculated by the above method, the target distance whose longitudinal distance is close to the self-vehicle is selected according to the longitudinal distance between the object and the vehicle compared to the distance of the vehicle. After accurately obtaining the position information (i.e., the target distance), the speed of the object is fitted using its 3-5 frames of historical information, and then the second Kalman filter is used for optimization to obtain the accurate speed, so that the position and speed of the object can be obtained in time in this scenario, so that the vehicle can respond to the avoidance control in time.
[0224] In general, the object detection solution provided by this application can be applied to object detection in intelligent driving scenarios, and can also be used for object detection in other scenarios. In the application of intelligent driving scenarios, since object detection is timely, accurate and stable, the vehicle can make a protective response in time according to the object detection results, thereby improving the safety performance of intelligent driving.
[0225] The present application also provides an object state detection device, the device comprising:
[0226] A first state detection module is used to obtain a state measurement value of the object and a first state prediction value of the image information according to the image information of the object, and obtain a first state detection result of the object, wherein the first state detection module is configured such that: a confidence level of the prediction value is greater than a confidence level of the measurement value;
[0227] A second state detection module is used to obtain the state measurement value of the object and the second state prediction value of the image information according to the image information of the object when the measured distance of the object determined based on the state measurement value is less than or equal to the preset distance, and obtain a second state detection result of the object, wherein the second state detection module is configured such that: the confidence of the measured value is greater than the confidence of the prediction value;
[0228] The fusion module is used to obtain a state output result of the object according to the first state detection result and the second state detection result, and the state output result includes a speed and a target distance.
[0229] The object state detection device provided in the embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be described in detail here.
[0230] The embodiment of the present application also provides an electronic device, the electronic device comprising: a memory, a processor;
[0231] The memory is used to store computer programs / instructions; the processor is used to execute the computer programs / instructions stored in the memory to implement the methods involved in the above embodiments.
[0232] The electronic device also includes a communication interface and a CAN bus, wherein the processor is used to provide computing power and control capabilities, and can be a GPU, CPU, NPU, MCU, FPGA, etc. The storage device includes an internal memory and a non-volatile memory. The non-volatile memory stores a computer program that implements the above method. The internal memory provides an environment for program startup and operation. The communication interface is used to communicate with an external terminal by wire or wireless.
[0233] In this embodiment, the electronic device may be an electronic component in a car, or may be a car.
[0234] The present invention also provides a computer-readable storage medium / computer program product, in which computer control instructions are stored / the computer program product includes computer control instructions, and when the computer control instructions are executed by a processor, they are used to implement the methods involved in the above-mentioned embodiments.
[0235] The above embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present invention is within the protection scope of the present invention.
Claims
1. A method for detecting an object state, characterized in that: include: Obtaining a first state detection result of the object based on a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information, wherein a confidence level of the first state prediction value is greater than a confidence level of the state measurement value; When it is determined according to the state measurement value that the measured distance of the object is less than or equal to the preset distance, a second state detection result of the object is obtained according to the state measurement value of the object obtained from the image information of the object and the second state prediction value of the image information, wherein the confidence of the state measurement value is greater than the confidence of the second state prediction value; A state output result of the object is obtained according to the first state detection result and the second state detection result, and the state output result includes a speed and a target distance.
2. The method according to claim 1, characterized in that The obtaining, according to the first state detection result and the second state detection result, a state output result of the object includes: Determine the information entropy of the second state detection result according to the second state detection result and multiple frames of historical second state detection results; When the information entropy of the second state detection result is greater than or equal to a preset threshold, obtaining a state output result of the object according to the first state detection result and the second state detection result, wherein a weight of the first state detection result is less than or equal to a preset weight; When the information entropy of the second state detection result is less than the preset threshold, the state output result of the object is obtained according to the first state detection result and the second state detection result.
3. The method according to claim 1, characterized in that The obtaining, according to the first state detection result and the second state detection result, a state output result of the object includes: Determine the information entropy of the second state detection result according to the second state detection result and multiple frames of historical second state detection results; Determine a second weight of the second state detection result according to the information entropy of the second state detection result, and determine a first weight of the first state detection result; The state output result of the object is obtained by fusing the first state detection result and the first weight, and the second state detection result and the second weight.
4. The method according to any one of claims 1 to 3, characterized in that: The obtaining the state output result of the object according to the first state detection result and the second state detection result includes: Determine a state difference corresponding to the first state detection result and the second state detection result; If the state difference meets a preset condition, obtaining a state output result of the object according to the first state detection result and the second state detection result; If the state difference does not satisfy the preset condition, the state output result of the object is obtained according to the first state detection result and the second state detection result, wherein the weight of the second state detection result is less than or equal to the preset weight.
5. The method according to claim 4, characterized in that The preset conditions include: The running direction of the object reflected by the first state detection result is the same as the running direction of the object reflected by the second state detection result; A difference between a speed of the object reflected by the first state detection result and a speed of the object reflected by the second state detection result is less than or equal to a preset difference.
6. The method according to any one of claims 1 to 3, characterized in that: The state measurement value of the object obtained according to the image information of the object includes: Acquire a detection frame of the object in the image information, and determine position information of the object in the image coordinate system according to the detection frame; Determine the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located according to the image information and the camera height; Using a monocular ranging algorithm, based on the position information, the camera height and the height difference, the object is measured to obtain an initial distance of the object; Obtaining an estimated height of the object according to the initial distance, the imaging height of the object in the camera imaging plane, and the focal length of the camera, and optimizing the estimated height to obtain a target height of the object; The measured distance of the object is determined according to the target height, the imaging height of the object in the camera imaging plane and the focal length of the camera, and the speed of the object is determined according to the measured distance obtained according to multiple frames of historical image information. The state measurement value includes the speed and measured distance of the object.
7. The method according to claim 6, characterized in that The step of determining the height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located according to the image information and the camera height includes: Determining a difference value based on the lane line detection width in the image information and the actual width of the lane line; The height difference between the horizontal plane where the object is located and the horizontal plane where the camera is located is determined according to the difference value and the camera height.
8. The method according to claim 6, characterized in that The step of optimizing the estimated height to obtain a target height of the object includes: Identify object attributes, wherein the object attributes include object type and / or object posture, the object type includes adult or child, and the object posture includes standing or squatting; The estimated height is optimized according to the object attributes to obtain a target height of the object.
9. The method according to claim 6, characterized in that The method of using a monocular ranging algorithm to measure the distance of the object based on the position information, the camera height, and the height difference to obtain an initial distance of the object includes: Determine the target grounding point of the object according to the position information, and according to the angle between the target grounding point and the optical center line of the camera; An initial distance of the object is obtained according to the camera height, the height difference, the camera pitch angle and the included angle.
10. The method according to claim 9, characterized in that Determining the target grounding point of the object according to the position information includes: When the detection frame is a human body detection frame, determining the intersection of the object and the horizon according to the position information as the target grounding point of the object; When the detection frame is a head detection frame, the target grounding point is determined by one or more of the following methods: estimating a target grounding point of the object according to a ratio of the head detection frame to the body detection frame, a ratio of the head to the body, and the position information; The height of the object is estimated according to the head detection frame and the height of adjacent vehicles, and the target grounding point of the object is determined according to the height of the object and the position information.
11. The method according to any one of claims 1 to 3, characterized in that: The state measurement value of the object obtained according to the image information of the object and the first state prediction value of the image information are used to obtain the first state detection result of the object, including: A first state detection result of the object is obtained by using a first Kalman filter according to a state measurement value of the object obtained from image information of the object and a first state prediction value of the image information; Alternatively, obtaining a second state detection result of the object based on a state measurement value of the object obtained from the image information of the object and a second state prediction value of the image information includes: A second state detection result of the object is obtained by using a second Kalman filter based on the state measurement value of the object obtained according to the image information of the object and the second state prediction value of the image information.
12. An object state detection device, characterized in that: The device comprises: A first state detection module is used to obtain a state measurement value of the object and a first state prediction value of the image information according to the image information of the object, and obtain a first state detection result of the object, wherein the first state detection module is configured such that: a confidence level of the prediction value is greater than a confidence level of the measurement value; a second state detection module, configured to obtain a state measurement value of the object and a second state prediction value of the image information according to the image information of the object when the measured distance of the object determined based on the state measurement value is less than or equal to a preset distance, and obtain a second state detection result of the object, wherein the second state detection module is configured such that: a confidence level of the measurement value is greater than a confidence level of the prediction value; A fusion module is used to obtain a state output result of the object according to the first state detection result and the second state detection result, wherein the state output result includes a speed and a target distance.
13. An electronic device, characterized in that: include: Memory, processor; The memory is used to store computer programs / instructions; The processor is configured to implement the method according to any one of claims 1 to 11 according to the computer program / instructions stored in the memory.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program / instruction, and the computer program / instruction is used to implement the method according to any one of claims 1 to 11 when executed by a processor.
15. A computer program product, characterized in that The computer program product comprises a computer program / instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.