An object detection method and an electronic device
By using V2X technology to obtain object location and category information in object detection, and combining object detection algorithm and NMS algorithm, the problems of missed detection and missed detection in the existing technology are solved, achieving higher detection accuracy and accuracy.
Patent Information
- Application Number
- CN202110221889.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-27
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-02-27
AI Technical Summary
The prior art often has missed and missed detection problems in object detection, which cannot meet user needs, especially under conditions such as weather impact, object occlusion or camera angle deviation.
An object detection method is adopted, using V2X technology to obtain the location and category information of the object, and combining the object detection algorithm to identify the object, suppress the adjustment of the NMS algorithm and candidate box information through maximum value, improve the accuracy of the detection box and reduce the chance of missed detection and missed detection.
It improves the accuracy of object detection, reduces the chance of missed and missed detection, and meets users' needs for high accuracy.
Smart Images

Figure CN115063733B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to an object detection method and an electronic device. Background Art
[0002] With the development of artificial intelligence (AI) technology, the application of AI image recognition is becoming more and more extensive. For example, cameras on the road collect images, and deep learning algorithms are used to detect objects in the images to obtain information such as the category and location of the objects.
[0003] Due to conditional limitations, such as visibility affected by weather conditions, objects in the image being blocked, and deviation of the camera shooting angle, problems such as missed detection and misdetection often occur in object detection, which cannot meet the user's needs. Summary of the Invention
[0004] To solve the above technical problems, this application provides an object detection method and an electronic device, which can improve the accuracy of object detection, reduce the probability of missed detection and misdetection, and meet the user's needs.
[0005] First aspect, an embodiment of the present application provides an object detection method, which is applied to an electronic device including a camera device. The electronic device is fixedly installed, and the electronic device communicates wirelessly with a mobile device. The mobile device is located within a certain range from the electronic device and is located on the object; the method includes: the electronic device uses the camera device to capture or record an image at a first angle, and the image includes at least one object; identify the objects included in the image, and obtain one or more candidate boxes for each object in the objects and a candidate box information corresponding to each candidate box; the candidate box information includes the position of the object in the image, the detection probability of the object, the first object category of the object, the maximum intersection over union (max IoU) between the other candidate boxes overlapping with the candidate box and the candidate box; within a preset time period before or after obtaining the image, receive a first message from the mobile device, and the first message includes the position information of the object and the second object category of the object; according to the position information of the object, the position information of the electronic device, the first angle, and the position of the object in the image, determine at least one first candidate box corresponding to the first message from one or more candidate boxes; obtain one or more second candidate boxes with a detection probability greater than or equal to a preset detection probability threshold from the at least one first candidate box; obtain a detection box of the object from the one or more second candidate boxes through the non-maximum suppression (NMS) algorithm; the candidate box information corresponding to the detection box reveals the information of the object; wherein, the first object category and the second object category are based on the same category division standard. For example, the first object category is the vehicle category, including sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc.; the second object category is also the vehicle category, including sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc. Or, the first object category is the human category, including infants, children, adults, the elderly, etc.; the second object category is also the human category, including infants, children, adults, the elderly, etc.
[0006] Wherein, when the object is not occluded, the candidate box of the object only includes the complete contour of the object.
[0007] In this method, the V2X technology (receiving the first message) is used to obtain information such as the position and category of the object. According to the obtained information such as the position and category of the object, combined with the information such as the position and category of the object identified by the object detection algorithm, the detection box is determined from multiple candidate boxes; compared with the method of identifying objects solely using the object detection algorithm, the accuracy of the detection box determined from the candidate boxes is improved; the probability of missed detection and misdetection in object detection is reduced.
[0008] According to the first aspect, the object includes at least one of the following: vehicle, human; the position information includes position coordinates.
[0009] Second aspect, an object detection method provided by an embodiment of the present application is applied to an electronic device including a camera device. The electronic device is fixedly installed, wirelessly communicates with a mobile device, the mobile device is within a certain range from the electronic device, and the mobile device is located on the object. The method includes: the electronic device captures or photographs an image at a first angle through the camera device, and the image includes at least one object; identifies the objects included in the image, and obtains one or more candidate boxes for each object in the objects and a candidate box information corresponding to each candidate box; the candidate box information includes the position of the object in the image, the detection probability of the object, the first object category of the object, other candidate boxes overlapping with the candidate box and the maximum intersection over union max IoU between the candidate box and the other candidate box; within a preset time period before or after obtaining the image, receives a first message from the mobile device, and the first message includes the position information of the object and the second object category of the object; according to the position information of the object, the position information of the electronic device, the first angle, and the position of the object in the image, determines at least one first candidate box corresponding to the first message from one or more candidate boxes; adjusts parameters corresponding to the at least one first candidate box, and the parameters are related to the candidate box information; obtains one or more second candidate boxes with a detection probability greater than or equal to a preset detection probability threshold from the at least one first candidate box and the candidate boxes other than the first candidate box; obtains a detection box of the object through the non-maximum suppression NMS algorithm; the candidate box information corresponding to the detection box is the information revealing the object; wherein, the first object category and the second object category are based on the same category division standard. For example, the first object category is the vehicle category, including sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc.; the second object category is also the vehicle category, including sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc. Or, the first object category is the human category, including infants, children, adults, the elderly, etc.; the second object category is also the human category, including infants, children, adults, the elderly, etc.
[0010] Wherein, when the object is not occluded, the candidate box of the object only includes the complete contour of the object.
[0011] In this method, the V2X technology (receiving the first message) is used to obtain information such as the position and category of the object. According to the obtained information such as the position and category of the object, relevant parameters in the process of determining the detection box from multiple candidate boxes are adjusted, so that the information of the detection box determined from the candidate boxes is more accurate and the accuracy is higher; the probability of missed detection and misdetection in object detection is reduced.
[0012] According to the second aspect, the object includes at least one of the following: vehicle, human; the position information includes position coordinates.
[0013] According to the second aspect, or any implementation manner of the above second aspect, adjust the parameters corresponding to at least one first candidate box; including: increasing the value of the detection probability corresponding to each first candidate box. Obtain one or more second candidate boxes whose detection probability is greater than or equal to a preset detection probability threshold; including: deleting or excluding the first candidate boxes whose detection probability is less than the preset detection probability threshold to obtain one or more second candidate boxes.
[0014] In this method, delete the candidate boxes whose detection probability is less than the detection probability threshold, that is, determine that the candidate boxes whose detection probability is less than the detection probability threshold are not detection boxes; the distance between the position of the object in the first message and the position of the object in the candidate box information of the first candidate box is less than a preset distance threshold d threshold , indicating that the probability that the object corresponding to the first candidate box and the object sending the first message are the same object is relatively large. Increasing the value of the detection probability in the candidate box information of the first candidate box increases the probability that the first candidate box is determined as a detection box and reduces the missed detection probability.
[0015] According to the second aspect, or any implementation manner of the above second aspect, increase the value of the detection probability corresponding to each first candidate box; including: after increasing Among them, d represents the distance between the position information in the first message and the position of the object in the image after coordinate conversion to the first position in the image; or, d represents the distance between the position of the object in the image and the position information in the first message after coordinate conversion to the first position in the geodetic coordinate system; d threshold represents a preset distance threshold; score oral represents the value of the detection probability corresponding to a first candidate box; score′ represents a set detection probability adjustment threshold; the detection probability threshold < score′ < 1.
[0016] According to the second aspect, or any implementation manner of the above second aspect, adjust the parameters corresponding to at least one first candidate box; including: increasing the intersection over union threshold corresponding to the first candidate box. Through the non-maximum suppression (NMS) algorithm, obtain a detection box of the object from one or more second candidate boxes; including: under the condition that the max IoU in the candidate box information of a second candidate box is less than the intersection over union threshold corresponding to the increased first candidate box, stop deleting or excluding the second candidate box from the candidate boxes and determine the second candidate box as the detection box of the object.
[0017] In this method, if the distance between the object position in the first message and the position of the object in the image in the candidate box information of the first candidate box is less than a preset distance threshold d threshold, it indicates that the probability that the object corresponding to the candidate box and the object sending the first message are the same object is relatively high, and increases the intersection over union threshold IoU corresponding to the first candidate box during the process of determining the detection box from multiple candidate boxes using the NMS algorithm th ; reduces the probability that the first candidate box is deleted, that is, increases the probability that the first candidate box is determined as the detection box, and reduces the missed detection probability.
[0018] In one implementation, increasing the intersection over union threshold corresponding to the first candidate box; includes: the corresponding intersection over union threshold of the first candidate box after increase Among them, d represents the distance between the position information in the first message, after being coordinate-transformed to the first position in the image, and the position of the object in the image; or, d represents the distance between the position of the object in the image, after being coordinate-transformed to the first position in the geodetic coordinate system, and the position information in the first message; d threshold represents a preset distance threshold; represents the preset intersection over union threshold in the NMS algorithm; IoU′ th represents the set intersection over union threshold adjustment threshold;
[0019] According to the second aspect, or any one of the above implementation manners of the second aspect, the method further includes: if the second object category of the object in the first message is inconsistent with the first object category of the object in the candidate box information of the first candidate box, determining the second object category of the object in the first message as the object category of the detection box. That is, determining the object category of the detection box according to the object category in the V2X message, and reducing the misdetection probability.
[0020] Among them, if the second object category of the object in the first message is inconsistent with the first object category of the object in the candidate box information of the first candidate box, determining the second object category of the object in the first message as the object category of the detection box; includes: the Among them, d represents the distance between the position information in the first message, after being coordinate-transformed to the first position in the image, and the position of the object in the image; or, d represents the distance between the position of the object in the image, after being coordinate-transformed to the first position in the geodetic coordinate system, and the position information in the first message; d threshold represents a preset distance threshold, class v represents the second object category of the object in the first message, class oral represents the first object category of the object in the candidate box information of the first candidate box.
[0021] According to a second aspect, or any implementation manner of the above second aspect, based on the position information of an object, the position information of an electronic device, and a first angle, and the position of the object in an image, determine at least one first candidate box corresponding to a first message from one or more candidate boxes; including: traversing all candidate box information corresponding to the object, comparing D obtained by converting according to the candidate box information with a preset distance threshold, and determining a candidate box with D less than the preset distance threshold as a first candidate box; where D represents the distance between the position information of the object, after being coordinate-converted to a first position in the image according to the position information of the electronic device and the first angle, and the position of the object in the image.
[0022] In this method, if the distance between the position of the object in the first message and the position of the object in the image in the candidate box information of the candidate box is less than a preset distance threshold d threshold , it indicates that the probability that the object corresponding to this candidate box and the object sending the first message are the same object is relatively high. Determine this candidate box as a first candidate box and adjust the parameters corresponding to this first candidate box.
[0023] According to a second aspect, or any implementation manner of the above second aspect, identify the objects included in the image; including: using the YOLO algorithm or the Fast_RCNN algorithm to identify the objects included in the image.
[0024] In a third aspect, an embodiment of the present application provides an electronic device. The electronic device includes: a processor; a memory; and a computer program, where the computer program is stored in the memory; when the computer program is executed by the processor, the electronic device is caused to execute the methods described in the above first aspect and any implementation manner of the first aspect, or the above second aspect and any implementation manner of the second aspect.
[0025] For the technical effects corresponding to the third aspect and any implementation manner of the third aspect, reference can be made to the technical effects corresponding to the above first aspect and any implementation manner of the first aspect, or the above second aspect and any implementation manner of the second aspect, which will not be elaborated here.
[0026] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device is caused to execute the methods described in the first aspect and any implementation manner of the first aspect, or the above second aspect and any implementation manner of the second aspect.
[0027] For the technical effects corresponding to the fourth aspect and any implementation manner of the fourth aspect, reference can be made to the technical effects corresponding to the above first aspect and any implementation manner of the first aspect, or the above second aspect and any implementation manner of the second aspect, which will not be elaborated here.
[0028] In a fifth aspect, a computer program product is provided. When running on a computer, it causes the computer to execute the methods in the first aspect and any implementation manner of the first aspect, or the methods in the second aspect and any implementation manner of the second aspect as described above.
[0029] For the technical effects corresponding to the fifth aspect and any implementation manner of the fifth aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect, or the technical effects corresponding to the second aspect and any implementation manner of the second aspect as described above, which will not be elaborated herein. Description of the Drawings
[0030] Figure 1 It is a schematic diagram of a scenario for AI image recognition;
[0031] Figure 2 It is a schematic diagram of object detection results;
[0032] Figure 3 It is a schematic diagram of a deployment instance of RSU and OBU;
[0033] Figure 4 It is a scenario example diagram applicable to the object detection method provided by the embodiment of the present application;
[0034] Figure 5 It is a schematic diagram of the architecture of an electronic device applicable to the object detection method provided by the embodiment of the present application;
[0035] Figure 6 It is a schematic diagram of the flow of the object detection method provided by the embodiment of the present application;
[0036] Figure 7 It is a scenario example diagram of the object detection method provided by the embodiment of the present application;
[0037] Figure 8 It is a schematic diagram of the conversion relationship between the geodetic coordinate system and the world coordinate system;
[0038] Figure 9 It is a schematic diagram of the conversion relationship between the world coordinate system and the pixel coordinate system;
[0039] Figure 10 It is a schematic diagram of the structure of the object detection electronic device provided by the embodiment of the present application. Detailed Embodiments
[0040] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the relationship between associated objects and indicates that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0041] Reference to "one embodiment" or "some embodiments" or the like described in this specification means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way. The term "connection" includes direct connection and indirect connection, unless otherwise stated.
[0042] Hereinafter, the terms "first" and "second" are only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0043] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0044] With the development of AI technology, the application of AI image recognition is becoming more and more extensive. Figure 1Shows a scenario of AI image recognition. An electronic device 100 is installed on a bracket beside the road. The electronic device 100 includes an AI camera, or the electronic device 100 is an AI camera. The electronic device 100 collects intersection images and performs image recognition on the collected images. For example, the electronic device 100 uses a deep learning algorithm to detect objects in the image and can obtain information such as the category and location of the objects. Figure 2 Shows an example of object detection results. The electronic device 100 uses an object detection algorithm to perform object detection on the collected image 110. The detected object is a vehicle, and information such as the location, category, and detection probability (i.e., the probability that the object is of this category) of the vehicle in the image is obtained. Exemplarily, the electronic device 100 obtains that the vehicle category of vehicle 111 in image 110 is a sedan, and the detection probability is 99%; obtains that the vehicle category of vehicle 112 in image 110 is a truck, and the detection probability is 98%; obtains that the vehicle category of vehicle 113 in image 110 is a bus, and the detection probability is 99%.
[0045] Due to conditional limitations, such as visibility affected by weather conditions, objects in the image being blocked, and deviation of the camera shooting angle, problems such as missed detection and misdetection often occur in object detection. Exemplarily, Figure 2 vehicle 114 in [the image] is not detected, that is, vehicle 114 is missed; Figure 2 the vehicle category of vehicle 115 in [the image] is a sedan, but it is detected as an off-road vehicle, that is, vehicle 115 is misdetected.
[0046] The embodiment of the present application provides an object detection method. The electronic device supports vehicle-to-everything (V2X) wireless communication technology, uses V2X to obtain vehicle information, and combines the vehicle information obtained by V2X for judgment in object detection to reduce misdetection and missed detection and improve the accuracy of object detection.
[0047] V2X is a new generation of information and communication technology that connects vehicles to everything; where V represents vehicles and X represents any object that interacts with vehicles, for example, X can include vehicles, people, roadside infrastructure, and networks. The information interaction modes of V2X include: information interaction between vehicle and vehicle (vehicle to vehicle, V2V), between vehicle and roadside infrastructure (vehicle to infrastructure, V2I), between vehicle and people (vehicle to pedestrian, V2P), and between vehicle and network (vehicle to network, V2N).
[0048] V2X based on Cellular technology, namely C-V2X. C-V2X is a vehicle wireless communication technology evolved from cellular network communication technologies such as 3G / 4G / 5G. C-V2X includes two communication interfaces: one is the short-range direct communication interface (PC5) between terminals such as vehicles, people, and roadside infrastructures, and the other is the communication interface (Uu) between terminals such as vehicles, people, and roadside infrastructures and the network, which is used to achieve long-distance and large-scale reliable communication.
[0049] The terminals in V2X include roadside units (RSUs) and on-board units (OBUs). Exemplarily, Figure 3 An example of the deployment of an RSU and an OBU is shown. An RSU is a static entity that supports V2X applications and is deployed on the roadside, such as on a gantry beside the road; it can exchange data with other entities that support V2X applications (such as RSUs or OBUs). An OBU is a dynamic entity that supports V2X applications and is usually installed in a vehicle and can exchange data with other entities that support V2X applications (such as RSUs or OBUs).
[0050] In one embodiment, the RSU can be Figure 1 or Figure 4 the electronic device 100 in
[0051] In another embodiment, Figure 1 or Figure 4 the electronic device 100 in
[0052] In one embodiment, the OBU can be referred to as a mobile device. Further, if the OBU is installed in a vehicle, it can be called an in-vehicle device.
[0053] In another embodiment, the mobile device can include an OBU. Further, if the mobile device is installed in a vehicle, it can be called an in-vehicle device.
[0054] It should be noted that the above various embodiments can be freely combined on the premise of not being contradictory.
[0055] The object detection method provided by the embodiments of this application can be applied to Figure 4The scenario shown. An OBU is installed in each of the electronic device 100 and the vehicle 200. The electronic device 100 and each vehicle 200 can communicate with each other via V2X. For example, the OBU of the vehicle 200 periodically broadcasts a basic safety message (BSM), which includes basic information of the vehicle, such as the vehicle's driving speed, heading, position, acceleration, vehicle type, predicted path and historical path, vehicle events, etc. The communication distance between OBUs is the first distance (such as 500 meters). The electronic device 100 can obtain the basic safety messages broadcast by the vehicles 200 within the first distance around it via V2X.
[0056] In some other examples, the electronic device 100 supports the RSU function, and the communication distance between the OBU and the RSU is the second distance (such as 1000 meters). The electronic device 100 can obtain the basic safety messages broadcast by the vehicles 200 within the second distance around it via V2X. The following embodiments of this application are described by taking the example of an OBU installed in the electronic device 100. It can be understood that the object detection method provided by the embodiments of this application is also applicable to the case where the electronic device supports the RSU function.
[0057] The electronic device 100 collects road images and performs object detection on the road images using an object detection algorithm. For example, the object detection algorithm includes the YOLO (you only look once) algorithm, Fast_RCNN, etc. In object detection, first, one or more candidate recognition regions (regions of interest, RoIs), called candidate boxes (RoIs), of each detected object are obtained through calculation; then, one RoI is selected from the one or more candidate boxes, that is, the detection box of the detected object is obtained. Among them, the detected objects can include vehicles, pedestrians, road signs, etc.
[0058] For the object detection method provided by the embodiments of this application, the electronic device 100 also obtains the broadcast messages of the vehicles 200 within the first distance around it via V2X; combines the obtained broadcast messages of the surrounding vehicles, and selects one RoI from multiple candidate boxes; to improve the detection accuracy of object detection.
[0059] Optionally, the electronic device 100 can be an AI camera, or an electronic device including an AI camera, or an electronic device including a camera or a camera (different from the AI camera).
[0060] Exemplarily, Figure 5The structural schematic diagram of an electronic device is shown. The electronic device 100 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna, a wireless communication module 250, a camera sensor 260, etc.
[0061] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0062] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent components or integrated in one or more processors. In some embodiments, the electronic device 100 may also include one or more processors 210. Among them, the controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0063] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface 230 is an interface that complies with the USB standard specification, and specifically may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 230 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and peripheral devices.
[0064] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0065] The charging management module 240 is used to receive a charging input from a charger. Among them, the charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 240 may receive the charging input from a wired charger through the USB interface 230. In some embodiments of wireless charging, the charging management module 240 may receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 242, the charging management module 240 can also supply power to the electronic device through the power management module 241.
[0066] The power management module 241 is used to connect to the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives inputs from the battery 242 and / or the charging management module 240 to supply power to the processor 210, the internal memory 221, the external memory interface 220, the wireless communication module 250, etc. The power management module 241 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be disposed in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 can also be disposed in the same device.
[0067] The wireless communication function of the electronic device 100 can be implemented by the antenna and the wireless communication module 250, etc.
[0068] The wireless communication module 250 can provide solutions for wireless communications applied to the electronic device 100, including Wi-Fi, Bluetooth (BT), wireless data transmission modules (e.g., 433 MHz, 868 MHz, 915 MHz), etc. The wireless communication module 250 can be one or more devices integrating at least one communication processing module. The wireless communication module 250 receives electromagnetic waves via the antenna, filters and frequency-modulates the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 250 can also receive the signals to be sent from the processor 210, frequency-modulate and amplify them, and convert them into electromagnetic waves via the antenna for radiation.
[0069] In the embodiments of this application, the electronic device 100 can receive broadcast messages (basic security messages) through the wireless communication module.
[0070] The external memory interface 220 can be used to connect to an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0071] The internal memory 221 can be used to store one or more computer programs, and the one or more computer programs include instructions. The processor 210 can execute the object detection method provided in some embodiments of this application, as well as various applications and data processing, by running the above instructions stored in the internal memory 221. The internal memory 221 can include a code storage area and a data storage area. Among them, the code storage area can store the operating system. The data storage area can store the data created during the use of the electronic device 100, etc. In addition, the internal memory 221 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc. In some embodiments, the processor 210 can cause the electronic device 100 to execute the object detection method provided in the embodiments of this application, as well as other applications and data processing, by running the instructions stored in the internal memory 221 and / or the instructions stored in the memory provided in the processor 210.
[0072] The camera sensor 260 is used to capture still images or videos. An object generates an optical image through a lens and projects it onto the camera sensor 260. The camera sensor 260 can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The camera sensor 260 converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 210 to be converted into a digital image signal. The DSP in the processor 210 converts the digital image signal into an image signal in standard formats such as RGB and YUV. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera sensor 260, where the optical signal is converted into an electrical signal. The camera sensor 260 transmits the electrical signal to the processor 210 for processing and converts it into an image visible to the naked eye. The processor 210 can also optimize the noise, brightness, and skin color of the image through algorithms; it can also optimize parameters such as the exposure and color temperature of the shooting scene.
[0073] In the embodiments of the present application, the electronic device 100 collects road images through the camera sensor 260, and the road images collected by the camera sensor 260 are transmitted to the processor 210. The processor 210 uses object detection algorithms such as YOLO and Fast_RCNN to calculate one or more candidate boxes (RoIs) for each object in the road image. The electronic device 100 also receives broadcast messages (basic safety messages) of surrounding vehicles through the wireless communication module 250. The processor 210 receives the basic safety messages of the surrounding vehicles, and combines the obtained basic safety messages of the surrounding vehicles to screen out one RoI from multiple candidate boxes, that is, the detected object is obtained.
[0074] It should be noted that the system applicable to the embodiments of the present application may include more electronic devices than Figure 4 shown. In some embodiments, the system further includes a server (such as a cloud server). In one example, the electronic device collects road images and uploads the road images to the server. The electronic device receives the basic safety messages periodically broadcast by the OBU within the first distance from it, and forwards the obtained basic safety messages to the server. On the server, object detection algorithms such as YOLO and Fast_RCNN are used to calculate one or more candidate boxes (RoIs) for each object in the road image through calculation; and in combination with the obtained basic safety messages of the vehicles, one RoI is screened out from multiple candidate boxes, that is, the detected object is obtained.
[0075] Next, in combination with the accompanying drawings, taking Figure 4 the scenario shown as an example, the object detection algorithm provided by the embodiments of the present application will be introduced in detail.
[0076] Exemplarily, Figure 6 shows a schematic flowchart of an object detection algorithm provided by the embodiments of the present application. The object detection algorithm provided by the embodiments of the present application may include:
[0077] S601. Collect images periodically at a first angle according to a cycle of a first duration; after each image is collected, use an object detection algorithm to identify the image, and obtain one or more candidate boxes for each object in the image, and candidate box information corresponding to each candidate box.
[0078] The electronic device uses a camera to collect images periodically at a first angle according to a cycle of a first duration (such as 1 s). After each image is collected, an object detection algorithm is used to identify the image. For example, the object detection algorithm is YOLO or Fast_RCNN, etc. The object for object detection is a vehicle.
[0079] The electronic device identifies each collected image, obtains one or more candidate boxes for each object in the image, and candidate box information corresponding to each candidate box. The candidate box information includes detection probability, image position, object category (for example, vehicle category), the maximum intersection over union (max IoU) of this candidate box and other candidate boxes, etc. Among them, the detection probability is the probability that the detected object is of this object category (for example, vehicle category), also known as the score; the image position is the pixel coordinates of this candidate box in the image. For example, it includes the pixel coordinate values (x min , y min ) of the upper left corner of the candidate box, the pixel coordinate values (x max , y max ) of the lower right corner of the candidate box, the pixel coordinate values (x p , y p ) of the center point of the candidate box, etc.; the vehicle category includes sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc. The intersection over union (IoU) is the ratio of the intersection of two candidate boxes to the union of the two candidate boxes.
[0080] Exemplarily, Figure 7 FIG. shows a scenario example diagram of the object detection method provided by an embodiment of the present application.
[0081] As Figure 7 shown in (a) of, object detection is performed on the first image, and candidate boxes 0, 1, 2, 3, 4, 5, 6, and 7 are obtained. The upper left pixel coordinate values (x min , y min ) of candidate boxes 0 - 7, the lower right pixel coordinate values (x max , y max ), the pixel coordinate values (x p , y p ) of the center point of the candidate box, the detection probability (score), and the maximum intersection over union (max IoU) information are shown in Table 1.
[0082] Table 1
[0083]
[0084]
[0085] S602. Receive a first message from the mobile device of the object; the first message includes the position information, category information, etc. of the object.
[0086] Taking the vehicle as the detection object and the OBU as the mobile device. The OBU in the vehicle periodically sends a first message at a second time interval (such as 100 ms). The first message includes information such as the identification of the vehicle, the vehicle position (latitude, longitude, and altitude), the vehicle category (including sedans, SUVs, buses, trucks, school buses, fire trucks, ambulances, police cars, etc.), and the message sending time. For example, the first message is a BSM message. Exemplarily, the BSM message includes the information shown in Table 2.
[0087] Table 2
[0088] BSM message Description vehicle ID Vehicle identification DSecond Coordinated Universal Time (UTC time) VehicleSize Vehicle size (including length, width, and height) AccelerationSet4Way Vehicle four-axis acceleration Heading Vehicle heading angle Speed Current vehicle driving speed VehicleClassification Vehicle category Position3D Vehicle position (latitude, longitude, altitude) … …
[0089] The electronic device receives the first messages periodically sent by the surrounding vehicles at the second time interval and saves the received first messages.
[0090] S603. Determine the first image corresponding to the first message according to whether the difference between the reception time of the first message and the acquisition time of the image is within a preset time interval.
[0091] In one implementation, the electronic device determines the acquisition time point of the image as the first moment, and determines the first messages whose message sending times are within the second time interval before or after the first moment as the first images corresponding to the first message.
[0092] Exemplarily, the first image is the image at 18:20:01 captured by the camera. The period for the vehicle to send the first message is 100 milliseconds (i.e., the second time interval is 100 milliseconds). The camera determines the first message whose sending time is between (18:20:01 - 100 milliseconds) and 18:20:01 as the first message corresponding to the first image.
[0093] Optionally, in some examples, if the first messages saved by the electronic device include multiple first messages of the same vehicle, the last first message among the multiple first messages of the same vehicle is retained.
[0094] S604. According to the position information of the object in the first message, combined with the position information of the electronic device and the first angle, obtain the first position of the object in the first image through coordinate transformation.
[0095] The electronic device obtains information such as the identification, vehicle position, vehicle category, and message sending time of the corresponding vehicle according to each first message in the first image corresponding to the vehicle. Exemplarily, based on the BSM message of vehicle 1, the camera obtains that the vehicle position of vehicle 1 is longitude L1, latitude B1, and altitude H1; the vehicle category is a sedan. Based on the BSM message of vehicle 2, the camera obtains that the vehicle position of vehicle 2 is longitude L2, latitude B2, and altitude H2; the vehicle category is a sedan. Based on the BSM message of vehicle 3, the camera obtains that the vehicle position of vehicle 3 is longitude L3, latitude B3, and altitude H3; the vehicle category is a sedan. Based on the BSM message of vehicle 4, the camera obtains that the vehicle position of vehicle 4 is longitude L4, latitude B4, and altitude H4; the vehicle category is a sedan. Based on the BSM message of vehicle 5, the camera obtains that the vehicle position of vehicle 5 is longitude L5, latitude B5, and altitude H5; the vehicle category is a sedan.
[0096] In one implementation, the latitude, longitude, and altitude (B, L, H) of the vehicle position in the first message are converted into pixel coordinates (x v , y v ).
[0097] I. Convert the latitude, longitude, and altitude values of the vehicle position into coordinate values in the world coordinate system.
[0098] As Figure 8 shown, the latitude, longitude, and altitude (B, L, H) are coordinate values in the geodetic coordinate system. The definition of latitude B is the angle between the normal of a certain point on the ground and the equatorial plane; starting from the equatorial plane, it is negative southward, with a range of -90° to 0°, and positive northward, with a range of 0° to 90°. Longitude L starts from the starting geodetic meridian plane, is positive eastward and negative westward, with a range of -180° to 180°. The distance from a certain point along the normal to the ellipsoid surface is called the altitude H of that point.
[0099] The world coordinate system takes the center O of the ellipsoid as the coordinate origin, the intersection line of the starting meridian plane and the equatorial plane as the X-axis, the direction orthogonal to the X-axis on the equatorial plane as the Y-axis, and the rotation axis of the ellipsoid as the Z-axis, with the three directions forming a right-handed system.
[0100] According to the following formula ①, convert the coordinate values (B, L, H) (B is latitude, L is longitude, and H is altitude) of the geodetic coordinate system into world coordinate system coordinate values (x, y, z).
[0101] x = (N + H)cosBcosL
[0102] y = (N + H)cosBsinL ①
[0103] z = (N(1 - e 2 ) + H)sinB
[0104] Among them, N is the radius of the prime vertical circle, and e is the first eccentricity of the Earth. Let the equatorial radius of the reference ellipsoid be a, and the polar radius of the reference ellipsoid be b. In the definition of the reference ellipsoid, a is greater than b. Then: e 2 =(a 2 -b 2 ) / a 2 , N = a / (1 - e 2 sin 2 B) 1 / 2 .
[0105] II. Convert the coordinate values in the world coordinate system into pixel coordinates in the image.
[0106] There are four coordinate systems in the camera, namely the world coordinate system, the camera coordinate system, the image coordinate system, and the pixel coordinate system. As Figure 9 shown, the coordinate system X c Y c Z c is the camera coordinate system. The camera coordinate system takes the camera's light point as the origin, the X-axis and Y-axis are parallel to two sides of the image, and the Z-axis coincides with the optical axis. The coordinate system X i Y i Z i is the image coordinate system. The image coordinate system takes the intersection point of the optical axis and the projection plane as the origin, the X-axis and Y-axis are parallel to two sides of the image, and the Z-axis coincides with the optical axis. The coordinate system uv is the pixel coordinate system. The pixel coordinate system and the image coordinate system are in the same plane, but with different origins; looking from the camera's light point towards the projection plane, the upper left corner of the projection plane is the origin of the pixel coordinate system.
[0107] Convert the world coordinate system coordinate values (x, y, z) into pixel coordinate system coordinate values (x v , y v ) according to the following formula ②.
[0108]
[0109] Among them, the matrix is the camera internal parameter matrix, which is related to the internal hardware parameters of the camera; the matrix is the camera external parameter matrix, which is related to the relative positions of the world coordinate system and the pixel coordinate system (for example, the shooting angle of the camera, which can also be called the first angle; the relative position between the camera and the vehicle, etc.). For example, a point in the world coordinate system is converted into a point in the pixel coordinate system through rotation translation ; s is the value of the object point in the Z-axis direction of the pixel coordinate system. For the content of r 11 -r 33 , t1 - t3, reference can be made to the relevant technologies in this field, and details will not be elaborated here.
[0110] The pixel coordinates in the converted image are the first position of the object in the first image.
[0111] The above implementation is introduced by taking the conversion of the vehicle position in the first message from the coordinate values (longitude, latitude, and altitude) in the geodetic coordinate system to the pixel coordinate values in the pixel coordinate system as an example. It can be understood that in some other implementations, the image position in the candidate box information can also be converted from the pixel coordinate values in the pixel coordinate system to the coordinate values in the geodetic coordinate system, and then compared with the vehicle position in the first message; the embodiments of the present application do not limit this.
[0112] S605. Determine at least one first candidate box related to the object in the image according to whether the difference between the first position and the position of the object in the image in the candidate box information is within a preset threshold.
[0113] Traverse the candidate boxes of the first image. If the first position (x v , y v ) and the image position (x p , y p ) in the candidate box information of the candidate box, the first distance d between them is less than the preset distance threshold d threshold (for example, 0.5 meters), determine that the candidate box is the first candidate box corresponding to the first message. Among them,
[0114] S606. Adjust the parameters corresponding to at least one first candidate box, and the parameters are related to the candidate box information.
[0115] In some embodiments, if d is less than d threshold , it means that the probability that the vehicle corresponding to the candidate box and the vehicle sending the first message is the same vehicle is relatively high, then increase the score of the candidate box corresponding to the first message. In this way, the probability that the candidate box is determined as the detection box is increased, and the missed detection probability is reduced.
[0116] In one implementation, the detection probability of the candidate box corresponding to the first message is increased according to the first distance d. Exemplarily, the adjusted detection probability score n of the candidate box corresponding to the first message is obtained according to the following formula ③. Among them, the adjustment coefficient is the confidence level of score n , the smaller d is, the greater the confidence level; score oral is the detection probability value in the candidate box information, score′ is the set detection probability adjustment threshold, and it satisfies score th <score′<1.
[0117]
[0118] In some embodiments, if d is less than d threshold , it indicates that the probability that the vehicle corresponding to the candidate box and the vehicle sending the first message is the same vehicle is relatively high, then increase the intersection over union threshold IoU corresponding to the candidate box in the process of determining the detection box from multiple candidate boxes using the NMS algorithm th ; in this way, the probability of deleting this candidate box is reduced, that is, the probability of determining this candidate box as the detection box is increased, and the probability of missed detection is reduced.
[0119] In one implementation, increase the intersection over union threshold corresponding to the candidate box of the first message according to the first distance d. Exemplarily, obtain the intersection over union threshold corresponding to the candidate box of the first message according to the following formula ④ where the adjustment coefficient is the confidence of, the smaller d is, the greater the confidence; is the preset IoU th value in the NMS algorithm, IoU′ th is the set intersection over union threshold adjustment threshold, satisfying
[0120]
[0121] S607. Obtain a detection box corresponding to the object according to the detection probability and in combination with a specific algorithm; the candidate box information of the detection box is the information of the object finally obtained.
[0122] In some embodiments, delete the candidate boxes with a detection probability (score) less than the detection probability threshold score th (for example, 0.85), that is, determine that the candidate boxes with a detection probability less than the detection probability threshold are not detection boxes.
[0123] Exemplarily, set score th to 0.85. The detection probabilities of the candidate boxes with serial numbers 2, 4, and 7 in Table 1 are less than score th . Receive a first message, and the first distance between the vehicle position in the first message and the image position (192, 223) of the candidate box with serial number 7 is less than d threshold , and this first message corresponds to the candidate box with serial number 7. Set score′ to 0.95. Exemplarily, as shown in Table 3, calculate that the score n of the candidate box with serial number 7 is 0.87 and the confidence is 0.5.
[0124] Table 3
[0125]
[0126]
[0127] In this way, the candidate boxes numbered 2 and 4 are deleted; the score of the candidate box numbered 7 n is greater than the score th , and the candidate box numbered 7 is retained; the probability of missing detection is reduced. After screening according to the detection probability, the remaining candidate boxes are the candidate boxes numbered 0, 1, 3, 5, 6, and 7.
[0128] In some embodiments, a specific algorithm, such as the non-maximum suppression (NMS) algorithm, is used to determine the detection box from multiple candidate boxes.
[0129] The NMS algorithm is introduced as follows:
[0130] 1. Select the one with the largest score from multiple RoIs, denoted as box_best, and retain it.
[0131] 2. Calculate the IoU between box_best and the remaining RoIs.
[0132] 3. If the IoU of a candidate box and box_best is greater than the intersection-over-union threshold IoU th (such as 0.5), delete the candidate box. Since the IoU of two candidate boxes is greater than IoU th , the probability that these two candidate boxes represent the same detection object is relatively high. Delete the candidate box with a smaller score and retain the candidate box with a higher score; in this way, the candidate box with a lower score of the same detection object can be removed from multiple candidate boxes, and the candidate box with a higher score of the same detection object is determined as the detection box.
[0133] 4. Select the one with the largest score from the remaining RoIs, and loop through steps 1-3.
[0134] Exemplarily, the detection box is determined from the candidate boxes shown in Table 1 according to the NMS algorithm. After screening according to the detection probability, the remaining candidate boxes in Table 1 are the candidate boxes numbered 0, 1, 3, 5, 6, and 7. The preset intersection-over-union threshold is 0.5. The maximum intersection-over-union max IoU of the candidate boxes numbered 0 and 1 is 0.542. Since the score value 0.91 of the candidate box numbered 1 is greater than the score value 0.87 of the candidate box numbered 0, and 0.542>0.5, the candidate box numbered 0 with a lower score value will be deleted. A first message is received, and the first distance between the vehicle position in the first message and the image position (13.5, 51) of the candidate box numbered 0 is less than d threshold , and the first message corresponds to the candidate box numbered 0; According to Calculate to obtain the confidence level to be 0.6. Exemplarily, as shown in Table 3, set IoU′ th to be 0.6, and calculate and obtain the intersection-over-union threshold corresponding to the candidate box with serial number 0 according to Formula ④ to be 0.56; since 0.542 < 0.56, the candidate box with serial number 0 will not be deleted. In this way, the candidate boxes with serial numbers 0 and 1 are both retained. The maximum intersection-over-union max IoU of the candidate boxes with serial numbers 5 and 6 is 0.501. Since the score value 0.98 of the candidate box with serial number 6 is greater than the score value 0.86 of the candidate box with serial number 5, and 0.501 > 0.5, the candidate box with serial number 5 with a lower score value is deleted. In this way, after screening using the NMS algorithm, the remaining candidate boxes in Table 1 are the candidate boxes with serial numbers 0, 1, 3, 6, and 7. As Figure 7 shown in (b) of, the candidate boxes with serial numbers 0, 1, 3, 6, and 7 are determined as detection boxes.
[0135] In some embodiments, the electronic device further determines the vehicle category of the detection object according to the first distance d.
[0136] In one implementation, if d is less than d threshold , it indicates that the probability that the vehicle corresponding to the candidate box and the vehicle sending the first message is the same vehicle is relatively high, and the vehicle category in the first message is determined as the vehicle category of the detection object. Since the accuracy rate of the vehicle category in the first message is relatively high, the accuracy of object detection is improved in this way.
[0137] Exemplarily, the vehicle category of the detection object is obtained according to the following Formula ⑤. Among them, the adjustment coefficient is the confidence level of class, the smaller d is, the greater the confidence level; class oral is the vehicle category in the candidate box information, class v is the vehicle category in the first message.
[0138]
[0139] In one example, the vehicle category in the first message is inconsistent with the vehicle category of the candidate box corresponding to the first message, and the vehicle category in the first message is determined as the vehicle category of the detection box. Exemplarily, the vehicle category in the candidate box information of the candidate box with serial number 0 is an off-road vehicle, and the vehicle category in the first message corresponding to the candidate box with serial number 0 is a sedan; then the vehicle category of the detection box with serial number 0 is determined to be a sedan.
[0140] In one example, when the first message corresponding to the candidate box is not received, the vehicle category of the candidate box is determined as the vehicle category of the detection box. Exemplarily, the vehicle category in the candidate box information of candidate box No. 1 is a sedan, and the first message corresponding to candidate box No. 1 is not received; then it is determined that the vehicle category of detection box No. 1 is a sedan.
[0141] In one example, when the vehicle category in the first message is consistent with the vehicle category of the candidate box corresponding to the first message, the vehicle category in the first message is determined as the vehicle category of the detection box. Exemplarily, the vehicle category in the candidate box information of candidate box No. 7 is a sedan, and the vehicle category in the first message corresponding to candidate box No. 7 is a sedan; then it is determined that the vehicle category of detection box No. 7 is a sedan.
[0142] The object detection method provided by the embodiments of the present application uses V2X technology to obtain information such as the position and vehicle category of a vehicle. According to the obtained information such as the position and vehicle category of the vehicle, the parameters in the process of determining the detection object using the object detection algorithm are adjusted, so that the information of the detection box determined from the candidate boxes is more accurate and has higher precision; the probability of missed detection and misdetection in object detection is reduced.
[0143] It should be noted that in each embodiment of the present application, the camera can be replaced by a camera, and the camera can be replaced by a camera.
[0144] It should be noted that all or part of each embodiment of the present application can be freely and arbitrarily combined. The combined technical solution is also within the scope of the present application.
[0145] It can be understood that in order to implement the above functions, the above electronic device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0146] The embodiments of the present application can divide the above electronic device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0147] In one example, please refer to Figure 10 , which shows a possible structural schematic diagram of the electronic device involved in the above embodiments. The electronic device 1000 includes: an image acquisition unit 1010, a processing unit 1020, a storage unit 1030, and a communication unit 1040.
[0148] Among them, the image acquisition unit 1010 is used to acquire images.
[0149] The processing unit 1020 is used to control and manage the operations of the electronic device 1000. For example, it can be used to perform object detection on images using an object detection algorithm to obtain candidate box information; determine the candidate box corresponding to each first message; adjust the candidate box parameters according to the first distance between the vehicle position in the first message and the image position in the candidate box information of the corresponding candidate box; determine the detection box from multiple candidate boxes according to the adjusted candidate box parameters; and / or be used for other processes of the technologies described herein.
[0150] The storage unit 1030 is used to save the program code and data of the electronic device 1000. For example, it can be used to save the received first message; or be used to save candidate box information, etc.
[0151] The communication unit 1040 is used to support the communication between the electronic device 1000 and other electronic devices. For example, it can be used to receive the first message.
[0152] Of course, the unit modules in the above electronic device 1000 include but are not limited to the above image acquisition unit 1010, processing unit 1020, storage unit 1030, and communication unit 1040. For example, the electronic device 1000 may further include a power supply unit, etc. The power supply unit is used to supply power to the electronic device 1000.
[0153] Among them, the image acquisition unit 1010 may be a camera sensor. The processing unit 1020 may be a processor or a controller. For example, it may be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The storage unit 1030 may be a memory. The communication unit 1040 may be a transceiver, a transceiver circuit, etc.
[0154] For example, the image acquisition unit 1010 is a camera sensor (such as Figure 5 the camera sensor 260 shown), the processing unit 1020 is a processor (such as Figure 5 the processor 210 shown), the storage unit 1030 can be a memory (such as Figure 5 the internal memory 221 shown), and the communication unit 1040 can be referred to as a communication interface, including a wireless communication module (such as Figure 5 the wireless communication module 250 shown). The electronic device 1000 provided by the embodiments of the present application can be Figure 5 the electronic device 100 shown. Among them, the above camera sensor, processor, memory, communication interface, etc. can be connected together, for example, connected through a bus.
[0155] The embodiments of the present application also provide a computer-readable storage medium, in which computer program code is stored. When the processor executes the computer program code, the electronic device executes the method in the above embodiments.
[0156] The embodiments of the present application also provide a computer program product. When the computer program product runs on a computer, the computer executes the method in the above embodiments.
[0157] Among them, the electronic device 1000, computer-readable storage medium or computer program product provided by the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0158] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the electronic device is divided into different functional modules to complete all or part of the functions described above.
[0159] In several embodiments provided by the present application, it should be understood that the disclosed electronic device and method can be implemented in other ways. For example, the electronic device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another electronic device, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the electronic device or unit can be electrical, mechanical or other forms.
[0160] In addition, in each embodiment of the present application, each functional unit may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0161] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which may be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROMs, magnetic disks, or optical discs that can store program codes.
[0162] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An object detection method, which is applied to an electronic device including a camera device. The electronic device is fixedly arranged, and the electronic device communicates wirelessly with a mobile device. The mobile device is within a certain range from the electronic device, and the mobile device is located on the object. It is characterized in that, The method includes: The electronic device captures or records an image at a first angle through the imaging device, and the image includes at least one object; Identify the objects included in the image, and obtain one or more candidate bounding boxes for each object in the objects and a candidate bounding box information corresponding to each candidate bounding box; the candidate bounding box information includes the position of the object in the image, the detection probability of the object, the first object category of the object, the maximum intersection over union max IoU between the other candidate bounding boxes overlapping with the candidate bounding box and the candidate bounding box; Within a preset time period before or after obtaining the image, receive a first message from the mobile device, and the first message includes the position information of the object and the second object category of the object; Based on the position information of the object, the position information of the electronic device, the first angle, and the position of the object in the image, determine at least one first candidate bounding box corresponding to the first message from the one or more candidate bounding boxes; Wherein, the first candidate bounding box is determined according to whether the difference between the first position and the position of the object in the image in the candidate bounding box information is within a preset threshold, and the first position is obtained through coordinate transformation based on the position information of the object in the first message, combined with the position information of the electronic device and the first angle; Obtain one or more second candidate bounding boxes with a detection probability greater than or equal to a preset detection probability threshold from the at least one first candidate bounding box; Through the non-maximum suppression NMS algorithm, obtain a detection bounding box of the object from the one or more second candidate bounding boxes; the candidate bounding box information corresponding to the detection bounding box is information revealing the object; Wherein, the first object category and the second object category are based on the same category division standard, and the first object category and the second object category are the same category.
2. The method according to claim 1, characterized in that The object includes at least one of the following: vehicle, person; the position information includes position coordinates.
3. An object detection method is applied to an electronic device including a camera device. The electronic device is fixedly arranged and wirelessly communicates with a mobile device. The mobile device is within a certain range from the electronic device and is located on the object. It is characterized in that, The method includes: The electronic device captures or records an image at a first angle through the imaging device, and the image includes at least one object; Identify the objects included in the image, and obtain one or more candidate bounding boxes for each object in the objects and a candidate bounding box information corresponding to each candidate bounding box; the candidate bounding box information includes the position of the object in the image, the detection probability of the object, the first object category of the object, the maximum intersection over union max IoU between the other candidate bounding boxes overlapping with the candidate bounding box and the candidate bounding box; Within a preset time period before or after obtaining the image, receive a first message from the mobile device, and the first message includes the position information of the object and the second object category of the object; Based on the position information of the object, the position information of the electronic device, the first angle, and the position of the object in the image, determine at least one first candidate bounding box corresponding to the first message from the one or more candidate bounding boxes; Among them, the first candidate box is determined according to whether the difference between the first position and the position of the object in the image in the candidate box information is within a preset threshold. The first position is obtained through coordinate transformation based on the position information of the object in the first message, combined with the position information of the electronic device and the first angle. Adjust the parameters corresponding to the at least one first candidate box, where the parameters are related to the candidate box information. Obtain one or more second candidate boxes with a detection probability greater than or equal to a preset detection probability threshold from the at least one first candidate box and the candidate boxes other than the first candidate box. Obtain a detection box of the object from the one or more second candidate boxes through the non-maximum suppression (NMS) algorithm; the candidate box information corresponding to the detection box reveals the information of the object. Among them, the first object category and the second object category are based on the same category division standard, and the first object category and the second object category are of the same category.
4. The method according to claim 3, characterized in that, The object includes at least one of the following: vehicle, person; the position information includes position coordinates.
5. The method according to claim 4, characterized in that The adjusting the parameters corresponding to the at least one first candidate box includes: increasing the value of the detection probability corresponding to each first candidate box.
6. The method according to any one of claims 3-5, characterized in that, The obtaining one or more second candidate boxes with a detection probability greater than or equal to a preset detection probability threshold includes: deleting or excluding the first candidate boxes with a detection probability less than the preset detection probability threshold to obtain one or more second candidate boxes.
7. The method according to claim 5, characterized in that, The increasing the value of the detection probability corresponding to each first candidate box includes: Among them, d represents the distance between the position information in the first message, after being coordinate-transformed to the first position in the image, and the position of the object in the image; d threshold represents a preset distance threshold; score oral represents the value of the detection probability corresponding to a first candidate box; score′ represents a set detection probability adjustment threshold; the detection probability threshold < score′ < 1.
8. The method according to claim 3, wherein The adjusting the parameters corresponding to the at least one first candidate box includes: increasing the intersection over union (IoU) threshold corresponding to the first candidate box.
9. The method according to claim 8, characterized in that, The increasing the intersection over union (IoU) threshold corresponding to the first candidate box includes: Among them, d represents the distance between the position information in the first message, after being coordinate-converted to the first position in the image, and the position of the object in the image; d threshold represents a preset distance threshold; represents the intersection over union (IoU) threshold preset in the NMS algorithm; IoU′ th represents a set IoU threshold adjustment threshold; 10. The method according to claim 8 or 9, characterized in that The obtaining a detection box of the object from the one or more second candidate boxes through the non-maximum suppression (NMS) algorithm includes: Through the NMS algorithm, when the max IoU in the candidate box information of a second candidate box is less than the intersection over union (IoU) threshold corresponding to the increased first candidate box, stop deleting or excluding the second candidate box from the candidate boxes, and determine the second candidate box as the detection box of the object.
11. The method according to claim 3, characterized in that, The method further includes: If the second object category of the object in the first message is inconsistent with the first object category of the object in the candidate box information of the first candidate box, determine the second object category of the object in the first message as the object category of the detection box.
12. The method according to claim 11, wherein The consistency or inconsistency between the second object category of the object in the first message and the first object category of the object in the candidate box information of the first candidate box is determined by the following formula: Among them, d represents the distance between the position information in the first message, after being coordinate-transformed to the first position in the image, and the position of the object in the image; d threshold represents a preset distance threshold, class v represents the second object category of the object in the first message, class oral represents the first object category of the object in the candidate box information of the first candidate box.
13. The method according to any one of claims 1 to 3, characterized in that, According to the position information of the object, the position information of the electronic device, and the first angle, the position of the object in the image, determine at least one first candidate box corresponding to the first message from the one or more candidate boxes, including: Traverse all the candidate box information corresponding to the object, compare the D obtained by converting according to the candidate box information with a preset distance threshold, and determine the candidate box with D less than the preset distance threshold as the first candidate box; wherein, D represents the position information of the object, which is the distance between the first position in the image after coordinate conversion according to the position information of the electronic device and the first angle and the position of the object in the image.
14. The method according to claim 13, wherein The recognition of the object included in the image includes: using the YOLO algorithm or the Fast_RCNN algorithm to recognize the object included in the image.
15. An electronic device, characterized in that, The electronic device includes: a processor; a memory; and a computer program, wherein the computer program is stored on the memory, and when the computer program is executed by the processor, the electronic device executes the method described in any one of claims 1-14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, and when the computer program runs on an electronic device, the electronic device executes the method described in any one of claims 1-14.
Citation Information
Patent Citations
Object detection method, device and system
CN108268869A
Image target detection method and device
CN108960266A