Element collection method and related device

By performing element fusion and noise filtering on low-confidence video frames in the terminal device, the problem of omission of identification results caused by transmission costs in the prior art is solved, and higher identification accuracy and efficiency are achieved.

CN120298939APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410034086.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the existing road element acquisition method, due to the transmission cost limitation between the acquisition terminal and the server, video frames with low confidence are discarded, which may lead to the omission of road element identification results.

Method used

In the terminal device, the low confidence video frame is fused to improve the confidence and then reported to the server. Combined with the motion state of the carrier and the camera external parameter calibration, noise is filtered to ensure the accuracy of the identification results.

Benefits of technology

It reduces the probability of missing road element recognition results, improves the accuracy and efficiency of element recognition, and reduces the dependence on network transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298939A_ABST
    Figure CN120298939A_ABST
Patent Text Reader

Abstract

The invention provides an element collection method and a related device. The embodiment of the invention can be applied to the traffic field. The method comprises the steps of obtaining a road video; performing element identification on video frames in the road video; according to the identification result, N element frames on the target road section are determined, and the element frames comprise target elements; if the confidence coefficients of the target elements in the N element frames are all lower than a preset value, fusing the target elements in the N element frames to obtain target fusion elements; if the confidence coefficient of the target fusion element is not lower than a preset value, target element information is reported to a server, and the target element information comprises the target element and the position information of the target road section. According to the method provided by the embodiment of the invention, the omission probability of the road element recognition result can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to an element acquisition method and related devices. Background Art

[0002] The acquisition of road elements is one of the important prerequisites for realizing applications such as autonomous driving and intelligent transportation. The current method for acquiring road elements is that the acquisition terminal obtains road videos. First, the acquisition terminal performs simple road element recognition on the video frames, reports the video frames containing road elements to the server, and the server further extracts the road elements in the video frames. The server obtains multiple groups of video frames through the acquisition terminal, and each group of video frames includes different road elements, so as to establish the relationship between each road element and its corresponding position on the road, thereby completing the acquisition process of road elements.

[0003] In practical applications, in order to reduce the transmission cost between the acquisition terminal and the server, the acquisition terminal will discard some video frames with low confidence, and this method may all lead to omissions of the reported road elements.

[0004] To solve the above problems, this application provides an element acquisition method to improve the accuracy and efficiency of road element acquisition. Summary of the Invention

[0005] The embodiments of this application provide an element acquisition method and related devices. By fusing the elements in the video frames with low confidence into elements with high confidence and then reporting them to the server, the omission probability of the recognition results of road elements is reduced.

[0006] The first aspect of this application provides an element acquisition method, including:

[0007] Obtain a road video, where the road video is collected by a shooting device disposed on a vehicle;

[0008] Perform element recognition on the video frames in the road video;

[0009] According to the recognition result, determine N element frames on the target section, where the element frames include target elements, where N is an integer greater than 1, and the element frames belong to the video frames of the road video;

[0010] If the confidence levels of the target elements in the N element frames are all lower than a preset value, then fuse the target elements in the N element frames to obtain a target fusion element;

[0011] If the confidence level of the target fusion element is not lower than the preset value, then report the target element information to the server, where the target element information includes the target element and the position information of the target section.

[0012] In a possible implementation method, after determining N element frames on the target road section according to the recognition result, the method further includes:

[0013] If at least one identification frame is included in the N element frames, report the target element information to the server, where the confidence level of the target element in the identification frame is not lower than the preset value.

[0014] In a possible implementation method, after determining N element frames on the target road section according to the recognition result, the method further includes:

[0015] Perform size recognition on the target elements in the N element frames;

[0016] Determine the element frame corresponding to the target element with the largest size as the recognition frame;

[0017] Establish a mapping relationship between the target element and the recognition frame.

[0018] In a possible implementation method, after reporting the target element in the recognition frame to the server, the method further includes:

[0019] In response to the recognition instruction for the target element issued by the server, send the recognition frame to the server based on the mapping relationship.

[0020] In a possible implementation method, after determining N element frames on the target road section according to the recognition result, the method further includes:

[0021] Perform character recognition on the target elements in the N element frames, and determine the number of characters of the target element in each element frame;

[0022] Determine M element frames with the largest number of characters as target frames, where M is an integer greater than 0 and not greater than N;

[0023] The step of determining the element frame corresponding to the target element with the largest size as the recognition frame specifically includes:

[0024] Determine the target frame corresponding to the target element with the largest size as the recognition frame.

[0025] In a possible implementation method, the element recognition of the video frames in the road video includes:

[0026] Determine the external parameter calibration of the shooting device according to the road video;

[0027] Combine the external parameter calibration to perform element recognition on the video frames.

[0028] In a possible implementation method, after determining N element frames on the target road section according to the recognition result, it further includes:

[0029] Determine the motion state of the vehicle on the target road section;

[0030] Determine the movement trajectory of the target element in the N element frames;

[0031] The reporting of the target element to the server specifically includes:

[0032] If the movement trajectory corresponds to the motion state, report the target element to the server.

[0033] A second aspect of the present application provides an element acquisition device, including:

[0034] A video acquisition module, configured to acquire a road video, where the road video is acquired by a shooting device disposed on a vehicle;

[0035] An element recognition module, configured to perform element recognition on video frames in the road video;

[0036] An element frame determination module, configured to determine N element frames on the target road section according to the recognition result, where the element frames include a target element, and N is an integer greater than 1;

[0037] A fusion module, configured to, if the confidence levels of the target element in the N element frames are all lower than a preset value, fuse the target elements in the N element frames to obtain a target fusion element;

[0038] A reporting module, configured to, if the confidence level of the target fusion element is not lower than the preset value, report the target element information to the server, where the target element information includes the target element and the position information of the target road section.

[0039] In a possible implementation method, it further includes:

[0040] The reporting module is specifically configured to, if at least one identification frame is included in the N element frames, report the target element information to the server, where the confidence level of the target element in the identification frame is not lower than the preset value.

[0041] In this embodiment, after the element acquisition device acquires multiple element frames, it is also necessary to determine the confidence level of the target element in the element frames. If there is at least one element frame among the multiple element frames, and the confidence level of its target element is higher than the preset value, then this at least one element frame is determined as an identification frame. In the recognition result of this identification frame, the confidence level of the target element is higher than the preset value, indicating that it can be determined that there is indeed a corresponding target element on the target road section. Therefore, this target element can be reported.

[0042] In a possible implementation method, it further includes:

[0043] A size recognition module, configured to recognize the size of a target element in N element frames;

[0044] An identification frame determination module, configured to determine the element frame corresponding to the target element with the largest size as the identification frame;

[0045] An establishment module, configured to establish a mapping relationship between the target element and the identification frame.

[0046] In a possible implementation method, it further includes:

[0047] A reporting module, further configured to, in response to an identification instruction for the target element sent by the server, send the identification frame to the server based on the mapping relationship.

[0048] In a possible implementation method, it further includes:

[0049] A character recognition module, configured to recognize characters of the target element in N element frames and determine the number of characters of the target element in each element frame;

[0050] A target frame determination module, configured to determine M element frames with the largest number of characters as target frames, where M is an integer greater than 0 and not greater than N;

[0051] The identification frame determination module is specifically configured to determine the target frame corresponding to the target element with the largest size as the identification frame.

[0052] After determining the target frame, performing size recognition on the target frame can ensure the clarity rate and integrity of the identification frame.

[0053] In a possible implementation method,

[0054] An element recognition module is specifically configured to determine the external parameter calibration of the shooting device according to the road video; and perform element recognition on the video frame in combination with the external parameter calibration.

[0055] In a possible implementation method, it further includes:

[0056] A motion state determination module, configured to determine the motion state of the vehicle on the target section;

[0057] A movement trajectory determination module, configured to determine the movement trajectory of the target element in N element frames;

[0058] The reporting module is specifically configured to report the target element to the server if the movement trajectory corresponds to the motion state.

[0059] The third aspect of the present application provides a computer device, including:

[0060] A memory, a transceiver, a processor, and a bus system;

[0061] Wherein, the memory is used to store programs;

[0062] The processor is used to execute the programs in the memory, including executing the methods in the above aspects;

[0063] The bus system is used to connect the memory and the processor, so that the memory and the processor can communicate.

[0064] The fourth aspect of this application provides a computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the methods in the above aspects.

[0065] The fifth aspect of this application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.

[0066] It can be seen from the above technical solutions that the embodiments of this application have the following advantages:

[0067] This application provides an element acquisition method and related devices. The method is executed by a terminal device and includes obtaining a road video, which is collected by a shooting device arranged on a vehicle; performing element recognition on video frames in the road video; according to the recognition result, determining N element frames on a target section, where the element frames include target elements and the element frames belong to the video frames of the road video; if the confidence levels of the target elements in the N element frames are all lower than a preset value, then fusing the target elements in the N element frames to obtain a target fusion element; if the confidence level of the target fusion element is not lower than the preset value, then reporting target element information to the server, where the target element information includes the target element and the position information of the target section. After the terminal device performs element recognition on the video frames, it will also fuse the elements in the video frames with low confidence levels. If the confidence level of the fused element is high, the element will be reported to the server. Compared with discarding the element information with low confidence levels in the prior art, the method provided by the embodiments of this application can reduce the omission probability of the recognition results of road elements. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a processing flow chart of an acquisition terminal in a road element acquisition method in the related art;

[0069] Figure 2 It is a schematic diagram of the video frame data received by the server per second;

[0070] Figure 3 Schematic diagram for selecting the best recognition frame for the server

[0071] Figure 4 Application environment diagram of the element acquisition method in the embodiment of the present application

[0072] Figure 5 Method flowchart of the element acquisition method provided by the embodiment of the present application

[0073] Figure 6 Schematic diagram for improving confidence by element fusion provided by the embodiment of the present application

[0074] Figure 7 Method flowchart of the element acquisition method provided by the embodiment of the present application

[0075] Figures 8a to 8c Schematic diagram of the movement trajectory of the target element provided by the embodiment of the present application

[0076] Figure 9 Schematic diagram of character recognition provided by the embodiment of the present application

[0077] Figure 10 Processing flowchart of the terminal device provided by the embodiment of the present application

[0078] Figure 11 Schematic diagram of an embodiment of the element acquisition device in the embodiment of the present application

[0079] Figure 12 Schematic diagram of a server structure provided by the embodiment of the present application Detailed implementation manners

[0080] The embodiment of the present application provides an element acquisition method. By fusing elements in video frames with low confidence, if the confidence of the fused elements is high, the elements are reported to the server, thereby reducing the omission probability of the recognition results of road elements.

[0081] In the description, claims, and above-mentioned drawings of the present application, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented, for example, in an order different from those illustrated or described herein. In addition, the terms "comprising", "corresponding to", and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0082] The Intelligent Traffic System (ITS), also known as the Intelligent Transportation System, effectively integrates advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthening the connection among vehicles, roads, and users, thereby forming an integrated transportation system that ensures safety, improves efficiency, protects the environment, and saves energy. Or;

[0083] The Intelligent Vehicle Infrastructure Cooperative Systems (IVICS), abbreviated as the vehicle-road collaborative system, is a development direction of the Intelligent Traffic System (ITS). The vehicle-road collaborative system uses advanced wireless communication and new-generation Internet technologies to comprehensively implement dynamic real-time information interaction between vehicles and between vehicles and roads, and conducts active vehicle safety control and road collaborative management based on the collection and fusion of full-time and full-space dynamic traffic information, fully realizing the effective collaboration of people, vehicles, and roads, ensuring traffic safety, improving traffic efficiency, and thus forming a safe, efficient, and environmentally friendly road traffic system.

[0084] The acquisition of road elements is one of the important prerequisites for realizing applications such as autonomous driving and intelligent transportation. The following is an introduction to the road element acquisition methods in related technologies. Please refer to Figure 1 , Figure 1 is the processing flow chart of the acquisition terminal in the road element acquisition method in related technologies.

[0085] As Figure 1As shown in the figure, the acquisition terminal first initializes the program to prepare for operations such as collecting data, identifying elements, and uploading them to the server. To reduce the transmission cost between the acquisition terminal and the server, the acquisition terminal usually obtains video frames at a low-frequency capture rate of one frame per second (i.e., the capture frequency is 1 Hz). The acquisition terminal performs preliminary image recognition on the captured video frames and selects video frames that include road elements. Among them, road elements can include the following types, such as traffic lights (including red, green, yellow lights, etc.), markings (including lane lines, turning arrows, stop lines, etc.), road signs (including milestone signs, intersection signs, road construction signs, etc.), guardrails (including isolation belts, crash barriers, etc.), traffic signs (including speed limit signs, no-parking signs, etc. indication signs), etc. It can be understood that at this time, the acquisition terminal only judges the general type of road elements. For example, if a blue square appears above the right side of the road in a certain video frame, then this video frame is determined to be a video frame carrying a road sign or a traffic sign. When the acquisition terminal recognizes the above road elements in the captured video frames, it records the recognition information of the type of road elements carried, and at the same time records the positioning information when the video frame of the corresponding road element is captured, and combines the positioning information and the recognition information to obtain video frame data. The acquisition element reports the combined video frame data to the server for subsequent processing.

[0086] The following introduces the processing flow after the server receives the video frame data. As Figure 2 shown, Figure 2 Figure shows a schematic diagram of the video frame data received by the server per second. The information carried by the video frame data can include time information, location information, element information, etc. Among them, the time information can be the system time when the video frame is captured, or the appearance time of the video frame in the road video. In short, the time information is used to represent the shooting time of the corresponding video frame. The location information is used to represent the corresponding location information when the video frame is captured, such as longitude, latitude, global positioning system (GPS) time, etc. The element information is used to represent the type of road element and the coordinate position in the video frame, where the coordinate position can be the range enclosed by at least two vertices ((x1, y1), (x2, y2), etc.) of the road element in the picture coordinate system.

[0087] In a possible case, further, the server will determine the best recognition frame for a certain road element among multiple video frames. Please refer to Figure 3 , Figure 3 Figure shows a schematic diagram of the server selecting the best recognition frame.

[0088] As Figure 3 shown, based on Figure 2For the acquired video frame data, the server will calculate the size of the road element in the video frame based on the two vertex coordinates of the recognition frame of the road element in each video frame data, and compare whether the range defined by this size is the largest. As Figure 3 shown, where the road element is a lane information sign, and the server will select one of the multiple video frames including the lane information sign as the best recognition frame.

[0089] It can be understood that the upper right corners of the first three figures are all lane information signs, and the area they occupy in the video frame gradually increases; although there is also a rectangular sign in the upper right corner of the fourth figure, after the server's recognition and judgment, the rectangular sign in the fourth figure is not a lane information sign but a road sign, that is to say, at this time the lane information sign has gone out of the field of view and there is no lane information sign. Then, the third figure, as the video frame including the lane information sign and with the largest recognition frame size, is determined as the best recognition frame. Finally, the server records the shooting position of the best recognition frame as the position of the corresponding road element, thus completing the collection of road elements, and the collection results can be applied to technologies such as autonomous driving, intelligent transportation, and path navigation.

[0090] In the above process, in order to reduce the transmission cost between the acquisition terminal and the server, limited by the network transmission data volume, the acquisition terminal reports information at a frequency of one frame per second. In many cases, when the moving speed of the acquisition device is large, this frequency cannot obtain the most accurate position; when the acquisition terminal reports, it only reports the recognition element types "exceeding a certain confidence threshold", and the content with low confidence will be discarded, so there may be omissions in the reported road elements.

[0091] To solve the above problems, the present application provides an element acquisition method and related device. After the acquisition terminal performs element recognition on the video frame, it will fuse the elements in the video frames including the same element but with low confidence in the recognition result, so as to improve the confidence of the elements; if the confidence of the fused elements meets the preset requirements, the elements will be reported to the server, thereby reducing the omission probability of the recognition results of road elements and improving the accuracy of the reported elements.

[0092] For easy understanding, please refer to Figure 4 , Figure 4 which is the application environment diagram of the element acquisition method in the embodiment of the present application. As Figure 1As shown in the figure, the element collection method in the embodiments of the present application is applied to an element collection system. The element collection system includes: a server and a terminal device; among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. In the embodiments of the present application, the terminal device is used to collect or obtain the road video collected by the shooting device, and perform image recognition processing on the video frames, which may specifically include but are not limited to the following three categories: driving record devices: collect road element data through cameras and sensors, complete road element recognition, complete data processing and report; high-end car machine devices: (refer to car machine devices with local computing capabilities) can receive camera and various sensor data and perform rough data processing, complete road element recognition, and report road data; professional collection devices: (special equipment for collecting road data as a whole) during the collection process, can collect detailed road element data through various sensors, complete road element recognition, and record it. In the above terminal devices, the process of road element collection and reporting will occur, and the logic of the element collection method provided by the embodiments of the present application is included in both the collection and reporting processes.

[0093] The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and the embodiments of the present application do not limit this.

[0094] Next, the element collection method in the present application will be introduced from the perspective of the terminal device. Please refer to Figure 5 , Figure 5 which is the method flow chart of the element collection method provided by the embodiments of the present application, including:

[0095] 501. Obtain a road video, which is collected by a shooting device arranged on a vehicle carrier.

[0096] It can be understood that to collect elements on the road, it is first necessary to obtain a road video, which is collected by a shooting device on a vehicle carrier.

[0097] The vehicle carrier refers to an object moving on the road, specifically a vehicle. The shooting device can be a driving recorder or a shooting device specifically used for collecting road videos. The shooting device is arranged on the vehicle carrier, and as the vehicle carrier moves, it records the road video on the moving path.

[0098] 502. Identify elements in the video frames of the road video.

[0099] It can be understood that after obtaining the road video, the road video is intercepted into video frames, and then the elements in the video frames can be identified.

[0100] 503. According to the recognition result, determine N element frames on the target section. The element frames include target elements, where N is an integer greater than 1, and the element frames belong to the video frames of the road video.

[0101] After identifying the elements in the video frames of the road video, according to the recognition result, multiple element frames including target elements can be determined, and the section where these multiple element frames are captured is determined as the target section.

[0102] It can be understood that the multiple element frames should be continuous in time sequence. Since in the road video, for a certain road element, the process from its entry into the field of view to its exit from the field of view should be continuous. For example: in the road video, the target element is recognized in the section corresponding to the nth to the (n + 3)th frames, not recognized in the (n + 4)th frame, and recognized again in the section corresponding to the (n + 5)th to the (n + 7)th frames. Then it can be determined that the target element may be blocked by an obstacle (such as plants and animals, other vehicles, pedestrians, etc.) in the (n + 4)th frame and not recognized. That is to say, the element frames including the target element are the nth to the (n + 7)th frames, and at the same time, the section where the nth to the (n + 7)th frames are captured is the target section.

[0103] In this embodiment, the operation of element recognition can be improved to different video frame interception frequencies according to the capabilities of the terminal device. The higher the frequency and the more video frames that can be used for element recognition, the more guaranteed the recognition accuracy.

[0104] 504. If the confidence levels of the target elements in the N element frames are all lower than the preset value, fuse the target elements in the N element frames to obtain a target fused element.

[0105] It can be understood that when the terminal device recognizes road elements, it may occur that the confidence levels of the target elements recognized in an entire section are all lower than the preset value. There are many possibilities for the low confidence of the recognition result, such as the accuracy problem caused by algorithm defects, insufficient information of the content to be recognized itself, etc. In the embodiments of this application, if road elements are recognized in a certain section but the overall confidence level is low, the method of element fusion can be used to improve the confidence level.

[0106] Please refer to Figure 6 , Figure 6Schematic diagram of improving confidence by element fusion provided by an embodiment of this application. During the process of road element recognition, it often occurs that due to occlusion, it is impossible to determine road elements in a single frame. Please refer to Image 601, Image 602, and Image 603 in sequence. The vehicle changes lanes from the right to the left, and the road element type recognized by the terminal device is a lane marking.

[0107] In Image 601, since the vehicle occludes the right half of the lane marking, the recognition result of the lane marking by the terminal device may be: "Left-turn lane" with low confidence, or "Lane for straight and left turns" with low confidence.

[0108] In Image 602, the vehicle occludes the upper half of the lane marking. Therefore, the recognition result of the lane marking by the terminal device may be: "Left-turn lane" with low confidence, "Lane for straight and left turns" with low confidence, "Straight lane" with low confidence, "Lane for straight and right turns" with low confidence, "Right-turn lane" with low confidence, "U-turn lane" with low confidence, etc.

[0109] In Image 603, the vehicle occludes the left half of the lane marking. Therefore, the recognition result of the lane marking by the terminal device may be: "Straight lane" with low confidence, "Lane for straight and left turns" with low confidence.

[0110] The terminal device combines the recognition results in Image 601, Image 602, and Image 603, and finds that all include the recognition information of "Lane for straight and left turns" with low confidence. Then it can be judged that the reason for the low-confidence recognition result of the video frame may be due to occlusion. Therefore, the terminal device fuses the lane markings in Image 601, Image 602, and Image 603 to obtain a target fusion element, that is, Image 604.

[0111] It can be understood that the terminal device can also combine scene data to judge whether the reason for the low confidence is occlusion. The scene data includes scene elements in the current video frame and the moving state of the vehicle carrier. Taking Image 601, Image 602, and Image 603 as examples, the scene data includes: "Vehicle information is recognized in the current video frame", "The current movement condition is slow speed". Among them, the moving state of the vehicle carrier can be obtained by acquiring the vehicle speed.

[0112] 505. If the confidence of the target fusion element is not lower than the preset value, report the target element information to the server. The target element information includes the target element and the position information of the target road section.

[0113] It can be understood that after fusing the target elements in the N element frames to obtain the target fusion element, if the confidence level of the target fusion element is increased to a medium-high confidence level, that is, exceeding the preset threshold, it means that the target element can be reported to the server. The data reported to the server is the target element information, including the position information of the target element and the target road section. For example, if the road section from the nth frame to the (n + 7)th frame is the target road section, then when the terminal device reports the target element, it will also report the position information when shooting the nth frame to the (n + 7)th frame.

[0114] In the embodiments of the present application, after the terminal device performs element recognition on the video frames, it will also fuse the elements in the video frames with low confidence levels. If the confidence level of the fused element is high, the element will be reported to the server. Compared with the prior art of discarding the element information with low confidence levels, the method provided by the embodiments of the present application can reduce the omission probability of the road element recognition results.

[0115] In the Figure 5 corresponding embodiment of the element acquisition method provided by the present application, in an alternative embodiment, please refer to Figure 7 , Figure 7 which is the method flowchart of the element acquisition method provided by the embodiments of the present application, including:

[0116] 701. Obtain a road video, which is collected by a shooting device arranged on a vehicle.

[0117] It can be understood that step 701 is similar to step 501 in the above Figure 5 corresponding embodiment, and will not be elaborated here.

[0118] 702. Determine the external parameter calibration of the shooting device according to the road video;

[0119] 703. Combine the external parameter calibration to perform element recognition on the video frames.

[0120] It can be understood that after obtaining the road video, to perform element recognition on the video frames in the road video, it can be achieved by combining the external parameter calibration of the shooting device. The external parameter calibration of the shooting device includes the calibration of the pitch angle, yaw angle, and roll angle of the shooting device. The calibration methods include pre-static calibration based on a special target with geometric features, dynamic self-calibration based on existing features such as the vanishing point of traffic signs and parallel lines, etc. Those skilled in the art can select the calibration algorithm according to the actual situation and experience, and the present application will not elaborate here. After determining the external parameter calibration and then performing element recognition, the accuracy of the recognition result can be improved.

[0121] 704. According to the recognition result, determine N element frames on the target road section, where the element frames include target elements, N is an integer greater than 1, and the element frames belong to the video frames of the road video.

[0122] It is understandable that step 704 is similar to the above Figure 5 Step 503 in the corresponding embodiment is similar and will not be described again here.

[0123] 705 , determining whether the N element frames include at least one identification frame, wherein the confidence level of the target element in the identification frame is not less than a preset value.

[0124] It is understandable that after acquiring multiple element frames, it is also necessary to determine the confidence of the target element in the element frame. If there is at least one element frame among the multiple element frames whose confidence of the target element is higher than a preset value, then the at least one element frame is determined as an identification frame.

[0125] If the judgment result of step 705 is "no", the process proceeds to step 706 to fuse the target elements in the N element frames to obtain a target fused element.

[0126] It is understandable that step 706 is similar to the above Figure 5 Step 504 in the corresponding embodiment is similar and will not be described again here.

[0127] 707, judging whether the confidence of the target fusion element is not less than a preset value.

[0128] It can be understood that after the target elements in the N element frames are fused to obtain the target fused element, the target fused element is identified again to obtain the corresponding confidence level to determine whether the fused target element meets the reporting requirement.

[0129] If the judgment result of step 707 is "No", the process proceeds to step 712 and the target element is not reported to the server.

[0130] It is understandable that if the fused target element still has low confidence, it means that the reason for the low confidence of the target element may be misrecognition caused by algorithm defects, rather than insufficient recognition content information caused by occlusion. Therefore, the target element is not reported to the server.

[0131] If the judgment result of step 705 is "yes", or the judgment result of step 707 is "yes", then go to step 708 to determine the movement state of the vehicle on the target road section.

[0132] If the judgment result of step 705 is "yes", it means that there is an identification frame, and in the recognition result of the identification frame, the confidence of the target element is higher than the preset value, which means that it can be determined that the corresponding target element does exist on the target section.

[0133] If the judgment result of step 707 is "yes", it also means that the corresponding target element does exist on the target road section.

[0134] It is understandable that the methods for the terminal device to determine the motion state of the vehicle on the target road section include: directly obtaining the speed information and steering information of the vehicle to obtain the motion state of the vehicle; analyzing the video segment corresponding to the target road section to obtain the motion state of the vehicle.

[0135] 709. Determine the movement trajectory of the target element in N element frames;

[0136] It is understandable that by combining the positions of the target element in N element frames, the movement trajectory of the target element frame in the video picture can be determined.

[0137] 710. Determine whether the movement trajectory corresponds to the motion state. If so, proceed to step 711 and report the target element information to the server. The target element information includes the position information of the target element and the target road section. If not, proceed to step 712.

[0138] It is understandable that after obtaining the movement trajectory of the target element and the motion state of the vehicle, according to the motion state of the vehicle, it can be determined whether the movement trajectory of the target element conforms to the visual field change law in this motion state. For details, please refer to Figures 8a to 8c .

[0139] In the scenario of real driving, during the vehicle driving process, the road elements on both sides of the visual field will continuously change from small to large and then out of the visual field as the distance gets closer. For example, Figure 8a As shown, road elements A and B appear in three consecutive video frames. As the distance gets closer, the range of the elements in the picture becomes larger and larger until part of the road elements goes out of the visual field.

[0140] If the recognized movement trajectory of the element does not conform to the motion state, for example, when the vehicle is moving normally, if an element does not move according to the normal situation or remains stationary all the time, then this element is probably a noise point. For example, Figure 8b As shown, road element A normally changes from small to large until it goes out of the visual field, while road element B remains unchanged. Then road element B is probably a noise point.

[0141] In a possible implementation method, to determine whether the movement trajectory corresponds to the motion state, the external parameter calibration and installation position of the shooting device can also be combined. For example, when there is a certain offset between the shooting picture of the shooting device and the moving direction of the vehicle, this offset angle can be combined to determine whether it conforms to the visual field change law. For example, Figure 8c As shown, when the shooting angle of the shooting device deviates to the right, the left road element A moves normally to the left, but the right road element B moves to the upper left, which does not conform to the visual field change law. Therefore, road element B is probably a noise point.

[0142] If the movement trajectory of the target element conforms to the motion state of the vehicle carrier, it indicates that the target element is not noise. Therefore, the terminal device can report the target element information to the server. Among them, the target element information includes the position information of the target element and the target road section. For specific details, please refer to Figure 5 the relevant description of step 505 in the corresponding embodiment, which will not be elaborated here.

[0143] In this embodiment, the movement trajectory of the target element is analyzed by combining the motion state of the vehicle carrier to eliminate noise. Compared with the prior art of reporting target elements with high confidence without discrimination, the method provided in this embodiment can reduce the transmission cost between the terminal device and the server.

[0144] 713. Identify the size of the target element in N element frames;

[0145] 714. Determine the element frame corresponding to the target element with the largest size as the recognition frame.

[0146] 715. Establish a mapping relationship between the target element and the recognition frame.

[0147] It can be understood that after determining that the target element in the N element frames is a valid element, the terminal device can also select the best recognition frame on behalf of the server.

[0148] The terminal device's identification of the size of the target element in the N element frames refers to determining the range where the target element is located in the element frame. For example, determining the respective vertex coordinates of the target element in each element frame, and taking the area framed according to the respective vertex coordinates as the size of the target element. Determine the element frame corresponding to the target element with the largest size as the recognition frame.

[0149] It can be understood that the method for the terminal device to determine the recognition frame is similar to the method for the server to determine the best recognition frame in the related art. In the embodiment of the present application, the execution operation of this method is changed to be processed in the terminal device. Completing the determination of the best recognition frame in the terminal device does not need to consider the transmission cost of uploading video frame data. Therefore, it can be unrestricted by the limitation of recognizing data at one frame per second in the conventional way. The terminal device can continuously increase the video capture frequency according to its own computing power, so as to identify the best recognition point closer to the actual road conditions.

[0150] After the terminal device determines the recognition frame, it can save the recognition frame and establish a mapping relationship between the target element and the recognition frame. That is to say, the data actively reported by the terminal device is only the target element information, and the target element information can be in the form of vector data rather than pictures. When the server needs to confirm the target element, it can send an identification instruction for the target element to the terminal device based on the target element information, for instructing the terminal device to send a picture carrying the target element.

[0151] In a possible implementation manner, before step 713, it further includes:

[0152] 716. Perform character recognition on the target element in the N element frames, and determine the number of characters of the target element in each element frame;

[0153] 717. Determine the M element frames with the most characters as the target frames, where M is an integer greater than 0 and not greater than N.

[0154] It can be understood that among the road elements on the real road, many contain text information. During the process from far to near, the number of text information that can be seen and clearly seen by the field of vision also changes from less to more; however, the number will become less and less during the process of the text information leaving the field of vision. Therefore, through the optical character recognition (OCR) technology, character recognition can be performed on the target element. The more characters are recognized, the clearer the element frame is. Determine the element frame with the most characters as the target frame.

[0155] For example, when the element frames including the target element are from the nth frame to the n + 7th frame, and through the character recognition result, the number of characters in the target elements of the n + 5th and n + 6th frames is both m, while the number of characters in the target elements of other frames is less than m, then both the n + 5th and n + 6th frames are target frames.

[0156] Step 714 specifically includes: determining the target frame corresponding to the target element with the largest size as the recognition frame.

[0157] It can be understood that after determining the target frame, then perform size recognition on the target frame, and determine the target frame corresponding to the target element with the largest size as the recognition frame, which can ensure the clarity rate and integrity of the recognition frame.

[0158] As Figure 9 shown Figure 9A schematic diagram of character recognition provided by an embodiment of the present application. For the road elements in the figure, the change process of the OCR recognition content is AB → ABC → ABCD → ABC. During the continuous change of the recognition frame, the video frame with the most recognized content is the one including ABCD. Even if the subsequent recognition frame is larger than this video frame, since the content is not as complete as the video frame with ABCD, it cannot be used as the best recognition frame. Instead, the frame with ABCD is used as the best recognition frame.

[0159] 718. In response to the recognition instruction for the target element sent by the server, send the recognition frame to the server based on the mapping relationship.

[0160] It can be understood that since the terminal device has already determined the best recognition frame in the video frame and established the mapping relationship between the target element and the recognition frame, the terminal device can, based on this mapping relationship, find the recognition frame at the corresponding storage location and send it to the server for verification.

[0161] The element acquisition method provided by the embodiment of the present application enables the recognition frequency to be independent of network transmission data by recognizing road elements on the terminal device. The terminal device can increase the recognition frequency according to its own computing power; the terminal device can also utilize the element recognition results with low confidence to fuse multiple frames of low-confidence data into high-confidence data, reducing the probability of missing the recognition results of road elements; the terminal device can also directly obtain the external calibration of the camera and the motion state of the vehicle carrier, and determine the correct element movement law in combination with the relationship of the field of view, thereby filtering out a lot of noise; the terminal device can also obtain the text quantity of the current element through OCR recognition, and characterize whether the current element has exited the field of view through the change of the text quantity, so as to better determine the recognition frame.

[0162] Please refer to Figure 10 , Figure 10 It is a processing flow chart of the terminal device provided by the embodiment of the present application.

[0163] Based on the element acquisition method provided by the embodiment of the present application, the processing flow of the terminal device is as follows:

[0164] 1) After completing the program initialization, perform the calculation of the external calibration of the camera. The calculation method can be simplified to the calibration in a simple scenario: for example, when the vehicle is moving forward, by calculating the change direction of the lane line, the vertical and horizontal angles between the current camera and the horizontal vehicle forward direction can be obtained.

[0165] 2) Enter the judgment and repetition process of "acquire video frame, image recognition, OCR recognition, update recognition point, exit the field of view", and the process is as follows:

[0166] Frequency increase to obtain video frames: This operation can be improved to different frequencies according to the device capabilities. The higher the frequency, the higher the accuracy of the best recognition point;

[0167] Image recognition: Complete the recognition of road elements in the current video frame. The recognition content includes: element type, confidence level, position and size;

[0168] OCR recognition: Recognize the text within the framed area of the element and record it in the OCR recognition results of each road element;

[0169] Best recognition frame update: With the help of the external camera calibration, for the multiple "image recognition results and OCR recognition results" of each road element, combine the motion trajectory to complete the judgment and denoising of road element recognition, and update the position of the best recognition frame of the current element.

[0170] Out-of-view judgment: Obtain the road elements that have not been updated in the database (DB), and judge that the elements that have not been updated for more than a certain number of consecutive frames (set threshold) indicate that the current element has gone out of view and enter step 3); other elements are not processed.

[0171] For each road element that has gone out of view, combine the "positioning data and recognition information" of the entire process from the appearance to the disappearance of the road element. The data is sorted in time series in groups of one second, and the combined information is reported to the server; combine the best recognition point information of the road element and upload the best recognition point after responding to the element acquisition instruction of the server.

[0172] The following will describe the element acquisition device in the present application in detail. Please refer to Figure 11 . Figure 11 is a schematic diagram of an embodiment of the element acquisition device 1100 in the embodiment of the present application. The element acquisition device 1100 includes:

[0173] A video acquisition module 1101, configured to acquire a road video, and the road video is acquired by a shooting device disposed on a vehicle;

[0174] An element recognition module 1102, configured to recognize elements in the video frames of the road video;

[0175] An element frame determination module 1103, configured to determine N element frames on a target road section according to the recognition results. The element frames include target elements, where N is an integer greater than 1;

[0176] A fusion module 1104, configured to, if the confidence levels of the target elements in the N element frames are all lower than a preset value, fuse the target elements in the N element frames to obtain a target fusion element;

[0177] A reporting module 1105, configured to report target element information to a server if the confidence level of a target fusion element is not lower than a preset value, where the target element information includes the target element and the location information of the target road section.

[0178] After the element collection device provided by the embodiment of the present application recognizes elements in a video frame, it will also fuse the elements in the video frame with low confidence levels. If the confidence level of the fused element is high, the element will be reported to the server. Compared with the prior art of discarding element information with low confidence levels, the method provided by the embodiment of the present application can reduce the omission probability of the recognition results of road elements.

[0179] In a possible implementation method, it further includes:

[0180] The reporting module 1105 is specifically configured to report target element information to the server if at least one identification frame is included in N element frames, where the confidence level of the target element in the identification frame is not lower than a preset value.

[0181] In this embodiment, after the element collection device obtains multiple element frames, it is also necessary to determine the confidence level of the target element in the element frames. If there is at least one element frame among the multiple element frames whose confidence level of the target element is higher than the preset value, then the at least one element frame is determined as an identification frame. In the recognition result of this identification frame, the confidence level of the target element is higher than the preset value, indicating that it can be determined that the corresponding target element actually exists on the target road section. Therefore, the target element can be reported.

[0182] In a possible implementation method, it further includes:

[0183] A size recognition module 1106, configured to recognize the size of the target element in N element frames;

[0184] An identification frame determination module 1107, configured to determine the element frame corresponding to the target element with the largest size as the identification frame;

[0185] An establishment module 1108, configured to establish a mapping relationship between the target element and the identification frame.

[0186] In this embodiment, the data actively reported by the element collection device is only the target element information, and the target element information can be in the form of vector data rather than pictures. When the server needs to confirm the target element, it can issue an identification instruction for the target element to the element collection device based on the target element information, for instructing the element collection device to send a picture carrying the target element. The selection of the picture can be determined by the size of the element frame.

[0187] In a possible implementation method, it further includes:

[0188] The reporting module 1105 is further configured to send an identification frame to the server based on the mapping relationship in response to an identification instruction for a target element issued by the server.

[0189] In this embodiment, since the element acquisition device has determined the optimal identification frame in the video frame and established the mapping relationship between the target element and the identification frame, the element acquisition device can find the identification frame at the corresponding storage location based on this mapping relationship and send it to the server for verification.

[0190] In a possible implementation method, it further includes:

[0191] The character recognition module 1109 is configured to perform character recognition on the target element in the N element frames to determine the number of characters of the target element in each element frame;

[0192] The target frame determination module 1110 is configured to determine the M element frames with the most characters as the target frames, where M is an integer greater than 0 and not greater than N;

[0193] The identification frame determination module 1107 is specifically configured to determine the target frame corresponding to the target element with the largest size as the identification frame.

[0194] In this embodiment, since among the road elements in the real road, many contain text information. When the text information approaches from far to near, the number that can be seen and clearly seen in the field of view also changes from less to more; however, the number will become less and less when the text information exits the field of view. Therefore, the character recognition technology can perform character recognition on the target element. The more characters are recognized, the clearer the element frame is. The element frame with the most characters is determined as the target frame.

[0195] After determining the target frame, performing size recognition on the target frame can ensure the clarity rate and integrity of the identification frame.

[0196] In a possible implementation method, the element recognition module 1102 is specifically configured to determine the external parameter calibration of the shooting device according to the road video; and perform element recognition on the video frame in combination with the external parameter calibration.

[0197] In this embodiment, after obtaining the road video, performing element recognition on the video frames in the road video can be achieved by combining the external parameter calibration of the shooting device. Performing element recognition after determining the external parameter calibration can improve the accuracy of the recognition result.

[0198] In a possible implementation method, it further includes:

[0199] The motion state determination module 1111 is configured to determine the motion state of the vehicle on the target section;

[0200] The movement trajectory determination module 1112 determines the movement trajectory of the target element in N element frames;

[0201] The reporting module 1105 is specifically configured to report the target element to the server if the movement trajectory corresponds to the motion state.

[0202] In this embodiment, by determining the movement trajectory of the target element and the motion state of the vehicle carrier, it is possible to judge whether the movement trajectory of the target element conforms to the visual field change law in this motion state according to the motion state of the vehicle carrier, thereby filtering out a lot of noise.

[0203] Figure 12 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 300 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 322 (for example, one or more processors) and a memory 332, and one or more storage media 330 for storing application programs 342 or data 344 (for example, one or more mass storage devices). Among them, the memory 332 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 322 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the server 300.

[0204] The server 300 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.

[0205] The server in the above embodiment may be based on the Figure 12 server structure shown.

[0206] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.

[0207] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0208] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0209] In addition, each functional unit in various embodiments of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0211] The above, the above embodiments are only used to illustrate the technical solution of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of this application.

Claims

1. An element collection method, characterized in that, Including: Obtain a road video, which is collected by a shooting device arranged on a vehicle; Perform element recognition on the video frames in the road video; According to the recognition result, determine N element frames on the target section, where the element frames include target elements, N is an integer greater than 1, and the element frames belong to the video frames of the road video; If the confidence levels of the target elements in the N element frames are all lower than the preset value, then fuse the target elements in the N element frames to obtain a target fusion element; If the confidence level of the target fusion element is not lower than the preset value, report the target element information to the server, where the target element information includes the target element and the position information of the target section.

2. The method according to claim 1, wherein After determining the N element frames on the target section according to the recognition result, it further includes: If at least one identification frame is included in the N element frames, report the target element information to the server, where the confidence level of the target element in the identification frame is not lower than the preset value.

3. The method according to claim 1 or 2, wherein After determining the N element frames on the target section according to the recognition result, it further includes: Perform size recognition on the target elements in the N element frames; Determine the element frame corresponding to the target element with the largest size as the recognition frame; Establish a mapping relationship between the target element and the recognition frame.

4. The method according to claim 3, wherein After reporting the target element in the recognition frame to the server, it further includes: In response to the recognition instruction for the target element issued by the server, send the recognition frame to the server based on the mapping relationship.

5. The method according to claim 3, wherein After determining the N element frames on the target section according to the recognition result, it further includes: Perform character recognition on the target elements in the N element frames, and determine the number of characters of the target element in each element frame; Determine M element frames with the largest number of characters as target frames, where M is an integer greater than 0 and not greater than N; The step of determining the element frame corresponding to the target element with the largest size as the recognition frame specifically includes: Determine the target frame corresponding to the target element with the largest size as the recognition frame.

6. The method according to claim 1, wherein The step of performing element recognition on the video frames in the road video includes: Determine the external parameter calibration of the shooting device according to the road video; Combined with the external parameter calibration, perform element recognition on the video frames.

7. The method according to claim 1, wherein After determining the N element frames on the target section according to the recognition result, it further includes: Determine the motion state of the vehicle on the target section; Determine the movement trajectory of the target element in the N element frames; The step of reporting the target element to the server specifically includes: If the movement trajectory corresponds to the motion state, report the target element to the server.

8. An element collection device, characterized in that, Including: A video acquisition module for acquiring a road video, which is collected by a shooting device arranged on a vehicle; An element recognition module for performing element recognition on the video frames in the road video; An element frame determination module, configured to determine N element frames on a target road section according to an identification result, where the element frames include target elements, and N is an integer greater than 1; A fusion module, configured to fuse the target elements in the N element frames to obtain a target fusion element if the confidence levels of the target elements in the N element frames are all lower than a preset value; A reporting module, configured to report the target element information to a server if the confidence level of the target fusion element is not lower than the preset value, where the target element information includes the target element and the location information of the target road section.

9. A computer device, characterized in that, Comprising: A memory, a transceiver, a processor, and a bus system; Wherein, the memory is used to store programs; The processor is configured to execute the programs in the memory, including executing the element acquisition method according to any one of claims 1 to 7; The bus system is used to connect the memory and the processor to enable communication between the memory and the processor.

10. A computer-readable storage medium, characterized in that, Comprising instructions that, when running on a computer, cause the computer to execute the element acquisition method according to any one of claims 1 to 7.