Method, apparatus, storage medium, and electronic device for determining recognition result
By storing and fusing multi-frame image features in license plate recognition, the problem of inaccurate single-frame recognition results is solved, and the accuracy and stability of license plate recognition is improved, especially in complex traffic scenarios.
Patent Information
- Application Number
- CN202210259167.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-16
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-03-16
AI Technical Summary
In the prior art, the license plate recognition results are inaccurate, especially in complex traffic scenarios, single frame recognition results are prone to errors, and the license plate voting method may lead to wrong voting results when the fuzzy frame exists for a long time.
By identifying the target object in the image, obtaining the recognition results and their confidence, and storing them in the target queue under the predetermined conditions, fusing the image features of all the recognition results in the queue to determine the final recognition results, and filtering high-quality historical frame features using CNN network and tracking algorithm for fusing.
It improves the accuracy of license plate recognition results, reduces the impact of single frame recognition errors in complex traffic scenarios, and enhances the stability and accuracy of recognition results.
Smart Images

Figure CN114580563B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing, and in particular, to a method, apparatus, storage medium, and electronic device for determining an identification result. Background Art
[0002] After an image is acquired, it is often necessary to identify the image to determine the identification result. The following takes the license plate image as an example for illustration:
[0003] License plate recognition technology has been widely applied in various fields of traffic scenarios, such as road monitoring, illegal speed measurement and capture, electronic police discrimination of violations, and so on. Due to the extremely complex traffic scenarios and many interference factors, such as lights, noise, weather, pedestrians, vehicles, buildings, advertisements, etc., it is extremely inaccurate to rely on the license plate recognition result of a single frame to capture and penalize vehicles.
[0004] In related technologies, a license plate voting strategy is mostly adopted to ensure the correctness of the image recognition result in the captured frame. The license plate voting method is a method for processing the license plate recognition results of multiple frames to ensure that there is one frame in the output results. However, through the license plate voting rules of string comparison, when the existence time of the fuzzy frame is long, the license plate recognition results in the queue may all be incorrect. This will also cause the final voting result to be incorrect.
[0005] It can be seen that there is a problem of inaccurate determination of the recognition result in related technologies.
[0006] In view of the above problems existing in related technologies, no effective solution has been proposed yet. Summary of the Invention
[0007] Embodiments of the present invention provide a method, apparatus, storage medium, and electronic device for determining an identification result, so as to at least solve the problem of inaccurate determination of the identification result existing in related technologies.
[0008] According to an embodiment of the present invention, a method for determining an identification result is provided, including: identifying a target object included in a first image to obtain a first identification result and a first confidence level of the first identification result; in response to the first identification result and the first confidence level satisfying a predetermined condition, storing the first identification result in a target queue, where the target queue is further used to store a second identification result having the same identifier as the target object, the second identification result being a result obtained by identifying an object included in a second image, and the second image being an image in an image sequence where the first image is located and where an object having the same identifier as the target object is located; fusing image features corresponding to all the identification results included in the target queue to obtain a fused feature; and determining a target identification result corresponding to the target queue based on the fused feature.
[0009] According to another embodiment of the present invention, there is provided a device for determining an identification result, including: an identification module, configured to identify a target object included in a first image, to obtain a first identification result and a first confidence level of the first identification result; a storage module, configured to, in response to the first identification result and the first confidence level satisfying a predetermined condition, store the first identification result into a target queue, wherein the target queue is further configured to store a second identification result having the same identifier as that of the target object, the second identification result being a result obtained by identifying an object included in a second image, and the second image being an image in an image sequence where the first image is located and where an object having the same identifier as that of the target object is located; a fusion module, configured to fuse image features corresponding to all the identification results included in the target queue, to obtain a fusion feature; and a determination module, configured to determine a target identification result corresponding to the target queue based on the fusion feature.
[0010] According to still another embodiment of the present invention, there is further provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0011] According to still another embodiment of the present invention, there is further provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0012] By means of the present invention, an object included in a first image is identified to obtain a first identification result and a first confidence level of the first identification result. When the first identification result and the first confidence level satisfy a predetermined condition, the first identification result is stored into a target queue, where the target queue is further configured to store a second identification result, the second identification result being a result obtained by identifying an object included in a second image, and the second image being an image in an image sequence where the first image is located and where an object having the same identifier as that of the target object is located. Image features corresponding to all the identification results included in the target queue are fused to obtain a fusion feature, and a target identification result corresponding to the target queue is determined according to the fusion feature. Since when determining the target identification result, the fusion feature obtained by fusing images corresponding to all the identification results in the target queue is identified, and the identification results included in the target queue are all results satisfying the predetermined condition, ensuring the accuracy of the identification results in the target queue, thus, the problem of inaccurate determination of the identification result in the related art can be solved, achieving the effect of improving the accuracy rate of determining the identification result. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1This is a hardware structure block diagram of a mobile terminal for a method for determining a recognition result according to an embodiment of the present invention;
[0014] Figure 2 is a flowchart of a method for determining a recognition result according to an embodiment of the present invention;
[0015] Figure 3 is a flow chart of adding a first recognition result to a target queue according to an exemplary embodiment of the present invention;
[0016] Figure 4 is a flow chart of determining an identification result of a removal queue according to an exemplary embodiment of the present invention;
[0017] Figure 5 is a schematic diagram of determining a target recognition result based on fusion features according to an exemplary embodiment of the present invention;
[0018] Figure 6 is a flow chart of a method for determining a recognition result according to a specific embodiment of the present invention;
[0019] Figure 7 4 is a structural block diagram of a device for determining a recognition result according to an embodiment of the present invention. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0022] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a mobile terminal for determining a recognition result according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0023] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the recognition result in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0025] In this embodiment, a method for determining a recognition result is provided. Figure 2 is a flowchart of the method for determining the recognition result according to the embodiment of the present invention, as Figure 2 shown, the process includes the following steps:
[0026] Step S202, identify the target object included in the first image, and obtain the first recognition result and the first confidence level of the first recognition result;
[0027] Step S204, in response to the first recognition result and the first confidence level meeting a predetermined condition, store the first recognition result into a target queue, where the target queue is also used to store a second recognition result having the same identifier as the target object, and the second recognition result is the result obtained by identifying the object included in the second image, and the second image is the image where the object having the same identifier as the target object is located in the image sequence where the first image is located;
[0028] Step S206, fuse the image features corresponding to all the recognition results included in the target queue to obtain a fused feature;
[0029] Step S208, determine the target recognition result corresponding to the target queue based on the fusion feature.
[0030] In the above embodiment, the first image may be an image in an image sequence, and the image sequence may be a sequence intercepted from a video stream. For example, in a video stream, a sub-video with a predetermined duration is taken, and the sequence composed of each frame image in the sub-video is determined as the image sequence. The first image may include a target object, and the target object may be a person, a vehicle, an item, etc. The video stream may be a video stream collected by a monitoring device, such as a road traffic video stream, a shopping mall monitoring video stream, etc.
[0031] In the above embodiment, each frame in the image sequence may be input into the value network model respectively, that is, the target object included in the first image may be recognized through the network model to obtain the first recognition result and the first confidence of the first recognition result. Among them, when the target object is a vehicle, the first recognition result may be information such as a license plate number and a vehicle model. When the target object is a person, the first recognition result may be the identity information of the person, such as an ID number, or other identification information, such as an employee number, etc. The first confidence may represent the accuracy rate or the probability of correct recognition of the first recognition result. Among them, the network model may be a CNN convolutional neural network model, and the CNN network is used to extract features from the input image, and convolutional layers, normalization layers, activation layers, pooling layers, etc., will be used. Finally, a high-dimensional abstract semantic feature will be obtained, and this feature needs to be saved and will be used by the subsequent voting network. Using this feature, the recognition result of this frame can be obtained through the recognition network head. Then, it is sent to the tracking algorithm to obtain the identification of the target object, marking each different license plate.
[0032] In the above embodiment, after the first recognition result is determined, the identification of the target object indicated by the first recognition result may be determined. The identification of the target object may be an identification assigned to the target object. When determining the identification, the recognition results of other images before the first image in the image sequence may be determined, and it may be determined whether there is a recognition result identical to the first recognition result, that is, it may be determined whether there is an object identical to the target object in the images before the first image. When it exists, the identification of the object identical to the target object is determined as the identification of the target object. When it does not exist, a unique identification may be assigned to the target object. Among them, the identification of each target object may be arranged in order. For example, for the first frame image in the image sequence, the object identification in the first frame image may be determined as 001. For the second frame image, if the recognition result of its object is the same as that of the object in the first frame image, the identification of the object in the second frame image is determined as 001. If the recognition result of the object in the second frame image is different from that of the object in the first frame image, the identification of the object in the second frame image may be determined as 002.
[0033] In the above embodiments, since it is necessary to vote on the results of multiple frames, a cache needs to be established to save the useful information in the historical frames, which is called a queue. What needs to be saved includes the identification information, the feature map of the image, the confidence of the recognition result, and the identifier of the object output by the tracking algorithm, such as the license plate ID and the number of frames in which the license plate with the current ID appears in the video, that is, the image sequence number. Among them, each license plate, that is, the license plate with each ID, has its own voting queue, and the license plate information saved in the queue is considered to be the same license plate retained in different frames. The license plate queue has a capacity limit and is not infinite. Therefore, a rule for enqueueing and dequeueing needs to be established to maintain the current queue. After the first recognition result is recognized, the first recognition result can be stored in the target queue. Among them, the target queue is used to store the recognition results of objects with the same identifier. That is, the identifiers of the recognition results stored in each target queue are the same. Before storing the first recognition result in the target queue, it can be first determined whether the first recognition result and the first confidence meet the preset conditions. In the case of meeting the conditions, the first recognition result is stored in the target queue.
[0034] In the above embodiments, when there is a queue in the existing queues whose stored object identifier is the same as that of the target object, that queue is determined as the target queue. When there is no such queue, a target queue is created, and the identifier of the target queue is determined as the identifier of the target object.
[0035] In the above embodiments, after each frame image included in the image sequence is recognized, the image features corresponding to all the recognition results in the target queue can be fused to obtain a fused feature, and the fused feature is recognized to determine the target recognition result.
[0036] Optionally, the execution subject of the above steps may be a background processor, or other devices with similar processing capabilities, or may also be a machine integrated with at least a data processing device. Among them, the data processing device may include terminals such as computers and mobile phones, but is not limited thereto.
[0037] Through the present invention, an object included in a first image is recognized to obtain a first recognition result and a first confidence level of the first recognition result. When the first recognition result and the first confidence level meet a predetermined condition, the first recognition result is stored in a target queue, where the target queue is further used to store a second recognition result, where the second recognition result is a result obtained by recognizing an object included in a second image, and the second image is an image of an object with the same identifier as the target object in the image sequence where the first image is located. Image features corresponding to all recognition results included in the target queue are fused to obtain a fused feature, and a target recognition result corresponding to the target queue is determined according to the fused feature. Since when determining the target recognition result, the fused feature obtained after fusing images corresponding to all recognition results in the target queue is recognized, and the recognition results included in the target queue are all results that meet the predetermined condition, the accuracy of the recognition results in the target queue is ensured. Therefore, the problem of inaccurate determination of recognition results in the related art can be solved, and the effect of improving the accuracy rate of determining recognition results can be achieved.
[0038] In an exemplary embodiment, in response to the first recognition result and the first confidence level meeting a predetermined condition, storing the first recognition result in the target queue includes: in response to the target queue being an existing queue, determining a first sequence number of a third image corresponding to a third recognition result that was most recently stored in the target queue, where the first sequence number is the sequence number of the third image in the image sequence; determining a first difference between the first sequence number and a second sequence number of the first image in the image sequence; in response to the first difference being greater than a preset difference, determining whether the first confidence level of the first recognition result is greater than a preset confidence level; and in response to the first confidence level being greater than the preset confidence level, storing the first recognition result in the target queue. In this embodiment, when the target queue is an existing target queue, before storing the first recognition in the target queue, a first sequence number of a third image corresponding to a third recognition result that was most recently stored in the target queue in the image sequence can be determined, a second sequence number of the first image in the image sequence can be determined, a first difference between the second sequence number and the first sequence number can be determined, when the first difference is greater than the preset difference, it can be determined whether the first confidence level is greater than the preset confidence level, and when the first confidence level is greater than the preset confidence level, the first recognition result is stored in the target queue. Wherein, the image sequence includes a plurality of consecutive frames of images, and then the images can be numbered according to the order of the images in the image sequence to obtain the sequence number of each frame of image. For example, the sequence number of the first frame of image can be 1, the sequence number of the second frame of image can be 2... That is, the first sequence number and the second sequence number are the numbers of the images in the image sequence.
[0039] In the above embodiments, it is determined which queue the license plate ID of the current frame should be added to or a new queue should be created. If it should be added to a certain queue, it is further determined whether the difference between the sequence number of the current frame and the most recent sequence number in the queue is greater than a set threshold. If the frame difference is very small, it means that the moving position of the license plate is very small, and there will be no significant changes in the recognition result and the license plate features. If it is continuously enqueued, the license plate information in the queue will only be that of the previous few frames, and the final voting result will only refer to the information of the previous few frames. Therefore, when the frame difference is less than the preset difference, the first recognition result is discarded. If it is greater than the preset difference, it is further determined whether the recognition confidence of this frame is greater than the set confidence threshold. If it is greater, it is added to the queue; otherwise, the current frame is discarded. Among them, the flowchart of adding the first recognition result to the target queue can be seen in the appendix Figure 3 。
[0040] In the above embodiments, when the identifiers of the objects stored in the existing queues are all different from the identifier of the target object, a target queue can be directly created, and the first recognition result is stored in the target queue. It is also possible to determine whether the first confidence is greater than the preset confidence. If it is greater, a target queue is created, and the first recognition result is stored in the target queue. If it is less than or equal to the preset confidence, there is no need to create a target queue.
[0041] In the above embodiments, when storing the first recognition result in the target queue, the time and confidence are fully screened. Therefore, the recognition results of the image frames tilted at different times are screened and added to the queue, improving the accuracy of determining the target recognition result.
[0042] In an exemplary embodiment, the method further includes: in response to the target number of all the recognition results stored in the target queue being greater than a preset number, determining a target score for each of the recognition results stored in the target queue; determining a fourth recognition result with the lowest target score from all the recognition results stored in the target queue; and deleting the fourth recognition result. In this embodiment, if the recognition result of the current frame is added to the queue and exceeds the queue capacity limit, a license plate needs to be selected according to the dequeue rule to remove from the current queue. The judgment rule is as follows: sort the scores of all license plates in the queue from high to low, select the license plate with the lowest score, and remove it from the current queue. That is, when the target number of the recognition results existing in the target queue is greater than the preset number, the recognition results can be considered to be preferentially saved, that is, the recognition results existing in the target queue are scored to obtain the target score, and the fourth recognition result with the lowest target score is deleted from the target queue. Ensure that the recognition results existing in the target queue are all high-quality recognition results, thereby improving the accuracy of determining the target recognition result. That is, if the license plate of the current frame is added to the queue and exceeds the queue capacity limit, a license plate needs to be selected according to the dequeue rule to remove from the current queue. The judgment rule is as follows: sort the scores of all license plates in the queue from high to low, select the license plate with the lowest score, and remove it from the current queue.
[0043] In an exemplary embodiment, determining a target score for each recognition result stored in the target queue includes: determining a second confidence level for each recognition result; determining a third serial number of the image corresponding to each recognition result, and a fourth serial number of the currently recognized image in the image sequence when the target number is greater than the preset number; and determining the target score for each recognition result based on the second confidence level, the third serial number, and the fourth serial number. In this embodiment, when determining the target score, the second confidence level of each recognition result in the target queue and the number of frames in which the image corresponding to each recognition result already exists can be determined, and the target score is determined according to the second confidence level and the number of frames in which the image corresponding to each recognition result already exists. Among them, the number of frames in which the image corresponding to each recognition result already exists is the number of frames determined by the third serial number of the image and the fourth serial number of the currently recognized image.
[0044] In an exemplary embodiment, determining the target score of each recognition result based on the second confidence level, the third serial number, and the fourth serial number includes: determining a first product of the second confidence level of each recognition result and a first coefficient; determining a second difference between the fourth serial number and the third serial number of the image corresponding to each recognition result; normalizing all the second differences corresponding to the recognition results to obtain a third difference; determining a second product of each third difference and a second coefficient; and determining a fourth difference between each first product and the second product as the target score of each recognition result. In this embodiment, the target score can be determined by Score = 1.3 * confidence(i) - 0.3 * plate_nums(i), where i represents the serial number of the image corresponding to the recognition result in the image sequence, confidence(i) represents the confidence level of the license plate, and plate_nums(i) represents the value after normalization of the number of existing frames of the image, that is, the third difference. 1.3 represents the first coefficient, and 0.3 represents the second coefficient. It should be noted that the values of the above first coefficient and second coefficient are only an exemplary illustration, and the first coefficient and second coefficient can also take other values. The first coefficient and second coefficient can be determined according to the interval where the desired target score is located, and the present invention does not limit this. Among them, the flowchart for determining the recognition results in the removal queue can be seen in Appendix Figure 4 。
[0045] Calculating the target score in the manner of the above embodiment can ensure that the recognition results in the queue are all high-quality results. The limitation of the number of existing frames (i.e., the second difference) is because when recognizing, the frame to be captured, that is, the frame that finally needs to be voted on, is generally in a relatively appropriate position. The frame farthest from this position is the frame with the longest appearance time. When the object first appears in the image, it is generally of poor quality and has a large angle. Therefore, the longer the existence time, the more the score decreases.
[0046] In an exemplary embodiment, fusing the image features corresponding to all the recognition results included in the target queue to obtain a fused feature includes: normalizing the confidence levels of all the recognition results in the target queue to obtain a plurality of target confidence levels; determining each of the target confidence levels as the weight of the recognition result corresponding to the target confidence level; and fusing the image features corresponding to each recognition result based on the weights to obtain the fused feature. In this embodiment, when determining the fused feature, the weight of each recognition result can be determined according to the confidence levels of all the recognition results in the target queue. For example, all the confidence levels are normalized to obtain a plurality of target confidence levels, and each obtained target confidence level is determined as the weight of the recognition result corresponding to the target confidence level. After obtaining the weights, the image features corresponding to each recognition result are fused according to the weights to obtain the fused feature.
[0047] In the above embodiment, all the recognition results in the target queue can be processed to obtain the voting recognition result of the current frame, that is, the target recognition result. First, normalize the confidence levels of all the recognition results in the queue to obtain the weight of each recognition information. For a license plate with a high confidence level, it indicates that the image quality is good and the obtained recognition result is relatively reliable. Therefore, a larger weight can be assigned.
[0048] In an exemplary embodiment, fusing the image features corresponding to each recognition result based on the weights to obtain the fused feature includes: determining the product of the image feature corresponding to each recognition result and the weight corresponding to the recognition result to obtain a plurality of third products; and determining the sum of the plurality of third products as the fused feature. In this embodiment, after obtaining the weight of each recognition result, the feature map corresponding to each retained frame is operated on, multiplied by the calculated weight, and then merged into a feature of n * feature dimension, where n is the queue length. A trained network head for license plate recognition can be used to parse the fused feature and output the recognition result, which is also the voting result of the license plate, that is, the target recognition result. Among them, the schematic diagram of determining the target recognition result according to the fused feature can be seen in the appendix Figure 5 .
[0049] In the above embodiment, by fusing the features and confidence levels of historical frames, image features with high confidence levels in the historical frames can be selected for preservation, and all the features are fused and re-recognized by the network to obtain a comprehensive recognition result, rather than directly using the string for frequency judgment to generate the voting result. This can effectively avoid the problem of accumulating incorrect recognition results caused by poor image quality, occlusion, overexposure and other external factors in historical frames, thereby obtaining an incorrect voting result.
[0050] The method for determining the recognition result will be described below in conjunction with specific embodiments:
[0051] Figure 6 is a flowchart of a method for determining the recognition result according to a specific embodiment of the present invention. As Figure 6 shown, the method includes:
[0052] Step S602, input a video frame. For road monitoring, the input should be a video stream. Each frame of the video stream is separately fed into the entire network for processing.
[0053] Step S604, license plate recognition + tracking algorithm to obtain the recognition result of a single frame. License plate recognition network: Use a CNN network to extract features from the input image, which will use convolutional layers, normalization layers, activation layers, pooling layers, etc. Finally, a high-dimensional abstract semantic feature will be obtained, and this feature needs to be saved and will be used by the subsequent voting network. Using this feature, the license plate recognition result of this frame can be obtained through the recognition network head parsing. Then, it is fed into the tracking algorithm to obtain the license plate id, which marks each different license plate.
[0054] Step S606, determine the license plate features and license plate information.
[0055] Step S608, determine whether the enqueue condition is satisfied. If the judgment result is yes, execute step S610; if the judgment result is no, execute step S612.
[0056] Step S610, voting queue.
[0057] Step S612, discard the result of the current frame.
[0058] Step S614, when the queue is full, determine whether each recognition result in the queue satisfies the dequeue condition. If the judgment result is yes, execute step S616; if the judgment result is no, execute step S610.
[0059] Step S616, dequeue
[0060] Step S618, when the queue is not full, determine the multi-frame voting result through the voting network.
[0061] In the foregoing embodiments, based on the license plate voting strategy and queue maintenance strategy in the foregoing embodiments, it is possible to effectively avoid the serious impact of the error information in the historical frames on the voting result, and thus effectively improve the license plate recognition rate. By introducing confidence judgment, on the one hand, the queue can be maintained to ensure the quality of the license plate images in the queue. On the other hand, the features can be weighted, so that the historical frames with clear images can have a great decisive influence on the final voting result, and will not be affected by other historical frames that may have a large number of frames but poor quality. Instead of simply comparing the license plate recognition results of each frame for strings, re-recognition is performed using weighted image features, which can effectively avoid the adverse impact of blurred historical frame images on the recognition result. It can also avoid the complex judgment of strings and the disadvantages of easy mis-voting and the need to rewrite different voting rules for different scenarios. It has good stability and maintainability.
[0062] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0063] In this embodiment, a device for determining the recognition result is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0064] Figure 7 is a structural block diagram of a device for determining the recognition result according to an embodiment of the present invention. As Figure 7 shown, the device includes:
[0065] A recognition module 72, configured to recognize a target object included in the first image, obtain a first recognition result and a first confidence level of the first recognition result;
[0066] A storage module 74, configured to store the first recognition result into a target queue in response to the first recognition result and the first confidence level meeting a predetermined condition, where the target queue is further configured to store a second recognition result that is the same as the identifier of the target object, and the second recognition result is a result obtained by recognizing an object included in a second image, and the second image is an image where an object with the same identifier as the target object is located in the image sequence where the first image is located;
[0067] A fusion module 76, configured to fuse image features corresponding to all recognition results included in the target queue to obtain a fused feature;
[0068] A determination module 78, configured to determine a target recognition result corresponding to the target queue based on the fused feature.
[0069] In an exemplary embodiment, the storage module 74 may implement storing the first recognition result into the target queue in response to the first recognition result and the first confidence level meeting a predetermined condition in the following manner: in response to the target queue being an existing queue, determining a first serial number of a third image corresponding to a third recognition result that was most recently stored into the target queue, where the first serial number is the serial number of the third image in the image sequence; determining a first difference between the first serial number and a second serial number of the first image in the image sequence; in response to the first difference being greater than a preset difference, determining whether the first confidence level of the first recognition result is greater than a preset confidence level; and in response to the first confidence level being greater than the preset confidence level, storing the first recognition result into the target queue.
[0070] In an exemplary embodiment, the apparatus is further configured to, in response to the target number of all recognition results stored in the target queue being greater than a preset number, determine a target score for each recognition result stored in the target queue; determine a fourth recognition result with the smallest target score from all the recognition results stored in the target queue; and delete the fourth recognition result.
[0071] In an exemplary embodiment, the apparatus may implement determining the target score for each recognition result stored in the target queue in the following manner: determining a second confidence level for each recognition result; determining a third serial number of an image corresponding to each recognition result, and a fourth serial number of the currently recognized image in the image sequence when the target number is greater than the preset number; and determining the target score for each recognition result based on the second confidence level, the third serial number, and the fourth serial number.
[0072] In an exemplary embodiment, the device may determine the target score of each of the recognition results based on the second confidence level, the third serial number, and the fourth serial number in the following manner: determine a first product of the second confidence level of each of the recognition results and a first coefficient; determine a second difference between the fourth serial number and the third serial number of the image corresponding to each of the recognition results; normalize all the second differences corresponding to the recognition results to obtain a third difference; determine a second product of each of the third differences and a second coefficient; and determine the fourth difference between each of the first products and the second products as the target score of each of the recognition results.
[0073] In an exemplary embodiment, the fusion module 76 may fuse the image features corresponding to all the recognition results included in the target queue to obtain a fused feature in the following manner: normalize the confidence levels of all the recognition results in the target queue to obtain a plurality of target confidence levels; determine each of the target confidence levels as the weight of the recognition result corresponding to the target confidence level; and fuse the image features corresponding to each of the recognition results based on the weights to obtain the fused feature.
[0074] In an exemplary embodiment, the fusion module 76 may fuse the image features corresponding to each of the recognition results based on the weights to obtain the fused feature in the following manner: determine the product of the image feature corresponding to each of the recognition results and the weight corresponding to the recognition result to obtain a plurality of third products; and determine the sum of the plurality of third products as the fused feature.
[0075] It should be noted that the above-mentioned respective modules may be implemented by software or hardware. For the latter, it may be implemented in the following manner, but not limited thereto: all the above-mentioned modules are located in the same processor; or, the above-mentioned respective modules are located in different processors in any combined form.
[0076] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0077] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0078] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0079] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device. The transmission device is connected to the processor, and the input / output device is connected to the processor.
[0080] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0081] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0082] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for determining an identification result, characterized in that, Including: Identifying a target object included in a first image to obtain a first recognition result and a first confidence level of the first recognition result; In response to the first recognition result and the first confidence level satisfying a predetermined condition, storing the first recognition result in a target queue, where the target queue is further used to store a second recognition result having the same identifier as the target object, and the second recognition result is a result obtained by identifying an object included in a second image, and the second image is an image in an image sequence where the object having the same identifier as the target object is located in the first image; Fusing image features corresponding to all recognition results included in the target queue to obtain a fused feature; Determining a target recognition result corresponding to the target queue based on the fused feature; In response to the first recognition result and the first confidence level satisfying a predetermined condition, storing the first recognition result in the target queue includes: in response to the target queue being an existing queue, determining a first sequence number of a third image corresponding to a third recognition result that was most recently stored in the target queue, where the first sequence number is the sequence number of the third image in the image sequence; determining a first difference between the first sequence number and a second sequence number of the first image in the image sequence; in response to the first difference being greater than a preset difference, determining whether the first confidence level of the first recognition result is greater than a preset confidence level; and in response to the first confidence level being greater than the preset confidence level, storing the first recognition result in the target queue; The method further includes: in response to a target quantity of all recognition results stored in the target queue being greater than a preset quantity, determining a target score for each recognition result stored in the target queue; determining a fourth recognition result having the smallest target score from all the recognition results stored in the target queue; and deleting the fourth recognition result.
2. The method according to claim 1, wherein Determining the target score for each recognition result stored in the target queue includes: Determining a second confidence level for each recognition result; Determining a third sequence number of an image corresponding to each recognition result, and a fourth sequence number of the currently recognized image in the image sequence when the target quantity is greater than the preset quantity; Determining the target score for each recognition result based on the second confidence level, the third sequence number, and the fourth sequence number.
3. The method according to claim 2, wherein Determining the target score for each recognition result based on the second confidence level, the third sequence number, and the fourth sequence number includes: Determining a first product of the second confidence level of each recognition result and a first coefficient; Determining a second difference between the fourth sequence number and the third sequence number of the image corresponding to each recognition result; Normalizing all the second differences corresponding to the recognition results to obtain a third difference; Determining a second product of each third difference and a second coefficient; Determining a fourth difference between each first product and the second product as the target score for each recognition result.
4. The method according to claim 1, wherein Fusing image features corresponding to all recognition results included in the target queue to obtain a fused feature includes: Normalize the confidence levels of all the recognition results in the target queue to obtain multiple target confidence levels; Determine each of the target confidence levels as the weight of the recognition result corresponding to the target confidence level; Fuse the image features corresponding to each of the recognition results based on the weights to obtain the fused feature.
5. The method according to claim 4, characterized in that, Fusing the image features corresponding to each of the recognition results based on the weights to obtain the fused feature includes: Determine the product of the image feature corresponding to each of the recognition results and the weight corresponding to the recognition result to obtain multiple third products; Determine the sum of the multiple third products as the fused feature.
6. An apparatus for determining an identification result, characterized in that, Includes: A recognition module, configured to recognize a target object included in a first image, to obtain a first recognition result and a first confidence level of the first recognition result; A storage module, configured to, in response to the first recognition result and the first confidence level satisfying a predetermined condition, store the first recognition result into a target queue, where the target queue is further configured to store a second recognition result having the same identifier as the target object, the second recognition result being a result obtained by recognizing an object included in a second image, and the second image being an image in the image sequence where the first image is located and where the object having the same identifier as the target object is located; A fusion module, configured to fuse the image features corresponding to all the recognition results included in the target queue to obtain a fused feature; A determination module, configured to determine a target recognition result corresponding to the target queue based on the fused feature; The storage module realizes storing the first recognition result into the target queue in response to the first recognition result and the first confidence level satisfying a predetermined condition in the following manner: in response to the target queue being an existing queue, determine a first sequence number of a third image corresponding to a third recognition result that was most recently stored into the target queue, where the first sequence number is the sequence number of the third image in the image sequence; determine a first difference between the first sequence number and a second sequence number of the first image in the image sequence; in response to the first difference being greater than a preset difference, determine whether the first confidence level of the first recognition result is greater than a preset confidence level; in response to the first confidence level being greater than the preset confidence level, store the first recognition result into the target queue; The apparatus is further configured to: in response to the target number of all the recognition results stored in the target queue being greater than a preset number, determine a target score for each of the recognition results stored in the target queue; determine a fourth recognition result having the smallest target score from all the recognition results stored in the target queue; and delete the fourth recognition result.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, where when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. An electronic device, comprising a memory and a processor, characterized in that The memory stores a computer program, and the processor is configured to run the computer program to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Improved plate number voting method and device
CN106778743A
Video identification method and device and computer equipment
CN111027347A
License plate voting method and device, computer equipment and storage medium
CN111860590A
License plate recognition method, device and equipment and storage medium
CN112651417A
License plate character recognition method and apparatus, and device and storage medium
WO2022205018A1