Method and device for determining imaging quality of video acquisition equipment, electronic equipment, storage medium and train

By acquiring multi-frame images and calculating the confidence dynamic mean using the target image recognition model, combining preset sliding windows and environmental information, the imaging quality of the video acquisition device is automatically detected, and the problem of high manual detection costs is solved, and accurate and real-time imaging quality detection is achieved.

CN120281894APending Publication Date: 2025-07-08CRRC QINGDAO SIFANG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585119.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the imaging quality detection of video acquisition equipment relies on manual inspection, resulting in high labor costs and low real-time and accuracy.

Method used

By acquiring multi-frame images, using a pre-constructed target image recognition model for object recognition, calculating the confidence dynamic mean, and combining preset sliding windows and environmental information, the imaging quality of the video acquisition device is automatically determined.

Benefits of technology

It realizes accurate and real-time detection of the imaging quality of video acquisition equipment, reduces labor costs, eliminates the need for additional equipment, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281894A_ABST
    Figure CN120281894A_ABST
Patent Text Reader

Abstract

The invention provides a method for determining the imaging quality of video acquisition equipment, and the method comprises the steps: obtaining a plurality of frames of images corresponding to a plurality of time windows, and enabling the plurality of frames of images to be obtained through the shooting of a target object through target video acquisition equipment; performing target object recognition processing on the multiple frames of images by using a pre-constructed target image recognition model to obtain a plurality of confidence coefficients corresponding to the multiple frames of images, the confidence coefficients being used for representing the recognition accuracy of the target image recognition model on the target object in the image; determining at least one target confidence coefficient from the multiple confidence coefficients by using a preset sliding window; based on the at least one target confidence coefficient, calculating to obtain a confidence coefficient dynamic mean value; and determining the imaging quality of the target video acquisition equipment by using the confidence coefficient dynamic mean value. The invention further provides a device for determining the imaging quality of the video acquisition equipment, electronic equipment, a computer readable storage medium and a train.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of rail transit and image processing, and more particularly, to a method, apparatus, electronic device, storage medium, and train for determining the imaging quality of a video acquisition device. Background Art

[0002] By using a video acquisition device to capture a detection object and identifying the captured image, it is possible to determine whether the detection object has a fault and determine the fault location. The imaging quality of the video acquisition device affects the accuracy of fault identification. For example, when the lens of the video acquisition device is dirty (e.g., contaminated by dust, rain, snow, frost, etc.) or the light is insufficient, the quality of the captured image may be low, resulting in low fault identification accuracy and easy false alarms.

[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: The related methods usually require manual inspection of the video acquisition device at regular intervals, which takes a large amount of time and labor costs of the staff. Summary of the Invention

[0004] In view of this, the present disclosure provides a method, apparatus, electronic device, storage medium, and train for determining the imaging quality of a video acquisition device.

[0005] One aspect of the present disclosure provides a method for determining the imaging quality of a video acquisition device, including: obtaining multiple frames of images corresponding to multiple time windows, where the multiple frames of images are captured by a target video acquisition device for a target object; using a pre-constructed target image recognition model to perform target object recognition processing on the multiple frames of images to obtain multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the target object in the image by the target image recognition model; using a preset sliding window to determine at least one target confidence level from the multiple confidence levels; calculating a confidence level dynamic mean based on the at least one target confidence level; and determining the imaging quality of the target video acquisition device using the confidence level dynamic mean.

[0006] According to an embodiment of the present disclosure, using a preset sliding window to determine at least one target confidence level from the multiple confidence levels includes: determining the environmental information corresponding to the target video acquisition device; adjusting the window width of the preset sliding window based on the environmental information; and determining at least one target confidence level from the multiple confidence levels according to the window width, where the environmental information includes at least one of the following: weather information, optical environment information, and operating environment information.

[0007] According to an embodiment of the present disclosure, among them, the multiple time windows include a current time window and at least one historical time window, and the multiple frames of images include at least one current image corresponding to the current time window and at least one historical image corresponding to at least one historical time window; using a preset sliding window to determine at least one target confidence from multiple confidences includes: adding the confidences corresponding to at least one current image to an initial confidence queue to obtain an updated confidence queue, where the initial confidence queue includes at least one confidence corresponding to at least one historical image; adjusting the positions of multiple confidences in the updated confidence queue according to the timestamps corresponding to the confidences to obtain a target confidence queue; using a preset sliding window to determine at least one target confidence from the target confidence queue.

[0008] According to an embodiment of the present disclosure, among them, using a preset sliding window to determine at least one target confidence from the target confidence queue includes: determining at least one target confidence from a predetermined position of the updated confidence queue using a preset sliding window according to the timestamp order in the target confidence queue.

[0009] According to an embodiment of the present disclosure, the method for determining the imaging quality of a video acquisition device further includes: determining a target loss function based on a classification loss function and a regression loss function; training an initial image recognition model based on the target loss function to obtain a target image recognition model.

[0010] According to an embodiment of the present disclosure, among them, obtaining multiple frames of images corresponding to multiple time windows includes: obtaining a video stream corresponding to a target object, where the video stream is obtained by shooting the target object using a target video acquisition device; performing frame division processing on the video stream to obtain an image queue; screening out multiple frames of images from the image queue according to the running state information corresponding to the target object.

[0011] Another aspect of the present disclosure provides a device for determining the imaging quality of a video acquisition device, including: an acquisition module, configured to acquire multiple frames of images corresponding to multiple time windows, where the multiple frames of images are obtained by shooting a target object using a target video acquisition device; an identification module, configured to perform target object identification processing on the multiple frames of images using a pre-constructed target image recognition model to obtain multiple confidences corresponding to the multiple frames of images, where the confidence is used to characterize the recognition accuracy of the image recognition model for the target object in the image; a first determination module, configured to determine at least one target confidence from the multiple confidences using a preset sliding window; a calculation module, configured to calculate a confidence dynamic mean based on at least one target confidence; a second determination module, configured to determine the imaging quality of the target video acquisition device using the confidence dynamic mean.

[0012] Another aspect of the present disclosure provides an electronic device, including:

[0013] One or more processors;

[0014] A memory for storing one or more programs,

[0015] wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0016] Another aspect of the present disclosure provides a train including a video acquisition device, and the imaging quality of the video acquisition device is determined by the method for determining the imaging quality of the video acquisition device as described above.

[0017] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions that are used to implement the method as described above when executed.

[0018] Another aspect of the present disclosure provides a computer program product, which includes computer-executable instructions that are used to implement the method as described above when executed.

[0019] According to an embodiment of the present disclosure, by using a target image recognition model to perform target object recognition processing on multiple frames of images, multiple confidence levels are obtained, and then a target confidence level is determined from the multiple confidence levels. Based on the target confidence level, a dynamic mean of the confidence levels is calculated, and the imaging quality of the target video acquisition device is determined by using the dynamic mean of the confidence levels. Technologies such as deep learning and computer vision in the field of artificial intelligence can be used to accurately and real-time determine the imaging quality of the video acquisition device, reducing labor costs and without the need to additionally increase other devices, further reducing costs. Therefore, at least partially, it overcomes the problems of low real-time performance and low accuracy brought by the related methods of relying on the experience and knowledge of professional technicians and visual inspection and maintenance operations, and realizes the technology of accurately, real-time, and low-cost determining the imaging quality of the video acquisition device. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0021] Figure 1 Schematically shows an exemplary system architecture to which the method, apparatus, electronic device, storage medium, and train for determining the imaging quality of a video acquisition device according to the present disclosure can be applied;

[0022] Figure 2 Schematically shows a flowchart of the method for determining the imaging quality of a video acquisition device according to an embodiment of the present disclosure;

[0023] Figure 3A schematic diagram showing the prediction boxes obtained when the camera is not dirty according to an embodiment of the present disclosure;

[0024] Figure 4 A schematic diagram showing the prediction boxes obtained when the camera is dirty according to an embodiment of the present disclosure;

[0025] Figure 5 A flowchart showing a method for determining the imaging quality of a video acquisition device according to another embodiment of the present disclosure;

[0026] Figure 6 A block diagram showing a device for determining the imaging quality of a video acquisition device according to an embodiment of the present disclosure; and

[0027] Figure 7 A block diagram showing a schematic diagram suitable for implementing a method for determining the imaging quality of a video acquisition device according to an embodiment of the present disclosure. Detailed implementation manners

[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0032] In the embodiments of the present disclosure, with respect to aspects such as the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the involved data (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0033] In the embodiments of the present disclosure, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.

[0034] In related methods, it is possible to identify the image of the detection object captured by the video acquisition device to determine the fault condition of the detection object. For example: The detection object may include the pantograph of a train. The pantograph is a key component of the high-voltage system of the multiple unit train, installed on the top of the multiple unit train vehicle, and is a mechanical component that obtains and transmits current using the overhead catenary. Due to the harsh operating conditions of the pantograph, with the increase in the operating mileage of the multiple unit train and the interference of the external environment, various types of faults are likely to occur in the pantograph, affecting the safety of train operation and the order of railway operation. The video acquisition device may include a pantograph video monitoring system, which is used to monitor the working status of the overhead pantograph and the catenary in real time during train operation, and also takes into account the working status of the high-voltage equipment attached to the pantograph, providing monitoring videos and analysis images for identifying whether the pantograph has failed.

[0035] When the lens of the video acquisition device is dirty (for example, contaminated by dust, rain, snow, ice, etc.) or the light is insufficient, the quality of the captured image may be low, resulting in a low accuracy of fault identification and prone to false alarms. For example: For the pantograph video monitoring system, when the camera screen of the monitoring system is dirty, it will cause the quality of the monitoring video to decline, and then the model cannot effectively identify the pantograph and other high-voltage equipment areas, resulting in a low accuracy of fault identification and prone to misjudgment and false alarms of faults.

[0036] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: In related methods, the video acquisition device is usually manually inspected regularly, resulting in a high labor cost. For example: When identifying the firmware of the pantograph, related methods mostly rely on the on-board mechanic to determine and record in real time, which occupies a large amount of the mechanic's time cost.

[0037] In view of this, embodiments of the present disclosure provide a method for determining the imaging quality of a video acquisition device, including: obtaining multiple frames of images corresponding to multiple time windows, where the multiple frames of images are obtained by using a target video acquisition device to photograph a target object; using a pre-constructed target image recognition model to perform target object recognition processing on the multiple frames of images to obtain multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the target object in the image by the target image recognition model; using a preset sliding window to determine at least one target confidence level from the multiple confidence levels; calculating a dynamic mean of the confidence levels based on the at least one target confidence level; and determining the imaging quality of the target video acquisition device by using the dynamic mean of the confidence levels.

[0038] Figure 1 Schematically shows an exemplary system architecture 100 to which the method, apparatus, electronic device, storage medium, and train for determining the imaging quality of a video acquisition device according to embodiments of the present disclosure can be applied. It should be noted that, Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0039] As Figure 1 shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0040] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0042] Server 105 may be a server that provides various services, such as a background management server (only for example) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0043] It should be noted that the method for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the system for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can generally be set in server 105. The method for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the system for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the method for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can also be executed by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the system for determining the imaging quality of the video acquisition device provided by the embodiments of the present disclosure can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or can be set in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0044] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0045] Figure 2 is merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0046] As Figure 2 shown, the method includes operations S210 to S250.

[0047] In operation S210, multiple frames of images corresponding to multiple time windows are acquired, where the multiple frames of images are obtained by capturing a target object using a target video acquisition device.

[0048] In operation S220, using a pre-constructed target image recognition model, the multiple frames of images are subjected to target object recognition processing to obtain multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the target object in the image by the target image recognition model.

[0049] In operation S230, using a preset sliding window, at least one target confidence level is determined from the multiple confidence levels.

[0050] In operation S240, based on the at least one target confidence level, a dynamic mean of the confidence levels is calculated.

[0051] In operation S250, using the dynamic mean of the confidence levels, the imaging quality of the target video acquisition device is determined.

[0052] According to an embodiment of the present disclosure, in operation S210, the video acquisition device is used to capture a target object. Based on the captured image or video, it can be determined whether the target object has a fault. The video acquisition device may include a surveillance camera, a video camera, etc., and may also include other devices for acquiring the video of the target object, which is not limited herein. For example: The target object may include the base area and the bow head area of a pantograph. The video acquisition device may include a video camera installed on a train for capturing the pantograph. The pantograph can be captured in real time by the video camera to obtain a video stream corresponding to the pantograph. By performing frame division processing on the video stream, multiple frames of images corresponding to multiple time windows can be obtained.

[0053] According to an embodiment of the present disclosure, in operation S220, using a pre-constructed target image recognition model, the target object recognition processing of the multiple frames of images may include: inputting the multiple frames of images into the target image recognition model so that the target image recognition model recognizes the target object in the image, generates a prediction box corresponding to the target object, and generates a confidence level corresponding to the prediction box, where the confidence level can characterize the recognition accuracy of the target image recognition model. When the confidence level is high, it indicates that the probability that the object in the prediction box is the target object is higher. On the contrary, when the imaging quality of the video acquisition device is low, for example, when the camera is dirty due to rain or snow weather, it will cause an inability to obtain a clear image, and the confidence level of the prediction box generated by the image recognition model is low, and the recognition accuracy is low.

[0054] Figure 3 Schematically shows a schematic diagram of the prediction box obtained when the camera is not dirty according to an embodiment of the present disclosure.

[0055] Figure 4A schematic diagram showing the predicted bounding boxes obtained when the camera is dirty according to an embodiment of the present disclosure.

[0056] As Figure 3 shown, the image obtained by the camera without dirt is clearer. Therefore, the target image recognition model can accurately recognize the pantograph head area in the image and obtain a relatively high confidence level, for example, it can be 1.00. When the camera is dirty (as Figure 4 shown), the obtained image is relatively blurred and cannot clearly display the head area, so that the target image recognition model is difficult to accurately recognize the pantograph head area in the image. Therefore, the confidence level that can be obtained is relatively low, only 0.01.

[0057] According to an embodiment of the present disclosure, in operation S230, if the confidence levels of all images are directly calculated, when the number of images is too large, the calculation complexity and latency will be significantly increased, making it difficult to detect the imaging quality in real time. Therefore, a preset sliding window can be used to select at least one target confidence level that meets the preset conditions from multiple confidence levels. By only calculating and processing some of the target images in the multiple images, the calculation speed can be increased and the calculation complexity can be reduced.

[0058] According to an embodiment of the present disclosure, in operations S240 to S250, the mean value of at least one target confidence level can be calculated to obtain the dynamic mean confidence level. The calculation method of the dynamic mean confidence level is shown, for example, in the following formula (1).

[0059] (1)

[0060] Where is the dynamic mean confidence level, and C n is the nth target confidence level.

[0061] According to an embodiment of the present disclosure, the dynamic mean confidence level can be directly used to characterize the imaging quality of the target video acquisition device. For example, when the dynamic mean confidence level is greater than or equal to the preset threshold, it can be considered that the imaging quality of the target video acquisition device is relatively high. When the dynamic mean confidence level is less than the preset threshold, it can be considered that the imaging quality of the target video acquisition device is relatively low. In the case where the dynamic mean confidence level is less than the preset threshold, an alarm message can be generated so that the operator can timely discover that the imaging quality of the target video acquisition device is relatively low based on the alarm message, and thus the target video acquisition device can be timely repaired.

[0062] For example, the target object may include the bow head part and the base part of the pantograph. The confidence dynamic mean corresponding to the bow head part and the confidence dynamic mean corresponding to the base part can be calculated respectively. When the confidence dynamic mean corresponding to the bow head part is less than a preset threshold, an alarm message of "the image imaging quality of the pantograph bow head area is low" can be generated. When the confidence dynamic mean corresponding to the base part is less than a preset threshold, an alarm message of "the image imaging quality of the pantograph base area is low" can be generated.

[0063] According to an embodiment of the present disclosure, a target object recognition model is used to perform target object recognition processing on multiple frames of images to obtain multiple confidences, and then a target confidence is determined from the multiple confidences. Based on the target confidence, a confidence dynamic mean is calculated, and the imaging quality of the target video acquisition device is determined using the confidence dynamic mean. Technologies such as deep learning and computer vision in the field of artificial intelligence can be used to accurately and real-time determine the imaging quality of the video acquisition device, reduce labor costs, and avoid the problems of low real-time performance and low accuracy brought by the related methods relying on the experience and knowledge of professional technicians and visual inspection and maintenance operations. And there is no need to additionally add other devices, so the cost is relatively low.

[0064] According to an embodiment of the present disclosure, using a preset sliding window to determine at least one target confidence from multiple confidences includes: determining the environmental information corresponding to the target video acquisition device; adjusting the window width of the preset sliding window based on the environmental information; and determining at least one target confidence from the multiple confidences according to the window width; wherein, the environmental information includes at least one of the following: weather information, optical environment information, and operating environment information.

[0065] Among them, the multiple confidences can be in the form of a confidence queue, as shown in Table 1 for example.

[0066] Table 1

[0067] Name 1 2 … N <![CDATA[Bow region confidence CP i > <![CDATA[CP1]]> <![CDATA[CP2]]> … <![CDATA[CP N > <![CDATA[Base area confidence CB i > <![CDATA[CB1]]> <![CDATA[CB2]]> … <![CDATA[CB N >

[0068] Specifically, the confidence queue can be a queue with a fixed length of N, and N can represent the number of confidences in the confidence queue. The window width of the preset sliding window can be n, and n is less than or equal to N. The preset sliding window can slide in the confidence queue, and each time it slides, it will select n consecutive confidence values in the confidence queue for subsequent calculation of the confidence dynamic mean based on the n confidence values.

[0069] For example: The confidence of the pantograph bow head area and the confidence of the pantograph base area are detected respectively to obtain the bow head area confidence queue and the base area confidence queue.

[0070] According to embodiments of the present disclosure, when the target video acquisition device captures a target object, it may be subject to short-term interference, resulting in short-term fluctuations in confidence. For example, during shooting, it may be interfered by light changes (such as entering or exiting a tunnel), weather changes (such as rain, snow or other weather conditions), or other factors, resulting in a decrease in imaging quality and short-term fluctuations in confidence. If the confidence dynamic mean is directly calculated based on all the confidences in the confidence queue, these short-term fluctuations may be overly smoothed, thus masking the actual decrease in imaging quality. By dynamically selecting multiple target confidences from the confidence queue through a preset sliding window, if the confidences of the recent few frames of images fluctuate, such as a sudden decrease in confidence, the preset sliding window can quickly detect this fluctuation without being diluted by earlier historical confidence data.

[0071] For example: Multiple frames of images for the pantograph base area can be collected, and a confidence queue of length 10 can be obtained. The confidence queue of length 10 can be expressed as: [0.96, 0.95, 0.95, 0.94, 0.91, 0.69, 0.70, 0.95, 0.95, 0.91]. Among them, the first 5 confidences are relatively high, indicating high imaging quality; the 6th and 7th confidences suddenly drop to 0.69 and 0.70, which may be due to the train entering the tunnel, resulting in a darker light, and thus a lower imaging quality and a decrease in confidence; the 8th, 9th, and 10th confidences recover to 0.90 and 0.91, which may be due to the train exiting the tunnel and the light recovering, so the confidence is relatively high.

[0072] For the above-mentioned confidence queue of length 10, if the confidence dynamic mean is directly calculated based on the entire confidence queue, the obtained confidence dynamic mean is relatively high, which is 0.894, and it cannot reflect the short-term fluctuations of the 6th and 7th confidences. By setting a sliding window, the width of the sliding window is, for example, 3, and the sliding window slides sequentially in the queue, and the confidence dynamic mean is calculated sequentially. First, the sliding window will select the first 3 confidences in the confidence queue, that is, 0.96, 0.95, 0.95, and calculate the confidence dynamic mean. The first confidence dynamic mean obtained is 0.94; further, the sliding window will make the first slide, select the 2nd, 3rd, and 4th confidences in the confidence queue, that is, 0.94, 0.93, 0.92, and calculate the confidence dynamic mean. The second confidence dynamic mean obtained is 0.933; when the sliding window makes the fourth, fifth, and sixth slides, the dynamic mean decreases from 0.92 to 0.893, 0.853, 0.833 in sequence, so that the dynamic fluctuations of the confidence can be accurately captured.

[0073] According to embodiments of the present disclosure, when the window width is small, it is vulnerable to instantaneous interference, and when the window is large, short-term fluctuations in confidence cannot be captured in a timely manner. Therefore, in order to enable the window width to respond to changes in confidence in a timely manner and accurately reflect the dynamic fluctuations in confidence, the window width of a preset sliding window can be dynamically adjusted according to environmental information. The environmental information may include at least one of the following: weather information, optical environment information, and operating environment information.

[0074] Among them, the weather information may include good weather such as sunny days, and may also include bad weather such as rainy, snowy, and foggy days, and may also include other types of weather, which are not limited herein. The window width can be dynamically adjusted according to the weather information. For example, a smaller window width is selected in good weather such as sunny days, for example, a window width less than a preset width threshold is selected; a larger window width is selected in bad weather, for example, a window width greater than or equal to the preset width threshold is selected.

[0075] The optical environment information may include light intensity, light direction, etc. The window width can be dynamically adjusted according to the light intensity. For example, a window width less than the preset width threshold can be selected under the condition that the light intensity is higher than the preset light intensity threshold; a window width greater than or equal to the preset width threshold can be selected under the condition that the light intensity is lower than or equal to the preset light intensity threshold.

[0076] The operating environment information may include the train running position. The train running position may include a normal position and a special position. Among them, the special position may include special positions such as tunnels, and the normal position may be other positions except for special positions. When the train enters a tunnel, the window width can be increased to smooth the confidence fluctuations caused by rapid changes in light.

[0077] According to embodiments of the present disclosure, by comprehensively considering the influences of weather information, optical environment information, and operating environment information, and adjusting the window width of the preset sliding window according to the above information, when short-term fluctuations in imaging quality occur due to light changes, weather changes, or other interference factors, these changes can be quickly detected, so that the imaging quality of the video acquisition device can be accurately and real-time determined.

[0078] According to embodiments of the present disclosure, multiple time windows include a current time window and at least one historical time window, and multiple frames of images include at least one current image corresponding to the current time window and at least one historical image corresponding to at least one historical time window.

[0079] According to embodiments of the present disclosure, using a preset sliding window to determine at least one target confidence from multiple confidences includes operations 11 to 13.

[0080] In operation 11, add the confidence corresponding to at least one current image to the initial confidence queue to obtain an updated confidence queue, where the initial confidence queue includes at least one confidence corresponding to at least one historical image.

[0081] In operation 12, adjust the positions of multiple confidences in the updated confidence queue according to the timestamps corresponding to the confidences to obtain a target confidence queue.

[0082] In operation 13, use a preset sliding window to determine at least one target confidence from the target confidence queue.

[0083] According to an embodiment of the present disclosure, in operation 11, adding the confidence of the current image to the initial confidence queue can combine the confidence at the current moment with the historical confidence for subsequent calculation of the target confidence. Specifically, each of the current image and the historical image corresponds to a timestamp, which is used to represent the acquisition time of the current image and the historical image respectively. Through the timestamps, the chronological order of each frame of image in time can be determined.

[0084] In operation 12, adjusting the positions of multiple confidences in the updated confidence queue according to the timestamps corresponding to the confidences may include: sorting the multiple confidences in the updated confidence queue in the chronological order of the acquisition time. The positions of multiple confidences in the updated confidence queue can be adjusted by a preset sorting algorithm. For example, quicksort, mergesort, etc. can be used to sort in the chronological order of the acquisition time.

[0085] In operation 13, after adjusting the positions of multiple confidences in the updated confidence queue to obtain a target confidence queue, using a preset sliding window to determine the target confidence from the target confidence queue can ensure that the confidences covered by the preset sliding window correspond to continuous time periods, and can more accurately reflect the dynamic changes of the confidences.

[0086] According to an embodiment of the present disclosure, in the actual process of determining the imaging quality, image acquisition may be affected by various factors such as network latency, resulting in the actual reception order of image data being inconsistent with the actual acquisition time order, and further resulting in the confidence sorting in the confidence queue being inconsistent with the image acquisition time order. In the case where the confidence sorting is inconsistent with the image acquisition time order, the confidence dynamic mean calculated directly by selecting confidences from the queue may not be reasonable enough and cannot accurately reflect the short-term fluctuations of the confidences. By adjusting the positions of multiple confidences in the updated confidence queue according to the timestamps corresponding to the confidences to obtain a target confidence queue, and using a preset sliding window to determine at least one target confidence from the target confidence queue, the confidence dynamic mean calculated subsequently can more accurately reflect the dynamic changes of the confidences.

[0087] According to an embodiment of the present disclosure, determining at least one target confidence from a target confidence queue by using a preset sliding window includes: determining at least one target confidence from a predetermined position of an updated confidence queue according to the timestamp order in the target confidence queue by using the preset sliding window.

[0088] According to an embodiment of the present disclosure, specifically, determining at least one target confidence from a predetermined position of an updated confidence queue according to the timestamp order in the target confidence queue may include: determining at least one target confidence from a predetermined position of an updated confidence queue according to the timestamp order in the target confidence queue by using the first-in-first-out rule with the preset sliding window. Multiple confidences in the target confidence queue are sorted according to the image acquisition time. By using the first-in-first-out rule, the preset sliding window processes the confidence that enters the target confidence queue earliest in the order of image acquisition time, and removes the processed confidence from the target confidence queue. For example, confidences 0.95, 0.93, and 0.94 are the first three data to enter the target confidence queue. The preset sliding window will first process the above three data, remove the above three processed data from the target confidence queue, and then continue to process other data in the target confidence queue according to the timestamp order.

[0089] According to an embodiment of the present disclosure, determining the target confidence from a predetermined position of the updated confidence queue according to the timestamp order in the target confidence queue reduces the complexity of the processing process and can improve the accuracy and reliability of the analysis.

[0090] According to an embodiment of the present disclosure, the method for determining the imaging quality of a video acquisition device further includes: determining a target loss function based on a classification loss function and a regression loss function; training an initial image recognition model based on the target loss function to obtain an image recognition model.

[0091] According to an embodiment of the present disclosure, the initial image recognition model may adopt a lightweight neural network and use depthwise separable convolution to reduce model parameters, thereby improving the detection speed of the image recognition model to facilitate the implementation of real-time monitoring functions. Further, a feature pyramid network may be used to achieve the fusion of multi-scale information to detect target objects of different scales. On this basis, a context module may be further used to enhance the feature extraction ability, enhance the receptive field, improve the recognition accuracy, and increase the robustness of the network. The initial image recognition model may adopt a head structure (Head structure), and the head structure may be responsible for specific prediction tasks. Among them, the classification prediction result output by the classification head can be used to determine whether the prior box contains the detected object, and the bounding box regression head can be used to predict and adjust the position of the prediction box.

[0092] According to an embodiment of the present disclosure, an initial image recognition model can be pre-trained to obtain a target image recognition model. Among them, training the initial image recognition model can include: constructing a training sample set and training the initial image recognition model to obtain the target image recognition model.

[0093] Specifically, constructing the training sample set can include: dividing the image samples pre-acquired through a video acquisition device according to a preset scenario to obtain an image sample set corresponding to each scenario. For example: The real video of the pantograph monitoring of the multiple unit train can be obtained by the in-vehicle download method, and the pantograph monitoring video can be frame-processed to obtain multiple pantograph image samples. The multiple frames of images can be screened and classified according to a total of 5 preset weather scenarios including sunny, cloudy, rainy, snowy, and foggy days to obtain the training data set samples corresponding to the above 5 preset weather scenarios respectively.

[0094] According to an embodiment of the present disclosure, after obtaining the initial training data set samples, the initial training data set samples can be augmented to enrich the sample diversity and reduce the occurrence of overfitting. The augmentation process can include data augmentation methods such as flipping, scaling, stitching, etc., or other data augmentation methods, which are not limited herein. For example: When the target object is a pantograph, since the positions of the pantograph and other components in the pantograph monitoring video are relatively fixed and there is a problem of unbalanced sample quantity in some scenarios, overfitting is likely to occur. Therefore, the training data set samples can be augmented to increase the quantity of the training data set samples.

[0095] Further, after the initial training data set samples are augmented, a labeling tool can be used to label the target objects in the multiple frame image samples of the initial training data set samples. For example: The pantograph area in the image can be labeled, and the labeled area is the pantograph head area and the pantograph base area that have not been soiled. After labeling, the label is saved in a preset format, for example, it can include the Visual Object Classes (VOC) format. After all the initial training data set samples are labeled, the final training data set can be obtained.

[0096] According to an embodiment of the present disclosure, after the training data set is constructed, the initial image recognition model can be trained using the training data set to obtain the target image recognition model. For example, the initial image recognition model can be trained based on a target loss function to obtain the target image recognition model. Among them, the target loss function is a function that can be used to evaluate the difference between the prediction result and the real result of the model, and can adjust the internal parameters of the model by guiding the learning process of the model to reduce the prediction error of the model.

[0097] According to an embodiment of the present disclosure, specifically, the target loss function may be composed of a classification loss function and a regression loss function. Among them, the classification loss function may include a cross-entropy loss function, and the cross-entropy loss function may be used to measure the difference between the prediction result for the target object and the true label. The calculation formula of the cross-entropy loss function may be as shown in Equation (2).

[0098] (2)

[0099] Among them, L cls represents the cross-entropy loss function, N represents the number of samples, y i represents the true label (0 or 1) of the i-th sample, and pi represents the prediction probability of the i-th sample.

[0100] According to an embodiment of the present disclosure, the regression loss function may include a Complete Intersection over Union (CIOU) function. The regression loss function may be used to evaluate and optimize the similarity and overlap degree between the predicted bounding box and the true bounding box. The calculation formula of the regression loss function may be as shown in Equation (3).

[0101] (3)

[0102] Among them, L CIOU represents the regression loss function, IOU represents the intersection over union of the predicted bounding box and the true bounding box, v represents the aspect ratio similarity factor of the bounding box, c represents the distance factor between the centers of the predicted bounding box and the true bounding box, and α and β are adjustment coefficients.

[0103] According to an embodiment of the present disclosure, the target loss function may include a weighted sum of the classification loss function and the regression loss function. The calculation formula of the target loss function may be as shown in Equation (4).

[0104] (4)

[0105] Among them, L represents the target loss function, λ1 represents the weight of the cross-entropy loss function, λ2 represents the weight of the regression loss function, L cls and L CIOU have the same meanings as above and will not be elaborated here.

[0106] According to an embodiment of the present disclosure, further, in order to make the image recognition model converge more easily and reduce the training time, the aspect ratio of the predicted bounding box may be adjusted. For example: when the target object is the pantograph head region and the base region, the aspect ratios of the predicted bounding boxes of the pantograph head region and the base region in the image are relatively fixed. In order to reduce the training time, the aspect ratios of the predicted bounding boxes may be 5:1 and 4:1.

[0107] According to an embodiment of the present disclosure, after adjusting the aspect ratio of the prediction box, an image recognition model can be used to recognize a target object in an image, and a confidence level corresponding to the image can be obtained. An example of the calculation formula for the confidence level is shown in the following formula (5).

[0108] (5)

[0109] In the above formula (5), C represents the confidence level, P(Object) represents the probability that a target object exists within the prediction box, and IOU(pred, gt) represents the intersection over union of the area of the prediction box and the area of the ground truth box.

[0110] Among them, the calculation formula for IOU(pred, gt) is shown in the following formula (6).

[0111] (6)

[0112] In the above formula (6), the meaning of IOU(pred, gt) is as described above and will not be elaborated here. box pred represents the area of the prediction box, and box gt represents the area of the ground truth box.

[0113] According to an embodiment of the present disclosure, obtaining multiple frames of images corresponding to multiple time windows includes: obtaining a video stream corresponding to a target object, where the video stream is obtained by using a target video acquisition device to capture the target object; performing frame division processing on the video stream to obtain an image queue; and screening out multiple frames of images from the image queue according to the running state information corresponding to the target object.

[0114] According to an embodiment of the present disclosure, a target video acquisition device can be used to capture the target object in real time to obtain a video stream corresponding to the target object. By performing frame division processing on the video stream, the video stream can be decomposed into single-frame images to obtain an image queue.

[0115] According to an embodiment of the present disclosure, if object recognition processing is directly performed on all frame images in an image queue, it will consume a large amount of computing resources and reduce computing efficiency. Therefore, multiple frame images can be screened out from the image queue according to the operating state information corresponding to the target object, and object recognition processing is only performed on the multiple frame images in the image queue. Specifically, the operating state information corresponding to the target object may include the operating state information of the target object itself, or may include the operating state information of the device on which the target object is installed. For example, when the target object is a pantograph, the operating state information may include the operating state information of the train on which the pantograph is installed, such as whether the train is in an operating state; the operating state information may also include the operating state information of the pantograph itself, such as whether it is in a raised state, or whether it is in an emergency lowering state, and the vehicle speed, etc. Multiple frame images corresponding to the train being in an operating state and the pantograph being in a raised state can be screened out from the image queue, and subsequent object recognition processing is performed on the above multiple frame images. For other images in the image queue except for the multiple frame images, no subsequent processing is required.

[0116] Figure 5 The flowchart of the method for determining the imaging quality of a video acquisition device according to another embodiment of the present disclosure is schematically shown.

[0117] As Figure 5 shown, the method includes operations S510 to S550.

[0118] In operation S510, multiple frame images corresponding to the pantograph are acquired.

[0119] For example, the pantograph can be photographed in real time by a camera to obtain a video stream corresponding to the pantograph. The video stream is frame-divided to obtain multiple frame images corresponding to the pantograph.

[0120] In operation S520, for each frame image, it can be determined whether the train is running at the image shooting moment. If the train is not running, it is temporarily not necessary to determine the imaging quality of the video acquisition device for the pantograph.

[0121] In operation S530, if it is determined that the train is running, for each frame image, it can be determined whether there is a raising signal at the image shooting moment. If there is no raising signal, it is temporarily not necessary to determine the imaging quality of the video acquisition device for the pantograph.

[0122] In operation S540, if it is determined that there is a raising signal, for each frame image, it can be determined whether there is an emergency lowering signal at the image shooting moment. If there is an emergency lowering signal, it is temporarily not necessary to determine the imaging quality of the video acquisition device for the pantograph.

[0123] In operation S550, in the case where it is determined that there is no emergency lowering signal, at least one target image is determined from multiple frames of images.

[0124] According to an embodiment of the present disclosure, in the case where the target object is a pantograph, the operating state corresponding to the pantograph can be determined. Specifically, the shooting time of each frame of image can be determined, and the operating state corresponding to the pantograph at the shooting time can be determined, where the operating state corresponding to the pantograph can include: whether the train where the pantograph is located is running, whether there is a raising signal, and whether there is an emergency lowering signal. For each frame of picture, in the case where its shooting time corresponds to the train running, there is a raising signal and there is no emergency lowering signal, this frame of image is determined as the target image.

[0125] Specifically, when the pantograph is not raised, or when the train is not running, the pantograph is not in a working state, and even if the image quality of the picture is poor, it will not directly affect the contact state between the pantograph and the catenary. Therefore, there is no need to further identify and process the images in the non-raised state. Correspondingly, in the case where there is an emergency lowering signal, the pantograph may have been separated from the catenary, so there is no need to further identify and process the images in the non-raised state either.

[0126] According to an embodiment of the present disclosure, a large amount of image data is generated during the running of the train. If all images are detected, it will significantly increase the computing burden and resource consumption of the system. By screening out some images from all images based on the operating state information corresponding to the target object and only performing subsequent processing on some images, the computing burden of the system can be reduced and resource consumption can be saved.

[0127] In operation S560, the confidence levels corresponding to at least one target image are determined to obtain a target confidence level queue. For example: The target objects in each frame of target image can be recognized through an image recognition model to obtain the confidence levels corresponding to each frame of image. The target confidence level queue can include the confidence levels corresponding to each frame of target image.

[0128] In operation S570, at least one target confidence level is determined by using a preset sliding window. For example: At least one target confidence level can be screened out from the target confidence level queue by using a preset sliding window.

[0129] In operation S580, based on at least one target confidence level, the dynamic mean value of the confidence levels is calculated. The dynamic mean value of the confidence levels can be used to characterize the imaging quality of the video acquisition device. For example, when the dynamic mean value of the confidence levels is high, it indicates that the imaging quality of the video acquisition device is high.

[0130] Figure 6 A block diagram of a device for determining the imaging quality of a video acquisition device according to an embodiment of the present disclosure is schematically shown.

[0131] As Figure 6 shown, a determining device 600 for the imaging quality of a video acquisition device includes an obtaining module 610, an identifying module 620, a first determining module 630, a calculating module 640, and a second determining module 650.

[0132] The obtaining module 610 is configured to obtain multiple frames of images corresponding to multiple time windows, where the multiple frames of images are obtained by using a target video acquisition device to capture a target object. In one embodiment, the obtaining module 610 may be configured to perform the operation S210 described above, which will not be elaborated herein.

[0133] The identifying module 620 is configured to perform target object identification processing on the multiple frames of images by using a pre-constructed image recognition model, and obtain multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the image recognition model for the target object in the image. In one embodiment, the identifying module 620 may be configured to perform the operation S220 described above, which will not be elaborated herein.

[0134] The first determining module 630 is configured to determine at least one target confidence level from the multiple confidence levels by using a preset sliding window. In one embodiment, the first determining module 630 may be configured to perform the operation S230 described above, which will not be elaborated herein.

[0135] The calculating module 640 is configured to calculate a dynamic mean value of the confidence levels based on the at least one target confidence level. In one embodiment, the calculating module 640 may be configured to perform the operation S240 described above, which will not be elaborated herein.

[0136] The second determining module 650 determines the imaging quality of the target video acquisition device by using the dynamic mean value of the confidence levels. In one embodiment, the second determining module 650 may be configured to perform the operation S250 described above, which will not be elaborated herein.

[0137] According to an embodiment of the present disclosure, the first determining module includes a first determining sub-module, an adjusting sub-module, and a second determining sub-module.

[0138] The first determining sub-module is configured to determine environmental information corresponding to the target video acquisition device; the adjusting sub-module is configured to adjust the window width of the preset sliding window based on the environmental information; the second determining sub-module determines at least one target confidence level from the multiple confidence levels according to the window width; where the environmental information includes at least one of the following: weather information, optical environmental information, and operating environmental information.

[0139] According to an embodiment of the present disclosure, the first determining module includes an adding sub-module, an adjusting sub-module, and a third determining sub-module.

[0140] The adding sub-module is used to add the confidence corresponding to at least one frame of the current image to the initial confidence queue to obtain an updated confidence queue, where the initial confidence queue includes at least one confidence corresponding to at least one frame of historical images; the adjustment sub-module is used to adjust the positions of multiple confidences in the updated confidence queue according to the timestamps corresponding to the confidences to obtain a target confidence queue; the third determination sub-module is used to determine at least one target confidence from the target confidence queue by using a preset sliding window.

[0141] According to an embodiment of the present disclosure, the third determination sub-module includes a determination unit. The determination unit is used to determine at least one target confidence from the predetermined positions of the updated confidence queue by using a preset sliding window according to the timestamp order in the target confidence queue.

[0142] According to an embodiment of the present disclosure, the device for determining the imaging quality of a video acquisition device further includes a third determination module and an obtaining module.

[0143] The third determination module is used to determine a target loss function based on a classification loss function and a regression loss function; the obtaining module is used to train an initial image recognition model based on the target loss function to obtain an image recognition model.

[0144] According to an embodiment of the present disclosure, the obtaining module includes an obtaining sub-module, a frame processing sub-module, and a screening sub-module.

[0145] The obtaining sub-module is used to obtain a video stream corresponding to a target object, where the video stream is obtained by using a target video acquisition device to shoot the target object; the frame processing sub-module is used to perform frame processing on the video stream to obtain an image queue; the screening sub-module is used to screen out multiple frames of images from the image queue according to the operation state information corresponding to the target object.

[0146] Any of a plurality of modules, sub-modules, units, and sub-units according to embodiments of the present disclosure, or at least part of the functions of any of the foregoing may be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure may be at least partially implemented as a computer program module, which may perform corresponding functions when the computer program module is run.

[0147] For example, any of a plurality of the acquisition module 610, the recognition module 620, the first determination module 630, the calculation module 640, and the second determination module 650 may be combined and implemented in one module / unit / sub-unit, or any one of the foregoing module / unit / sub-unit may be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units may be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to embodiments of the present disclosure, at least one of the acquisition module 610, the recognition module 620, the first determination module 630, the calculation module 640, and the second determination module 650 may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of them. Alternatively, at least one of the acquisition module 610, the recognition module 620, the first determination module 630, the calculation module 640, and the second determination module 650 may be at least partially implemented as a computer program module, which may perform corresponding functions when the computer program module is run.

[0148] It should be noted that the data processing system part in the embodiments of the present disclosure corresponds to the data processing method part in the embodiments of the present disclosure. For the description of the data processing system part, reference may be made to the data processing method part, and details are not described herein again.

[0149] Figure 7 A block diagram of an electronic device suitable for implementing the method described above according to an embodiment of the present disclosure is schematically shown. Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0150] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0151] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in one or more memories.

[0152] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom can be installed into the storage section 708 as needed.

[0153] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0154] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiment; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0155] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, device, or device.

[0156] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.

[0157] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program includes program codes for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program codes are used to cause the electronic device to implement the method for determining the imaging quality of the video acquisition device provided by the embodiment of the present disclosure.

[0158] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, apparatus, module, unit, etc. may be implemented by computer program modules.

[0159] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0160] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure can be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in a variety of ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0162] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A method for determining the imaging quality of a video acquisition device, comprising: Obtaining multiple frames of images corresponding to multiple time windows, where the multiple frames of images are obtained by using a target video acquisition device to capture a target object; Using a pre-constructed target image recognition model to perform target object recognition processing on the multiple frames of images, obtaining multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the target object in the image by the target image recognition model; Using a preset sliding window to determine at least one target confidence level from the multiple confidence levels; Calculating a dynamic mean of the confidence levels based on the at least one target confidence level; Using the dynamic mean of the confidence levels to determine the imaging quality of the target video acquisition device.

2. The method according to claim 1, wherein, The step of using a preset sliding window to determine at least one target confidence level from the multiple confidence levels includes: Determining environmental information corresponding to the target video acquisition device; Adjusting the window width of the preset sliding window based on the environmental information; Determining the at least one target confidence level from the multiple confidence levels according to the window width; Wherein the environmental information includes at least one of the following: weather information, optical environmental information, operating environmental information.

3. The method according to claim 1, wherein, The multiple time windows include a current time window and at least one historical time window, and the multiple frames of images include at least one current image corresponding to the current time window and at least one historical image corresponding to the at least one historical time window; The step of using a preset sliding window to determine at least one target confidence level from the multiple confidence levels includes: Adding the confidence levels corresponding to the at least one current image to an initial confidence level queue to obtain an updated confidence level queue, where the initial confidence level queue includes at least one confidence level corresponding to the at least one historical image; Adjusting the positions of the multiple confidence levels in the updated confidence level queue according to the time stamps corresponding to the confidence levels to obtain a target confidence level queue; Using the preset sliding window to determine the at least one target confidence level from the target confidence level queue.

4. The method according to claim 3, wherein, The step of using the preset sliding window to determine at least one target confidence level from the target confidence level queue includes: Determining the at least one target confidence level from a predetermined position of the updated confidence level queue by using the preset sliding window according to the time stamp order in the target confidence level queue.

5. The method according to claim 1, wherein The method further includes: Determining a target loss function based on a classification loss function and a regression loss function; Training an initial image recognition model based on the target loss function to obtain the target image recognition model.

6. The method according to claim 1, wherein The step of obtaining multiple frames of images corresponding to multiple time windows includes: Obtaining a video stream corresponding to the target object, where the video stream is obtained by using the target video acquisition device to capture the target object; Performing frame splitting processing on the video stream to obtain an image queue; Screening out the multiple frames of images from the image queue according to the operating state information corresponding to the target object.

7. An apparatus for determining the imaging quality of a video acquisition device, comprising: An acquisition module, configured to acquire multiple frames of images corresponding to multiple time windows, where the multiple frames of images are obtained by using a target video acquisition device to capture a target object; An identification module, configured to perform target object identification processing on the multiple frames of images by using a pre-constructed target image recognition model, to obtain multiple confidence levels corresponding to the multiple frames of images, where the confidence level is used to characterize the recognition accuracy of the target object in the image by the target image recognition model; A first determination module, configured to determine at least one target confidence level from the multiple confidence levels by using a preset sliding window; A calculation module, configured to calculate a dynamic mean of the confidence levels based on the at least one target confidence level; A second determination module, configured to determine the imaging quality of the target video acquisition device by using the dynamic mean of the confidence levels.

8. A train, comprising: A video acquisition device; The imaging quality of the video acquisition device is determined by the method according to any one of claims 1 to 6.

9. An electronic device, comprising: One or more processors; A memory, configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, having stored thereon executable instructions, which when executed by a processor cause the processor to implement the method according to any one of claims 1 to 6.