Lens picture quality analysis method based on video image
By performing significance detection and dynamic threshold adjustment of video images, identifying picture quality indicators of key semantic objects, and generating local and global picture quality evaluations, the limitations of lens picture quality analysis in the existing technology are solved, and accurate evaluation and fault positioning of the local quality of lens picture are achieved.
Patent Information
- Application Number
- CN202510798305.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing lens picture quality analysis technology has local mass distortion, insufficient adaptability to dynamic scenes, weak semantic information utilization, and lack of spatial positioning capabilities for specific defect areas of the lens, resulting in high evaluation blind spots and operation and maintenance costs.
By receiving continuous video frames, visual significance detection is performed to divide dynamic and static significance areas, identify and track the mask area of key semantic objects, obtain picture quality indicators, and dynamically adjust the threshold according to the object type and significance area, generate local display quality scores, and finally generate the global picture quality evaluation results of the lens, and trigger alarms and defect position marking when local abnormalities are found.
It realizes accurate evaluation of the local quality of the lens picture, breaks through the limitations of global evaluation, provides accurate decision-making basis, and improves the meticulousness of the evaluation and troubleshooting efficiency.
Smart Images

Figure CN120298416A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and relates to a method for analyzing the quality of lens images based on video images. Background Art
[0002] In the era of rapid information dissemination, videos have become an important way for people to obtain information and are widely used in fields such as daily shooting, film and television production, security monitoring, and virtual reality. High-quality lens images are the key to enhancing the visual effect and directly affect the user experience and application effect. For example, in security monitoring, clear video images are the basis for accurate event recognition and analysis; in film and television production, high-quality images are the prerequisite for creating excellent works; in virtual reality and augmented reality, the lens image quality determines the immersion and interactivity of the virtual environment.
[0003] However, in actual shooting, image quality problems often occur due to environmental factors and the limitations of the lens itself. Therefore, it is particularly important to develop an effective method for analyzing the quality of lens images to ensure the quality of video images and meet diverse application requirements.
[0004] In existing lens image quality analysis technologies, there are already some mature solutions. For example, a method and system for analyzing the quality of lens images based on an image recognition algorithm, with the Chinese patent publication number CN117094965A, which obtains three frames of images of an image capture device at different preset times, performs fast feature point calculation on the first two frames of images, connects background feature points to form line segments according to the minimum proximity principle, and constructs an evaluation function to obtain an evaluation result for the image captured by the image capture device, and can evaluate the image quality of the lens in real time and accurately.
[0005] Another method, device, electronic device, storage medium, and product for detecting image quality with the Chinese patent publication number CN118747747A, which obtains the current frame image of a camera to be detected and determines its scene category, performs grayscale processing on the image according to the grayscale parameter corresponding to the scene, and then performs quality detection on the grayscale image to determine the image quality result. Abandoning multi-frame time-series analysis and relying only on a single frame of image to avoid interference from scene changes, it is applicable to general scenarios and is beneficial to improving the detection efficiency and accuracy.
[0006] However, existing lens image quality analysis technologies still have limitations, specifically: 1. Existing technologies generally use global image statistical features as the core evaluation basis, combined with preset fixed thresholds or coarse-grained scene classification for lens image quality determination, which is likely to cause problems such as local quality distortion and missed detection, insufficient adaptability to dynamic scenes, and weak utilization of semantic information.
[0007] 2. The existing technology lacks the ability to spatially locate specific defect areas of the lens, which may lead to inefficient fault troubleshooting and increased risk of quality misjudgment, further resulting in a double deterioration of the evaluation blind spot and operation and maintenance costs. Summary of the Invention
[0008] In view of this, to solve the problems proposed in the above background technology, a method for analyzing the quality of the lens screen based on video images is proposed.
[0009] The object of the present invention can be achieved through the following technical solutions: The present invention provides a method for analyzing the quality of the lens screen based on video images, including: receiving continuous video frame input, and performing visual saliency detection to divide the dynamic and static saliency regions in each frame of the image.
[0010] Identify and track the mask regions of each key semantic object in the continuous video frames, and obtain the screen quality indicators of the mask regions. The screen quality indicators at least include sharpness, noise level, contrast, and motion blur sensitivity.
[0011] Set differentiated benchmark compliance thresholds for each screen quality indicator according to the type of key semantic object, and dynamically adjust the thresholds based on the type of saliency region to which the mask region belongs.
[0012] Generate a single-frame display quality score based on the preset compliance threshold comparison rule, synthesize multiple-frame scores, and calibrate through the time fluctuation variance to generate the local display quality scores of each key semantic object in the continuous video frames.
[0013] Synthesize multiple local display quality scores to generate the global screen quality evaluation result of the lens, and trigger an alarm and mark the defect position of the lens display when the local score is lower than the preset score threshold.
[0014] Compared with the existing technology, the beneficial effects of the present invention are as follows: (1) By tracking the mask regions of key semantic objects, the present invention dynamically sets the compliance threshold standards in combination with the object type and the status of the saliency region to which it belongs. On this basis, after generating a single-frame score according to the preset rules, the local object-level display quality score is output through calibration of the time fluctuation variance of multiple frames, breaking through the limitation of the existing technology that only focuses on the overall picture quality, refining the evaluation to the semantic unit level, fully considering the characteristics of different objects and the visual focus of the picture, and helping to more accurately capture the local quality status of the picture.
[0015] (2) The present invention generates the global screen quality evaluation result of the lens by fusing the local object-level display quality scores, and triggers an alarm for marking the defect position of the lens display when there is a local anomaly, effectively developing the linkage mechanism between the local display quality score and the lens defect position marking, providing a precise decision-making basis for lens maintenance. Description of the Drawings
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of a method for analyzing the quality of a lens picture based on video images provided in the first embodiment of the present invention.
[0018] Figure 2 It is an example diagram of the specific execution logic of visual saliency detection provided in the first embodiment of the present invention.
[0019] Figure 3 It is a schematic diagram of the hardware structure of a computer provided in the second embodiment of the present invention. Detailed implementation manners
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0021] Embodiment 1
[0022] Please refer to Figure 1 As shown, a method for analyzing the quality of a lens picture based on video images provided in the first embodiment of the present invention includes steps 110 to 150: In step 110, continuous video frame input is received, and visual saliency detection is performed to divide the dynamic and static saliency regions in each frame of the image.
[0023] As Figure 2 shown, in a preferred embodiment of the present invention, the visual saliency detection refers to the following execution process: The input continuous video frames are standardized, and the video frame images are divided into each grid unit.
[0024] It should be added that the above standardization process includes image normalization, noise suppression, and illumination correction. Among them, image normalization linearly maps the pixel values of consecutive video frames to a standard range (such as [0, 1] or [0, 255]) to eliminate the influence of illumination differences on saliency detection. Noise suppression uses non-local means or bilateral filtering algorithms to eliminate random noise while retaining edge details. Illumination correction is obtained by histogram equalization and is used to compensate for local overexposure or underexposure caused by uneven illumination or backlighting. It should be specifically noted that noise suppression and illumination correction in the standardization process here will affect the acquisition accuracy of the picture quality index in the subsequent key semantic object mask area, so it is only used in the visual saliency detection process.
[0025] Extract the spatial features of each grid unit in the video frame image. The spatial features include color contrast, texture gradient, and edge intensity. Identify static saliency units based on the stability analysis of multi-frame spatial features.
[0026] It should be added that the above spatial feature extraction process involves calculations. Exemplarily, color contrast refers to the average Euclidean distance in the Lab color space between a grid unit and its adjacent domain units. Texture gradient is the gradient amplitude calculated by extracting multi-scale texture responses through a Gabor filter bank. Edge intensity is the proportion of edge pixels within a unit counted by a Canny edge detector.
[0027] The above process of identifying static saliency units based on the stability analysis of multi-frame spatial features specifically includes the following steps: Construct a sequence of spatial feature vectors for each grid unit in consecutive video frames, quantify the difference degree between each element in the spatial feature vector sequence and the mean of the sequence vectors one by one, obtain the spatial feature difference degree of each grid unit in consecutive video frames through mean calculation, use the difference between 1 and the spatial feature difference degree as the stability index, and thus obtain the spatial feature stability index of each grid unit in consecutive video frames. If the spatial feature stability index of a certain grid unit is greater than the preset spatial feature stability index compliance threshold (set artificially based on the requirements of the camera monitoring scenario, for example, 0.8 in the security monitoring scenario), then regard this grid unit as a static saliency unit.
[0028] It should be noted that the specific quantification process of the difference degree between each element in the spatial feature vector sequence and the mean of the sequence vectors is as follows: Quantify the deviation amplitude between each element in the spatial feature vector sequence and the mean of the sequence vectors through the Euclidean distance, use the deviation amplitude as the numerator, use the largest spatial feature vector in the spatial feature vector sequence as the denominator, perform normalization processing through ratio operation, and further use the operation result as the difference degree.
[0029] Extract the temporal features of each grid cell in the video frame image. The temporal features are obtained through inter-frame motion vector analysis and include the magnitude of the motion vector, the direction angle of the motion vector, and the cumulative displacement of the motion. When the magnitude of the motion vector in multiple consecutive frames exceeds the preset magnitude standard and there is a stable motion trajectory, it is determined as a dynamically significant unit.
[0030] It should be noted that the quantization standard for the above stable motion trajectory is as follows: the calculated result of the standard deviation of the direction angles of the motion vectors in consecutive video frames is less than or equal to the preset angle threshold, and the difference in the cumulative displacement of the motion between each frame and its previous frame is greater than or equal to the preset reference displacement. The preset angle threshold, the preset reference displacement, and the above preset magnitude standard are all calibrated through statistical experiments in typical scenarios.
[0031] Merge adjacent statically significant units or adjacent dynamically significant units to construct statically significant regions and dynamically significant regions, and perform morphological closing operations to smooth the region boundaries and fill internal holes.
[0032] Output the coordinate masks and classification labels of the dynamic and static significant regions in each frame image.
[0033] In step 120, identify and track the mask regions of each key semantic object in the consecutive video frames, and obtain the picture quality indicators of the mask regions. The picture quality indicators at least include sharpness, noise level, contrast, and motion blur sensitivity.
[0034] It should be added that the above method for obtaining the picture quality indicators includes: applying a gradient operator (such as Sobel) to the pixels in the mask region and calculating the average value of the pixel gradients to obtain the sharpness.
[0035] Quantify the noise level by analyzing the standard deviation of the pixel values in the smoothed region of the mask region image.
[0036] Quantify the contrast by calculating the dynamic range of the pixel values in the mask region (the difference between the maximum and minimum pixel values / the sum of the maximum and minimum pixel values).
[0037] Apply Canny edge detection to the mask region image, extract the edge pixels and calculate the gradient direction histogram of the edge pixels, determine the main motion direction, and calculate the diffusion width of the edge pixels along this motion direction as the motion blur sensitivity.
[0038] In addition, the picture quality indicators may also include color saturation, exposure, texture complexity, etc.
[0039] Special note: The supplementary calculation methods of the picture quality indicators are all existing technologies in the field of image processing and are widely used in image quality assessment, video coding optimization, and computer vision tasks. Therefore, no in-depth analysis will be carried out here.
[0040] In a preferred embodiment of the present invention, the mask regions of each key semantic object in the continuous video frames are as follows in the recognition and tracking process: Detect each key semantic object in each frame of image through a preset deep learning model, and output the bounding box, class label, and pixel-level binary mask of each key semantic object, where the pixels belonging to the key semantic object in the pixel-level binary mask are marked as 1, and the remaining pixels are marked as 0.
[0041] Assign a unique tracking identifier to the mask region of each key semantic object, and associate the same key semantic object in adjacent frames through feature matching.
[0042] It should be added that the above feature matching is based on the appearance features and motion features of the key semantic object, specifically including: Extracting the color histogram, texture features, and shape features of the mask region as appearance features, and calculating the similarity score of the relative appearance features between the mask regions of the front and back video frames (the calculation logic can be exemplified as: using the Bhattacharyya distance to measure the similarity of the color histogram, using the cosine similarity to measure the similarity of the texture features, using the Euclidean distance to measure the similarity of the shape features, and performing weighted fusion on the three types of feature similarity measurement values to obtain the similarity score of the appearance features, where the preset weights corresponding to the color histogram, texture features, and shape features are the same and the cumulative value is 1, for example, all are )
[0043] Extract the motion vector and trajectory consistency of the mask region as motion features, predict the position category of the mask region in the adjacent frame based on the assumption of brightness constancy between adjacent frames, and quantify the similarity score of the relative motion features between the mask regions of the front and back video frames based on the prediction displacement amount matching degree and the prediction trajectory consistency.
[0044] Take the sum of the similarity scores of the appearance features and the motion features as the overall matching score. For each mask region in the current frame, find the candidate region with the highest overall matching score in the next frame. If the overall matching score reaches the preset score standard, it is determined as the same key semantic object and the same tracking identifier is assigned, otherwise a new tracking identifier is assigned.
[0045] Dynamically update the position coordinates of the mask region of the key semantic object in each frame of image, and form a sequence of object motion trajectories between consecutive frames to complete cross-frame tracking.
[0046] In a preferred embodiment of the present invention, the preset deep learning model is trained and generated based on an annotated image dataset containing various key semantic objects, where the annotated image dataset includes the bounding box, class label, and pixel-level mask information of the object, and a multi-task loss function is used to optimize the model parameters. The multi-task loss function includes classification loss, bounding box regression loss, and mask generation loss.
[0047] In step 130, differential baseline compliance thresholds are set for each picture quality metric according to the critical semantic object type, and the thresholds are dynamically adjusted based on the saliency region type to which the mask region belongs.
[0048] In a preferred embodiment of the present invention, the dynamic threshold adjustment includes: if the mask region belongs to a static saliency region, the compliance threshold of the display quality metric is tightened based on the multi-frame stability score, and the amplitude of tightening the compliance threshold of the display quality metric is inversely proportional to the stability score.
[0049] If the mask region belongs to a dynamic saliency region, the compliance threshold of the display quality metric is relaxed based on the magnitude of the motion vector, and the amplitude of relaxing the compliance threshold of the display quality metric is proportional to the magnitude of the motion vector.
[0050] It should be noted that the stability score of the above mask region specifically refers to the calculated mean of the spatial feature stability indicators of each grid cell included in the mask region in consecutive video frames, and the magnitude of the motion vector of the mask region specifically refers to the hierarchical calculated mean of the magnitudes of the motion vectors of each grid cell included in the mask region in each video frame (divided into two layers, that is, the mean calculation process of the magnitudes of the motion vectors is performed in the dimension of each video frame and the dimension of each grid cell in sequence).
[0051] It should also be noted that the tightening of the compliance threshold of the display quality metric based on the multi-frame stability score specifically means taking the difference between 1 and the stability score of the mask region as the tightening ratio, and taking the product of the tightening ratio and the preset baseline tightening amplitude as the actual tightening amplitude. Relaxing the compliance threshold of the display quality metric based on the magnitude of the motion vector specifically means taking the ratio of the magnitude of the motion vector of the mask region to the preset maximum reference motion vector magnitude as the relaxation ratio, and taking the product of the relaxation ratio and the preset baseline relaxation amplitude as the actual relaxation amplitude.
[0052] If the mask region belongs to a non-saliency region, the compliance threshold of the display quality metric is maintained.
[0053] In step 140, a single-frame display quality score is generated based on a preset compliance threshold comparison rule, and the multi-frame scores are integrated and calibrated through the time fluctuation variance to generate the local display quality scores of each critical semantic object in the consecutive video frames.
[0054] In a preferred embodiment of the present invention, the preset compliance threshold comparison rule is as follows: for the mask region of the critical semantic object in a single-frame image, the actual monitored value of each picture quality metric is compared item by item with its corresponding compliance threshold to determine whether each picture quality metric meets the standard. If a certain picture quality metric is determined to meet the standard, the preset full score value is assigned to this picture quality metric. If a certain picture quality metric is determined not to meet the standard, the preset full score value of this picture quality metric is deducted proportionally according to the deviation between the actual monitored value of this picture quality metric and its corresponding compliance threshold.
[0055] It should be noted that the conditions for determining whether the above-mentioned picture quality indicators meet the standards are as follows: the actual monitored value needs to satisfy a specific directional relationship with the passing threshold, that is, according to the nature of the indicator, the actual value should be greater than or less than its passing threshold.
[0056] It should also be noted that the specific process of deducting the full score value preset for the picture quality indicator proportionally according to the deviation between the actual monitored value of the picture quality indicator and its corresponding passing threshold includes: performing a ratio operation on the absolute difference between the actual monitored value of the picture quality indicator and its corresponding passing threshold and the preset maximum deviation value to obtain the deviation ratio between the actual monitored value of the picture quality indicator and its corresponding passing threshold, and taking the product of the score deducted corresponding to the preset unit deviation ratio and the actually calculated deviation ratio as the actual deducted score, and thus performing a deduction process on the full score value preset for the picture quality indicator.
[0057] Accumulate the scores of each picture quality indicator to obtain the display quality score of the key semantic object mask area in a single-frame image, that is, the single-frame display quality score.
[0058] In a preferred embodiment of the present invention, the comprehensive multi-frame scoring and calibration through time fluctuation variance are as follows: arrange the single-frame display quality scores of each key semantic object in consecutive video frames in chronological order to form a scoring sequence.
[0059] For the scoring sequence, introduce a time decay factor and a variance sensitivity factor respectively to generate a reference scoring item and a fluctuation penalty item for the display quality of the key semantic object within the global time range, and further take the product of the reference scoring item and the fluctuation penalty item as the calibrated local display quality score, thereby generating the local display quality scores of each key semantic object in consecutive video frames.
[0060] As an example, the specific analysis process of the local display quality scores of each key semantic object in the above-mentioned consecutive video frames can be referred to the following formula: , where is the reference scoring item calculated based on time decay weighted average, is the display quality score of the th frame of the th key semantic object in consecutive video frames, is the number of each key semantic object, , is the number of each frame arranged in chronological order in consecutive video frames, and can also be used as a numerical value representing the number to participate in the formula calculation, , is the time decay weight, used to assign higher weights to recent frames, reflecting the timeliness of quality changes, and its specific expression is , is the preset time decay rate, which is used to control the weight decay rate. is the total number of consecutive video frames. In terms of the calculation of the scoring item, its mathematical meaning lies in the weighted sum of the numerator and the normalization of the denominator weights, ensuring the stability of the scoring range.
[0061] is the fluctuation penalty term. is the calculation variance of the relative single-frame display quality score in the scoring sequence of the key semantic object. is the preset reference variance, which can take the global variance mean or a fixed theoretical value. is the preset variance sensitivity coefficient, which is used to adjust the variance penalty intensity. The fluctuation penalty term uses an exponential function to enhance the penalty effect of high variance, fitting the non-linear perception sensitive to the fluctuation effect of the lens quality, and controlling the penalty ratio interval through the mathematical properties of the exponential function combined with the preset variance sensitivity coefficient to ensure that there is no excessive penalty.
[0062] Suppose the total number of consecutive video frames is 5, the preset time decay rate and the preset variance sensitivity coefficient are both 0.5, the preset reference variance is 100, and the scoring sequence of a certain key semantic object is Then, through calculation, it can be known that 、 、 、 、 According to the calculation of the numerator of the scoring item, the result is 165.295, the result of the denominator calculation is 2.332, the calculated value of the scoring item is 70.9, the calculation variance of the relative single-frame display quality score of the scoring sequence of this key semantic object is 116, the calculated value of the fluctuation penalty term is 0.56, and the local display quality score of this key semantic object in the final consecutive video frames is 39.704.
[0063] In the embodiment of the present invention, by tracking the masked area of the key semantic object and dynamically setting the compliance threshold standard in combination with the object type and the status of the significant area to which it belongs, after generating the single-frame score according to the preset rules, the local object-level display quality score is output through multi-frame time fluctuation variance calibration, breaking through the limitation of the prior art that only focuses on the overall picture quality, refining the evaluation to the semantic unit level, fully considering the characteristics of different objects and the visual focus of the picture, and helping to more accurately capture the local quality status of the picture.
[0064] In step 150, the global picture quality evaluation result of the lens is generated by comprehensively considering multiple local display quality scores, and an alarm and the marking of the lens display defect position are triggered when the local score is lower than the preset score threshold.
[0065] In a preferred embodiment of the present invention, the global picture quality evaluation result of the lens refers to the weighted fusion value of the local display quality scores of each key semantic object in the consecutive video frames.
[0066] It should be noted that the above weighted fusion aims to reflect the importance differences of different key semantic objects in the monitoring scenario through dynamic weight allocation, so as to more accurately evaluate the overall picture quality. Assuming that the key semantic objects include people, vehicles, and packages, then in the security monitoring scenario, to ensure the priority of the picture quality in the area where people are active, the weight allocation should be people > vehicles > packages; in the traffic monitoring scenario, to focus on the road traffic conditions, the weight allocation should be vehicles > people > packages.
[0067] In a preferred embodiment of the present invention, the lens display defect position marking includes the following content: Mark the key semantic objects with local display quality scores lower than the preset score threshold in consecutive video frames as abnormal objects.
[0068] Extract the display quality scores of the mask regions of each abnormal object in each frame of the image, locate the mask regions of the lowest score frame and the frames with a sudden drop in score of each abnormal object, integrate the mask regions of the lowest score frame and the frames with a sudden drop in score of each abnormal object and perform geometric intersection analysis, and output the minimum circumscribed rectangle coordinate sequence of the high-frequency overlapping region in a single frame of the image.
[0069] It should be added that the above frames with a sudden drop in score specifically refer to: Obtain the decrease rate of the display quality score of each frame and its previous frame for the mask region of the abnormal object in consecutive video frames, and retrieve the frame corresponding to the maximum decrease rate of the display quality score of the mask region of the abnormal object compared with its previous frame as the frame with a sudden drop in score.
[0070] It also needs to be added that the premise of the above geometric intersection analysis is that the camera position and perspective remain unchanged during video shooting, the coordinate systems of all frames are strictly aligned, the positions of the mask regions of abnormal objects in each frame are fixed in the image coordinate system, and the masks across frames or objects can be directly superimposed for analysis. Therefore, the common coverage region of the masks in the video frame map is retrieved through geometric intersection (union). If the number of repeated coverages of a certain mask common coverage region exceeds the preset number standard, it is regarded as the high-frequency overlapping region in a single frame of the image.
[0071] Based on the lens calibration parameters, map the minimum circumscribed rectangle coordinate sequence to the lens physical coordinates to perform the lens display defect position marking action.
[0072] In a preferred embodiment of the present invention, the execution process of the lens display defect position marking action adopts a streaming processing architecture, including: Parallelly execute the quality analysis calculation of the current input consecutive video frames and the lens display defect position marking action of the historical input consecutive video frames through a double-buffer mechanism, and the buffer size is dynamically allocated according to the video resolution and thread synchronization is achieved through a mutex.
[0073] Perform inter-frame differential coding compression on the minimum circumscribed rectangle coordinate sequence of the high-frequency overlapping region, record the change amount of the coordinates relative to the previous frame, and set the maximum allowable error.
[0074] Dynamically adjust the computational granularity of geometric intersection analysis according to the video frame rate. When the frame rate exceeds the preset frame rate threshold, downsample the coordinates of the mask region to grid block coordinates for approximate overlap detection.
[0075] In the embodiments of the present invention, a global picture quality evaluation result of the lens is generated by fusing local object-level display quality scores, and when a local anomaly occurs, an alarm for marking the position of the lens display defect is triggered, effectively developing a linkage mechanism between local display quality scores and lens defect position marking, providing an accurate decision-making basis for lens maintenance.
[0076] Embodiment 2
[0077] In the second embodiment of the present invention, the following technical solution is provided. A computer includes a memory 202, a processor 201, and a computer program stored on the memory 202 and executable on the processor 201. When the processor 201 executes the computer program, it implements a method for analyzing the picture quality of a lens based on video images as described above.
[0078] Specifically, the above-mentioned processor 201 may include a central processing unit (CPU), or a specific integrated circuit, or may be configured as one or more integrated circuits implementing the embodiments of the present application.
[0079] Among them, the memory 202 may include a mass storage for data or instructions. By way of example and not limitation, the memory 202 may include a hard disk drive, a floppy disk drive, a solid state drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus drive, or a combination of two or more of these. In a suitable case, the memory 202 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 202 may be inside or outside the data processing device. In a specific embodiment, the memory 202 is a non-volatile memory. In a specific embodiment, the memory 202 includes a read-only memory and a random access memory. In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM, an erasable PROM, an electrically erasable PROM, an electrically rewritable ROM, or a flash memory (FLASH), or a combination of two or more of these. In a suitable case, the RAM may be a static random access memory or a dynamic random access memory, where the DRAM may be a fast page mode dynamic random access memory, an extended data output dynamic random access memory, a synchronous dynamic random access memory, etc.
[0080] The memory 202 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 201.
[0081] The processor 201 implements the above-mentioned method for analyzing the quality of a lens image based on video images by reading and executing computer program instructions stored in the memory 202.
[0082] In some of these embodiments, the computer may further include a communication interface 203 and a bus 200. Among them, as Figure 3 shown, the processor 201, the memory 202, and the communication interface 203 are connected through the bus 200 and complete communication with each other.
[0083] The communication interface 203 is used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application. It can also implement data communication with other components, such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.
[0084] The bus 200 includes hardware, software, or both, and couples the components of the computer to each other. The bus 200 includes, but is not limited to, at least one of the following: data bus, address bus, control bus, expansion bus, local bus. By way of example and not limitation, the bus 200 may include a graphics acceleration interface or other graphics bus, an enhanced industry standard architecture bus, a front-side bus, a hypertransport interconnect, an industry standard architecture bus, a wireless bandwidth interconnect, a low pin count bus, a memory bus, a microchannel architecture bus, a peripheral component interconnect bus, a PCI-Express bus, a serial advanced technology attachment bus, a video electronics standards association local bus, or other suitable bus or a combination of two or more of these. In a suitable case, the bus 200 may include one or more buses. Although the embodiments of the present application describe and illustrate a specific bus, the present application contemplates any suitable bus or interconnect.
[0085] The above content is merely an example and illustration of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should fall within the protection scope of the present invention.
Claims
1. A method for analyzing the quality of a lens image based on video images, characterized in that The method includes: Receiving continuous video frame input, performing visual saliency detection to divide the dynamic and static saliency regions in each frame of the image; Identifying and tracking the mask regions of each key semantic object in the continuous video frames, and obtaining the picture quality indicators of the mask regions, where the picture quality indicators at least include sharpness, noise level, contrast, and motion blur sensitivity; Setting differentiated benchmark compliance thresholds for each picture quality indicator according to the key semantic object type, and dynamically adjusting the thresholds based on the saliency region type to which the mask region belongs; Generating a single-frame display quality score based on a preset compliance threshold comparison rule, integrating multi-frame scores and calibrating through the temporal fluctuation variance to generate the local display quality scores of each key semantic object in the continuous video frames; Integrating multi-local display quality scores to generate the global picture quality evaluation result of the shot, and triggering an alarm and marking the position of the shot display defect when the local score is lower than the preset score threshold.
2. The method for analyzing the quality of a lens image based on video images according to claim 1, wherein: The visual saliency detection refers to the following execution process: performing normalization processing on the input continuous video frames, and dividing the video frame images into each grid unit; Extracting the spatial features of each grid unit of the video frame image, where the spatial features include color contrast, texture gradient, and edge intensity, and identifying static saliency units based on the stability analysis of multi-frame spatial features; Extracting the temporal features of each grid unit of the video frame image, where the temporal features are obtained through inter-frame motion vector analysis and include motion vector amplitude, motion vector direction angle, and motion cumulative displacement. When the motion vector amplitudes of multiple consecutive frames exceed the preset amplitude standard and there is a stable motion trajectory, it is determined as a dynamic saliency unit; Merging adjacent static saliency units or adjacent dynamic saliency units to construct static saliency regions and dynamic saliency regions, and performing morphological closing operation to smooth the region boundaries and fill the internal holes; Outputting the coordinate masks and classification labels of the dynamic and static saliency regions in each frame of the image.
3. The method for analyzing the quality of a lens image based on video images according to claim 1, wherein: The mask regions of each key semantic object in the continuous video frames refer to the following identification and tracking process: detecting each key semantic object in each frame of the image through a preset deep learning model, and outputting the bounding box, class label, and pixel-level binary mask of each key semantic object, where the pixels belonging to the key semantic object in the pixel-level binary mask are marked as 1, and the remaining pixels are marked as 0; Assigning a unique tracking identifier to the mask region of each key semantic object, and associating the same key semantic object in adjacent frames through feature matching; Dynamically updating the position coordinates of the mask region of the key semantic object in each frame of the image to form an object motion trajectory sequence between consecutive frames to complete cross-frame tracking.
4. A method for analyzing the quality of a lens image based on video images according to claim 3, characterized in that: The preset deep learning model is trained and generated based on an annotated image dataset containing various key semantic objects, where the annotated image dataset includes the bounding box, class label, and pixel-level mask information of the object, and a multi-task loss function is used to optimize the model parameters, and the multi-task loss function includes classification loss, bounding box regression loss, and mask generation loss.
5. A method for analyzing the quality of a lens image based on video images according to claim 1, characterized in that: The dynamic adjustment threshold includes: if the masked area belongs to the static saliency area, tighten the passing threshold of the display quality index based on the multi-frame stability score, and the amplitude of tightening the passing threshold of the display quality index is inversely proportional to the stability score; if the masked area belongs to the dynamic saliency area, relax the passing threshold of the display quality index based on the magnitude of the motion vector, and the amplitude of relaxing the passing threshold of the display quality index is proportional to the magnitude of the motion vector; if the masked area belongs to the non-saliency area, keep the passing threshold of the display quality index.
6. The method for analyzing the quality of a lens image based on video images according to claim 1, wherein: The preset passing threshold comparison rule is as follows: for the masked area of the key semantic object in a single-frame image, compare the actual monitored value of each picture quality index with its corresponding passing threshold item by item to determine whether each picture quality index passes. If a certain picture quality index is determined to pass, assign the preset full score value to this picture quality index. If a certain picture quality index is determined not to pass, deduct the preset full score value of this picture quality index proportionally according to the deviation between the actual monitored value of this picture quality index and its corresponding passing threshold; Accumulate the scores of each picture quality index to obtain the display quality score of the masked area of the key semantic object in a single-frame image, that is, the single-frame display quality score.
7. A method for analyzing the quality of a lens image based on video images according to claim 6, characterized in that: The comprehensive multi-frame scoring and calibration through time fluctuation variance are as follows: arrange the single-frame display quality scores of each key semantic object in consecutive video frames in chronological order to form a scoring sequence; For the scoring sequence, introduce a time decay factor and a variance sensitivity factor respectively to generate a reference scoring item and a fluctuation penalty item for the display quality of the key semantic object within the global time range, and further use the product of the reference scoring item and the fluctuation penalty item as the calibrated local display quality score, thereby generating the local display quality scores of each key semantic object in consecutive video frames.
8. A method for analyzing the quality of a lens image based on video images according to claim 1, characterized in that: The global picture quality evaluation result of the lens refers to the weighted fusion value of the local display quality scores of each key semantic object in consecutive video frames.
9. A method for analyzing the quality of a lens image based on video images according to claim 6, characterized in that: The lens display defect position annotation includes the following: Mark the key semantic objects with local display quality scores lower than the preset scoring threshold in consecutive video frames as abnormal objects; Extract the display quality scores of the masked areas of each abnormal object in each frame of the image, locate the masked areas of the lowest scoring frame and the frame with a sudden drop in score of each abnormal object, integrate the masked areas of the lowest scoring frame and the frame with a sudden drop in score of each abnormal object and perform geometric intersection analysis, and output the minimum circumscribed rectangle coordinate sequence of the high-frequency overlapping area in a single-frame image; Based on the lens calibration parameters, map the minimum circumscribed rectangle coordinate sequence to the lens physical coordinates to perform the lens display defect position annotation action.
10. A method for analyzing the quality of a lens image based on video images according to claim 9, characterized in that: The execution process of the lens display defect position annotation action adopts a streaming processing architecture, including: parallelly execute the quality analysis calculation of the current input consecutive video frames and the lens display defect position annotation action of the historical input consecutive video frames through a double-buffer mechanism, and the buffer size is dynamically allocated according to the video resolution and thread synchronization is achieved through a mutex; Adopt inter-frame differential coding compression for the minimum circumscribed rectangle coordinate sequence of the high-frequency overlapping area, record the change amount of the coordinates relative to the previous frame, and set the maximum allowable error; Dynamically adjust the computational granularity of geometric intersection analysis according to the video frame rate. When the frame rate exceeds the preset frame rate threshold, downsample the coordinates of the mask region to grid block coordinates for approximate overlap detection.
Citation Information
Patent Citations
Lens picture quality analysis method and system based on image recognition algorithm
CN117094965A
Picture quality detection method and device, electronic equipment, storage medium and product
CN118747747A
Video picture quality evaluation method and device for automatic driving, equipment and storage medium
CN115439450A
User interest analysis method based on e-commerce big data
CN116805257A
User data efficient storage method and system for security evaluation system
CN118709004A
Cited By
Video dynamic quality evaluation and analysis method based on perception and memory
CN121305450A
Intelligent video acquisition and processing method for mine monitoring
CN121982619A
Video intelligent acquisition and processing method for mine monitoring
CN121982619B