Training method of discrimination model, method for discriminating breathing effect and device thereof

By training the respiratory effect discrimination model and using multi-dimensional features to discriminate the respiratory effect, the problems of inaccurate and highly subjective evaluation results in the existing technology are solved, and efficient and accurate respiratory effect testing is achieved.

CN113870186BActive Publication Date: 2025-09-09ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111016028.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-09-09
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

Existing breathing effect evaluation methods rely solely on clarity information, resulting in low evaluation results that are difficult to compare in real scenes. Subjective evaluation is time-consuming and labor-intensive, and the results vary from person to person.

Method used

By obtaining multi-dimensional features of video frames, such as structural similarity, brightness difference, texture difference, motion degree and average brightness, a breathing effect discrimination model is trained and multi-dimensional features are used for discrimination to reduce manual judgment errors.

Benefits of technology

The accuracy, stability and repeatability of respiratory effect testing are improved, and comparability and efficient testing are achieved in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870186B_ABST
    Figure CN113870186B_ABST
Patent Text Reader

Abstract

The present application provides a method for training a discriminant model, a method for discriminating a respiratory effect, and a device thereof. The training method includes: obtaining a video set to be trained, and the coding information and classification labels of the video frames in the video set; obtaining a continuous sampling frame group in the video set based on the coding information of all video frames, wherein the coding information of at least one video frame in the continuous sampling frame group is preset coding information; obtaining multi-dimensional features of the continuous sampling frame group; and training a respiratory effect discriminant model based on the continuous sampling frame group and its multi-dimensional features and classification labels. In the above manner, the training method of the present application trains the discriminant model through multi-dimensional respiratory effect influence indicators, so that the discrimination of the respiratory effect is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of camera imaging quality detection, and in particular to a training method for a discrimination model, a method for discriminating a breathing effect, and a device thereof. Background Art

[0002] With the widespread use of cameras, people are increasingly demanding higher quality images from surveillance cameras. The breathing effect introduced by image encoding and decoding algorithms has a serious impact on image stability. Therefore, objectively determining the severity of this breathing effect has become a pressing issue in camera image quality testing technology.

[0003] The current breathing effect evaluation method only relies on clarity information to measure the breathing effect of the test scene, without considering the characteristic information of other dimensions of the test scene, resulting in the accuracy of the evaluation results being inaccurate, and the single-dimensional evaluation indicators are difficult to be comparable in real scenes. Summary of the Invention

[0004] The present application provides a training method for a discriminant model, a method for discriminating respiratory effect, and a device thereof.

[0005] This application provides a training method for a discriminant model, the training method comprising:

[0006] Obtaining a video set to be trained, as well as encoding information and classification labels of video frames in the video set;

[0007] Obtaining a continuous sampling frame group in the video set based on the encoding information of all video frames, wherein the encoding information of at least one video frame in the continuous sampling frame group is preset encoding information;

[0008] Acquire multi-dimensional features of the continuous sampling frame group;

[0009] The breathing effect discrimination model is trained based on the continuous sampling frame group and its multi-dimensional features and classification labels.

[0010] The multi-dimensional features include at least four dimensional features of structural similarity, brightness difference, texture difference, motion degree, average brightness and average texture.

[0011] The step of obtaining a continuous sampling frame group in the video set based on the encoding information of all video frames includes:

[0012] Locating a position of a target video frame with preset coding information in the video set based on coding information of all video frames;

[0013] A group of continuous sampling frames is obtained based on the position of the target video frame in the video set.

[0014] The step of obtaining a continuous sampling frame group based on a position of the target video frame in the video set includes:

[0015] The position of the target video frame is used as the middle position, and the continuous sampling frame group is obtained according to the time sequence of the video set.

[0016] The number of video frames whose timing is before the target video frame in the continuous sampling frame group is the same as the number of video frames whose timing is after the target video frame.

[0017] The coding information of the video frame includes intra-frame coding and inter-frame coding, and the preset coding information is intra-frame coding.

[0018] The present application provides a method for determining a respiratory effect, the method comprising:

[0019] Get the video set to be judged;

[0020] Obtaining multi-dimensional features of the video collection;

[0021] The multi-dimensional features of the video collection are input into a pre-trained breathing effect discrimination model to obtain an output classification label.

[0022] The multi-dimensional features include at least four dimensional features of structural similarity, brightness difference, texture difference, motion degree, average brightness and average texture.

[0023] The present application also provides a terminal device, which includes an acquisition module, a sampling module, a feature module, and a training module;

[0024] The acquisition module is used to acquire a video set to be trained, as well as encoding information and classification labels of video frames in the video set;

[0025] The sampling module is configured to obtain a continuous sampling frame group in the video set based on the coding information of all video frames, wherein the coding information of at least one video frame in the continuous sampling frame group is preset coding information;

[0026] The feature module is used to obtain multi-dimensional features of the continuous sampling frame group;

[0027] The training module is used to train the breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels.

[0028] The present application also provides another terminal device, comprising a memory and a processor, wherein the memory is coupled to the processor;

[0029] The memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned training method of the discriminant model and / or the method for discriminating the breathing effect.

[0030] The present application also provides a computer storage medium, which is used to store program data. When the program data is executed by a processor, it is used to implement the above-mentioned training method of the discriminant model and / or the method for discriminating the respiratory effect.

[0031] The beneficial effects of the present application are as follows: the terminal device obtains a video set to be trained, as well as the encoding information and classification labels of the video frames in the video set; based on the encoding information of all video frames, a continuous sampling frame group is obtained in the video set, wherein the encoding information of at least one video frame in the continuous sampling frame group is preset encoding information; multi-dimensional features of the continuous sampling frame group are obtained; and a breathing effect discrimination model is trained based on the continuous sampling frame group and its multi-dimensional features and classification labels. In the above manner, the training method of the present application trains the discrimination model through multi-dimensional breathing effect influence indicators, making the discrimination of the breathing effect more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0033] Figure 1 This is a flow chart of an embodiment of a training method for a discriminant model provided in this application;

[0034] Figure 2 This is a structural diagram of an embodiment of a continuous sampling frame group provided by the present application;

[0035] Figure 3 This is a flow chart of an embodiment of a method for determining a respiratory effect provided by the present application;

[0036] Figure 4 This is a schematic structural diagram of an embodiment of a terminal device provided by this application;

[0037] Figure 5 is a structural diagram of another embodiment of the terminal device provided by this application;

[0038] Figure 6 This is a schematic structural diagram of another embodiment of the terminal device provided by the present application;

[0039] Figure 7It is a structural diagram of an embodiment of a computer storage medium provided by this application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0041] The main reason for the breathing effect is the difference in image distortion caused by the two coding modes, intra-frame coding (I-frame) and inter-frame coding (P-frame or B-frame), which creates a subjective perception of a sudden change in the morphology of the video frame. Existing testing schemes for the breathing effect are mostly subjective assessments. By visually observing a large number of videos in different scenes, the severity of the breathing effect is rated, for example, whether there is / is not a breathing effect, or whether the breathing effect is mild / normal / severe. This method is time-consuming and labor-intensive, and judging the severity of the breathing effect by visually observing the video screen is subjective. The test results vary from person to person, resulting in low accuracy. At the same time, because the evaluation environment and process cannot be reproduced, it is difficult to ensure stability and comparability.

[0042] In order to solve the problems of current breathing effect discrimination technology, the embodiment of the present application proposes a discrimination model training method to train a discriminant model with high accuracy to solve the problem of classifying the severity level of breathing effect in any real-life environment, so that the classification results are comparable in different environments, thereby improving the test efficiency, test accuracy, stability and repeatability of breathing effect.

[0043] Please refer to the following for details: Figure 1 , Figure 1 It is a flowchart of an embodiment of the training method of the discriminant model provided in this application.

[0044] The training method of the present application is applied to a terminal device, wherein the terminal device of the present application can be a server or a system comprising a server and a terminal device. Accordingly, the various components of the terminal device, such as the various units, subunits, modules, and submodules, can be all provided in the server or separately provided in the server and the terminal device.

[0045] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing distributed servers, or it can be implemented as a single software or software module, which is not specifically limited here. In some possible implementations, the action matching method of the embodiment of the present application can be implemented by a processor calling computer-readable instructions stored in a memory.

[0046] Specifically, if Figure 1 As shown, the training method of the discriminant model in the embodiment of the present application specifically includes the following steps:

[0047] Step S11: Obtain a video set to be trained, as well as encoding information and classification labels of video frames in the video set.

[0048] In an embodiment of the present application, a terminal device obtains a video set to be trained, wherein the video set to be trained may be derived from a packaged video file or a real-time video stream.

[0049] The terminal device further obtains a classification label of the video set, wherein the classification label of the video set can be a classification label of the entire video set, or a classification label of each video frame in the video set. The classification label can be qualified by experienced staff for the video set, and the classification label can reflect the classification result of the respiratory effect. For example, the classification label can be 0, 1, 2, 3, etc., where classification label 0 indicates no respiratory effect, classification label 1 indicates a mild level of respiratory effect, classification label 2 indicates a general level of respiratory effect, and classification label 3 indicates a severe level of respiratory effect. In other embodiments, the terminal device can use classification labels of other types or other values ​​to mark the respiratory effect of the video set, which are not listed here one by one.

[0050] The terminal device can also obtain the encoding information of each video frame in the video set. The video encoding information described in the embodiments of the present application is a sequence of encoding tags for each video frame, that is, the encoding method of the video frame. For example, if the encoding information of a video frame is intra-frame coding, the video frame is an I-frame; if the encoding information of a video frame is inter-frame coding, the video frame is a P-frame or B-frame.

[0051] Since the breathing effect is mainly caused by the difference in image distortion between intra-frame coding (I-frames) and inter-frame coding (P-frames or B-frames), the embodiment of the present application only needs to distinguish between I-frame video frames, and does not need to distinguish between P-frame and B-frame video frames. In other embodiments, video frames can also be classified as I-frames, P-frames, and B-frames, which will not be repeated here.

[0052] Step S12: obtaining a continuous sampling frame group in the video set based on the coding information of all video frames, wherein the coding information of at least one video frame included in the continuous sampling frame group is preset coding information.

[0053] In an embodiment of the present application, the terminal device downsamples the video set in the time dimension according to the coding information. For example, the terminal device can downsample the video frames near the I frame in the time dimension according to the coding information to obtain a continuous sampling frame group.

[0054] In the embodiment of the present application, intra-frame coding is defined as preset coding information. The terminal device locates the position of the I frame in the video set based on the preset coding information, and uses the position of the I frame as a standard to obtain other P frames or B frames that are continuous in the time dimension to form a continuous sampling frame group.

[0055] Specifically, if Figure 2 As shown, Figure 2 This is a structural diagram of an embodiment of a continuous sampling frame group provided by this application. Terminal device positioning Figure 2 The position of the I frame in the video set is obtained, and then 2N+1 consecutive video frames of two consecutive GOPs (Group of Pictures) are obtained. Among them, the middle frame of the consecutive 2N+1 frames is an I frame, and the other frames are P frames or B frames. The number of video frames before the I frame is the same as the number of video frames after the I frame. It should be noted that N is an integer greater than 0, and there is only one I frame in a continuous sampling frame group, and the I frame is located in the middle position of the continuous sampling frame group.

[0056] Step S13: Obtain multi-dimensional features of the continuous sampling frame group.

[0057] In an embodiment of the present application, the terminal device obtains multi-dimensional features of a continuous sampling frame group, wherein the multi-dimensional features of the continuous sampling frame group are respectively obtained by obtaining the SSIM value (structural similarity, StructuralSimilarity), brightness difference, texture difference, degree of motion, average brightness and average texture of the continuous sampling frame group. It should be noted that, in other embodiments, the multi-dimensional features do not necessarily adopt all of the above-mentioned dimensional features, and at least four dimensional features of structural similarity, brightness difference, texture difference, degree of motion, average brightness and average texture may also be adopted.

[0058] The following describes the calculation method and meaning of each dimension feature:

[0059] Among them, the SSIM value is the SSIM value of the middle I frame of the continuous sampling frame group and its previous frame. SSIM is an indicator to measure the similarity between two images. The brightness difference is the absolute value of the average brightness difference between the middle I frame and its previous frame. The degree of motion is the sum of the absolute values ​​of the frame differences between the middle I frame and its previous frame (SAD, Sum of absolute differences), which is mainly used for image block matching. The absolute value of the difference between the corresponding values ​​of each pixel is summed to evaluate the similarity of the two image blocks. The average brightness is the average brightness of the 2N+1 frames sampled continuously. The texture difference is the absolute value of the difference in texture between the middle I frame and its previous frame, and the average texture is the average texture of the 2N+1 frames sampled continuously. Among them, the calculation method of the image frame texture is:

[0060]

[0061] Among them, I i (x,y) is the pixel brightness, I mean is the average brightness of the image frame, and M×N is the video resolution, i.e. the number of pixels.

[0062] Step S14: training a breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels.

[0063] In the embodiment of the present application, the terminal device inputs the continuous sampling frame group and its multi-dimensional features and classification labels into the classification model for training. Among them, the multi-classification model includes but is not limited to SVM (Support Vector Machine), Naive Bayes, decision tree and other models. Specifically, the staff can select a specific multi-classification model based on the sample size and experimental classification results.

[0064] In an embodiment of the present application, a terminal device obtains a video set to be trained, as well as encoding information and classification labels of video frames in the video set; obtains a continuous sampling frame group in the video set based on the encoding information of all video frames, wherein the encoding information of at least one video frame in the continuous sampling frame group is preset encoding information; obtains multi-dimensional features of the continuous sampling frame group; and trains a breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels. In the above manner, the training method of the present application trains the discrimination model through multi-dimensional breathing effect influence indicators, making the discrimination of breathing effect more accurate.

[0065] Please continue reading Figure 3 , Figure 3 It is a flow chart of an embodiment of a method for determining the respiratory effect provided by the present application.

[0066] Specifically, if Figure 3As shown, the method for determining the breathing effect in the embodiment of the present application specifically includes the following steps:

[0067] Step S21: Obtain a video set to be identified.

[0068] Among them, step S21 of the embodiment of the present application is basically the same as step S11 in the above embodiment, and will not be repeated here.

[0069] Step S22: Obtain multi-dimensional features of the video set.

[0070] Among them, step S22 of the embodiment of the present application is basically the same as step S13 in the above embodiment, and will not be repeated here.

[0071] Step S23: Input the multi-dimensional features of the video set into a pre-trained breathing effect discrimination model to obtain an output classification label.

[0072] In an embodiment of the present application, a terminal device inputs a video collection and its multi-dimensional features into a pre-trained breathing effect discrimination model to classify the video collection using the breathing effect discrimination model, and outputs a breathing effect classification result. The classification result can be a classification label such as 0, 1, 2, or 3. For example, the classification label can be a classification label such as 0, 1, 2, or 3, where classification label 0 indicates no breathing effect, classification label 1 indicates a mild level of breathing effect, classification label 2 indicates a moderate level of breathing effect, and classification label 3 indicates a severe level of breathing effect. In other embodiments, the terminal device can use classification labels of other types or other values ​​to represent the breathing effect of the video collection, which are not listed here.

[0073] The automated breathing effect discrimination process of the embodiment of the present application reduces errors caused by manual and subjective judgment, and improves the testing efficiency, accuracy, stability and repeatability of the breathing effect.

[0074] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0075] In order to implement the training method of the breathing effect discrimination model of the above embodiment, this application also proposes a terminal device, please refer to Figure 4 , Figure 4 It is a structural diagram of an embodiment of a terminal device provided by this application.

[0076] like Figure 4 As shown, the terminal device 400 provided in this application includes an acquisition module 41, a sampling module 42, a feature module 43 and a training module 44.

[0077] The acquisition module 41 is used to acquire a video set to be trained, as well as encoding information and classification labels of video frames in the video set.

[0078] The sampling module 42 is configured to obtain a continuous sampling frame group in the video set based on the coding information of all video frames, wherein the coding information of at least one video frame in the continuous sampling frame group is preset coding information.

[0079] The feature module 43 is configured to obtain multi-dimensional features of the continuous sampling frame group.

[0080] The training module 44 is used to train the breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels.

[0081] In order to implement the method for determining the breathing effect of the above embodiment, this application also proposes another terminal device, for details, please refer to Figure 5 , Figure 5 It is a structural diagram of another embodiment of the terminal device provided by this application.

[0082] like Figure 5 As shown, the terminal device 500 provided in this application includes an acquisition module 51, a feature module 52 and a determination module 53.

[0083] The acquisition module 51 is used to acquire a video set to be identified.

[0084] The feature module 52 is configured to obtain multi-dimensional features of the video set.

[0085] The discrimination module 53 is configured to input the multi-dimensional features of the video set into a pre-trained breathing effect discrimination model to obtain an output classification label.

[0086] In order to implement the training method of the breathing effect discrimination model and / or the breathing effect discrimination method of the above embodiment, this application also proposes another terminal device, which can be found in detail. Figure 6 , Figure 6 It is a structural diagram of another embodiment of the terminal device provided by this application.

[0087] The terminal device 600 of the embodiment of the present application includes a memory 61 and a processor 62, wherein the memory 61 and the processor 62 are coupled.

[0088] The memory 61 is used to store program data, and the processor 62 is used to execute the program data to implement the breathing effect discrimination model training method and / or the breathing effect discrimination method described in the above embodiments.

[0089] In this embodiment, the processor 62 may also be referred to as a CPU (Central Processing Unit). The processor 62 may be an integrated circuit chip having signal processing capabilities. The processor 62 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 62 may be any conventional processor.

[0090] This application also provides a computer storage medium, such as Figure 7 As shown, the computer storage medium 700 is used to store program data 71. When the program data 71 is executed by the processor, it is used to implement the training method of the breathing effect discrimination model and / or the breathing effect discrimination method as described in the above embodiments.

[0091] This application also provides a computer program product, wherein the computer program product includes a computer program operable to cause a computer to execute the breathing effect discrimination model training method and / or the breathing effect discrimination method described in the embodiments of this application. The computer program product may be a software installation package.

[0092] The training method of the respiratory effect discrimination model and / or the respiratory effect discrimination method described in the above embodiments of the present application, when implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0093] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A training method for a respiratory effect discrimination model, characterized in that: The training method comprises: Obtaining a video set to be trained, as well as encoding information and classification labels of video frames in the video set; Obtaining a continuous sampling frame group in the video set based on the encoding information of all video frames, wherein the encoding information of at least one video frame included in the continuous sampling frame group is preset encoding information; Acquire multi-dimensional features of the continuous sampling frame group; Training a breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels; The obtaining of a continuous sampling frame group in the video set based on the encoding information of all video frames includes: Locating a position of a target video frame with preset coding information in the video set based on coding information of all video frames; The position of the target video frame is used as the middle position, and the continuous sampling frame group is obtained according to the time sequence of the video set.

2. The training method according to claim 1, characterized in that The multi-dimensional features include at least four dimensional features of structural similarity, brightness difference, texture difference, motion degree, average brightness and average texture.

3. The training method according to claim 1, characterized in that The number of video frames in the continuous sampling frame group that are time-series before the target video frame is the same as the number of video frames that are time-series after the target video frame.

4. The training method according to claim 1 or 3, characterized in that: The coding information of the video frame includes intra-frame coding and inter-frame coding, and the preset coding information is intra-frame coding.

5. A method for determining a breathing effect, characterized in that: The discrimination method includes: Get the video set to be judged; Obtaining multi-dimensional features of the video collection; The multi-dimensional features of the video set are input into a pre-trained breathing effect discrimination model to obtain an output classification label, wherein the breathing effect discrimination model is the breathing effect discrimination model according to any one of claims 1 to 4.

6. The identification method according to claim 5, characterized in that: The multi-dimensional features include at least four dimensional features of structural similarity, brightness difference, texture difference, motion degree, average brightness and average texture.

7. A terminal device, characterized in that: The terminal device includes an acquisition module, a sampling module, a feature module and a training module; The acquisition module is used to acquire a video set to be trained, as well as encoding information and classification labels of video frames in the video set; The sampling module is configured to obtain a continuous sampling frame group in the video set based on the coding information of all video frames, wherein the coding information of at least one video frame in the continuous sampling frame group is preset coding information; The feature module is used to obtain multi-dimensional features of the continuous sampling frame group; The training module is used to train the breathing effect discrimination model based on the continuous sampling frame group and its multi-dimensional features and classification labels; The sampling module is further used to locate the position of the target video frame with preset coding information in the video set based on the coding information of all video frames; and to obtain the continuous sampling frame group according to the time sequence of the video set with the position of the target video frame as the middle position.

8. A terminal device, characterized in that: The terminal device includes a memory and a processor, wherein the memory is coupled to the processor; The memory is used to store program data, and the processor is used to execute the program data to implement the training method according to any one of claims 1 to 4 and the discrimination method according to any one of claims 5 to 6.

9. A computer storage medium, characterized in that The computer storage medium is used to store program data, and when the program data is executed by the processor, it is used to implement the training method according to any one of claims 1 to 4 and the discrimination method according to any one of claims 5 to 6.

Citation Information

Patent Citations

  • Video coding method and device for suppressing respiratory effect

    CN111314700A

  • Video decoding method and device, equipment and storage medium

    CN112188215A