Video image processing method and device
Through the methods of decoding, identifying, blurring and encapsulating the video, the problem of privacy leakage during video acquisition in the prior art is solved, and effective privacy protection for the objects of interest is achieved.
Patent Information
- Application Number
- CN202510113887.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology fails to effectively protect privacy during video collection, resulting in the leakage of other information recorded in pet videos, and there is a risk of privacy leakage.
Provide a method for processing video images, through steps such as video decoding, object recognition of interest, blurring processing and packaging, information that needs to be protected is blurred to avoid privacy leakage.
Effectively identify and protect the privacy of the objects of interest, avoid the risk of privacy leakage, and ensure the normal playback and storage of videos.
Smart Images

Figure CN120091141A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to a method and apparatus for processing video images. Background Art
[0002] Currently, there are many video capture products for home pet care, and home surveillance cameras are very common. However, when capturing pets through a camera device, other information is often recorded as well, such as human behavior, resulting in a high risk of privacy leakage.
[0003] The prior art does not perform privacy protection processing on videos. Devices for viewing videos, such as mobile terminals like mobile phones, simply export the original video during the video export process, presenting a problem of privacy leakage. Summary of the Invention
[0004] In view of the problems existing in the prior art, this application provides a processing solution for video images that can perform blurring processing on information that needs to be protected.
[0005] According to the first aspect of this application, there is provided a method for processing video images, characterized by comprising:
[0006] Performing video decoding on the obtained video file to obtain decoded image frames, wherein the video file is a pet scene video file;
[0007] Identifying the decoded image frames to determine the objects of interest and their location information;
[0008] Determining the area to be subjected to blurring processing according to the location information;
[0009] Performing blurring processing on the area to be subjected to blurring processing to obtain blurred image frames; and
[0010] Encapsulating the blurred image frames.
[0011] According to the second aspect of this application, there is provided a device for processing video images, characterized by comprising:
[0012] A decoding module for performing video decoding on the obtained video file to obtain decoded image frames, wherein the video file is a pet scene video file;
[0013] An identification module for identifying the decoded image frames to determine the objects of interest and their location information;
[0014] A determination module for determining the area to be subjected to blurring processing according to the location information;
[0015] A blurring processing module, configured to perform blurring processing on the area to be blurred, and obtain a blurred image frame; and
[0016] An encapsulation module, configured to encapsulate the blurred image frame.
[0017] According to a third aspect of the present application, there is provided an electronic device, including:
[0018] A processor; and
[0019] A memory storing computer instructions, when the computer instructions are executed by the processor, the processor is caused to execute the method described in the first aspect.
[0020] According to a fourth aspect of the present application, there is provided a non-transitory computer storage medium storing a computer program, when the computer program is executed by a plurality of processors, the processors are caused to execute the method described in the first aspect.
[0021] According to the video image processing method and device provided by the present application, it is possible to identify the object of interest to be displayed and its position, and at the same time perform blurring processing on the information to be protected, avoiding the risk of privacy leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings according to these drawings, without exceeding the scope required to be protected by the present application.
[0023] Figure 1 It is a flowchart of a video image processing method according to an embodiment of the present application.
[0024] Figure 2 It is a flowchart of a video image processing method according to another embodiment of the present application.
[0025] Figure 3 It is a flowchart of a video image processing method according to still another embodiment of the present application.
[0026] Figure 4 It is a schematic diagram of a video image processing device according to an embodiment of the present application.
[0027] Figure 5 It is a schematic diagram of a video image processing device according to another embodiment of the present application.
[0028] Figure 6Schematic diagram of a video image processing device according to another embodiment of the present application.
[0029] Figure 7 Schematic diagram of the structure of an electronic device provided by the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0031] Throughout the specification and claims, terms may have subtle meanings implied or suggested in the context, rather than the explicitly stated meanings. Similarly, the phrase "in one embodiment" or "in some embodiments" used herein does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" used herein does not necessarily refer to different embodiments. The phrase "in one implementation" or "in some implementations" used herein does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" used herein does not necessarily refer to different implementations. For example, the claimed subject matter includes all or part of the combinations of the exemplary embodiments or implementations.
[0032] Generally speaking, terms can be at least partially understood from their use in the context. For example, terms such as "and", "or", and "and / or" used herein can include multiple meanings, which can at least partially depend on the context in which these terms are used. Generally, "or" if used to relate a list, such as A, B, or C, means A, B, and C, used herein in an inclusive sense, as well as A, B, or C, used herein only in an exclusive sense. In addition, the term "one or more" or "at least one" used herein can, at least partially depending on the context, be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, and properties in a plural sense. Similarly, terms such as "a", "an", or "the" can, at least partially depending on the context, also be understood to convey a singular or plural usage. In addition, also at least partially depending on the context, the terms "based on" or "determined by" can be understood not necessarily to mean a set of exclusive factors, but rather to allow for the existence of other factors that are not necessarily explicitly described.
[0033] According to one aspect of the present application, a method for processing video images is provided. Figure 1 Schematic flowchart of a method for processing video images according to an embodiment of the present application. AsFigure 1 As shown, the method includes the following steps.
[0034] Step S101: Perform video decoding on the acquired video file to obtain the decoded image frames.
[0035] In some embodiments, an imaging device in a household captures images of a scene containing a pet, and the acquired video file is a pet scene video file, which can be an MP4 format file. In one embodiment, in the iOS platform, the AVFoundation framework AVAssetReader can be used to read the video file and obtain the decoded image frame sampleBuffer. In other platforms, such as the Android platform, corresponding frameworks can be used for video decoding.
[0036] In some embodiments, the decoded image frames can include various information of the frame video, such as image width and height, image information, image format, etc.
[0037] Step S102: Identify the object of interest and its location information from the decoded image frames.
[0038] In some embodiments, a corresponding object detection algorithm is used to identify the object of interest and its location information from the decoded image frames. In a specific embodiment, the object detection algorithm can be the YOLO model, and the YOLO model can be deployed to the iOS platform. The model identifies the object of interest and its location information in the image frames. In the iOS platform, the YOLO model is loaded using Core ML, a detector identifier is created using VNCoreMLModel, and the sampleBuffer is input into the YOLO model for prediction inference. The inference output identifies the object of interest and its location information. During the process of using the object detection algorithm to identify the decoded image frames, the object of interest to be output can be specified, and the object detection algorithm outputs the object of interest accordingly.
[0039] In some embodiments, the object of interest can be at least one of a pet (including cats, dogs, etc.), a person, an item, etc. in the pet scene. The location information of the object of interest is the location of the object of interest in the image frame. For example, for an image frame with a size of 1920*1080, the location information of the object of interest is (100*100, 200*300), where 100*100 represents the position of the upper left corner of the rectangle enclosing the object of interest, and 200*300 represents the width and height of the rectangle enclosing the object of interest.
[0040] Step S103: Determine the area to be blurred according to the location information;
[0041] Step S104: Perform blurring processing on the area to be blurred, and obtain a blurred image frame.
[0042] In some embodiments, an area other than the position information in the decoded image frame can be determined as the area to be blurred. For example, if the object of interest is a pet, the area outside the pet's position can be determined as the area to be blurred, and the area outside the pet's position is blurred to obtain a blurred image frame, so that only the pet picture is displayed.
[0043] In some embodiments, the area corresponding to the position information can be determined as the area to be blurred. For example, if the object of interest is a person, the area corresponding to the position information of the person can be determined as the area to be blurred, and the area corresponding to the position information of the person is blurred to obtain a blurred image frame, so as not to disclose privacy information.
[0044] In a specific embodiment, OPENGLES can be used to blur the image and output the processed image pixcelBuffer.
[0045] Step S105: Package the blurred image frame.
[0046] In some embodiments, the blurred image frame is packaged to form a video file. In a specific embodiment, the data can be encoded and packaged into a video file through the packaging function modules AVAssetWriterInputPixelBufferAdaptor and AVAssetWriter provided by the iOS platform.
[0047] Figure 2 It is a schematic flowchart of a method for processing video images according to another embodiment of the present application. Compared with Figure 1 Figure 2 Steps S202 to S206 of Figure 1 are the same as steps S101 to S105 of Figure 2 The difference is that the method shown in
[0048] Step S201: Train the original YOLO model according to samples related to the pet scene to obtain a trained YOLO model.
[0049] The original YOLO model is trained based on existing publicly available datasets. However, the image recognition effect in specific scenarios is not ideal. In this application, a new dataset specifically for the pet scenario is created, and the original YOLO model is fine-tuned using this dataset. Using the trained YOLO model to recognize the decoded image frames can significantly improve the accuracy of model recognition. In some embodiments, the training process for the original YOLO model includes: data collection, data annotation, dataset allocation, and data fine-tuning.
[0050] In some embodiments, first, image data is collected. For example, 5000 pictures are collected as training data, and these images cover the situations of pets in different environments, lighting conditions, and various daily activities. In a specific embodiment, the data sources can include three parts. Among them, the first part of the data comes from the network. For example, the Web crawler technology Apify can be used to capture pictures and videos related to the pet camera footage on existing social media platforms, such as Xiaohongshu, Douyin, Google Search, YouTube, etc., and capture the pictures and videos locally; the second part of the data can be obtained by offline collection. For example, 10 real pet camera collection scenarios are prepared, with a recording duration of 2 weeks, and video data is collected; the third part of the data can be co-created with users. Cooperate with users who currently use pet care cameras. Considering the diversity of pets, collect the video data provided by users.
[0051] Then, the collected data is annotated to form a sample dataset. Each sample includes image data, the object of interest, and location information. In a specific embodiment, corresponding to the 5000 pictures collected, 5000 groups of data pairs can be sorted out to form a sample dataset. Each pair of data pairs contains picture data and the object of interest and its location information, such as pet type and its location information, person and its location information, etc., and can be in the form of a combination of ".png" + ".txt".
[0052] In some embodiments, relevant automation scripts can be written in Python to convert videos into pictures and screen out pictures with pets; the LabelImg tool can be used to annotate the pictures, and the YOLO type of annotation method can be selected to annotate all the pictures.
[0053] To achieve the richness and diversity of the sample dataset, the samples can include different environmental backgrounds (such as indoor, outdoor, different rooms), lighting conditions at different times (daytime, night, dark places, etc.), and different types of pet activities (such as eating, playing, resting, etc.). These diverse images make the dataset highly applicable in the field of household pet recognition.
[0054] In a specific embodiment, the sample data set can be allocated. For example, the sample data set can be divided into a training data set and a validation data set. For example, for a sample data set containing 5000 samples, the training data set can contain 4000 samples, and the validation data set can contain 1000 samples. The sample data set can be randomly divided into two groups and placed in the training data set folder and the validation data set folder respectively. By reasonably allocating the sample data set into a training data set and a validation data set, it can be ensured that the model can fully learn during training through the training data set, and the generalization ability and performance of the model can be evaluated through the validation data set.
[0055] In some embodiments, the training data set is input into the original YOLO model for training to obtain the trained YOLO model.
[0056] In a specific embodiment, the PyTorch framework can be used to train the original YOLO model, and transfer learning is performed on the collected data set to fine-tune the YOLO model. For example, the initial value of the learning rate is adjusted to 0.01, the batch size is 16, and the number of training epochs is 50. The model is trained using a graphics card, and then the learning rate is increased to 0.03, the batch size is 16, and the number of training epochs is 50. The training of the model can be optimized by adjusting training parameters such as the learning rate, batch size, and number of training epochs.
[0057] Figure 3 It is a schematic flowchart of a method for processing video images according to another embodiment of the present application. Compared with Figure 2 compared, Figure 3 Steps S301, S303 to S306 of Figure 2 are the same as steps S201 to S206 of Figure 3 The difference is that the method shown in
[0058] also includes:
[0059] Step S302, using the validation data set to verify the trained YOLO model to determine the recognition accuracy of the trained YOLO model.
[0060] In some embodiments, a trained YOLO model is used to predict a validation data set to obtain the objects of interest (such as pets or people) in each image, the location information of the objects of interest (such as a rectangular box encompassing the object of interest), and / or the confidence level. The precision and recall of the images of each category can be calculated: for each image, calculate the matching situation between the prediction result and the true label, including: the predicted location information matches the true location information and the category of the object of interest is correct, the predicted location information does not match the true location information and the category of the object of interest is incorrect, and the true location information is not predicted.
[0061] In a specific embodiment, for each type of object of interest, such as cats, dogs, parrots, people, etc., the recognition precision of the YOLO model corresponding to the object of interest of this type can be determined according to the predicted location information and the true location information in the corresponding annotated image. In another specific embodiment, according to the recognition precision of the YOLO model corresponding to each object of interest, the average recognition precision of the YOLO model is determined, for example, it can be calculated by means of weighted average.
[0062] According to another aspect of the present application, a video image processing device is provided. Figure 4 It is a schematic diagram of a video image processing device according to an embodiment of the present application. As Figure 4 shown, the device includes: a decoding module 401, an identification module 402, a determination module 403, a blurring processing module 404, and a packaging module 405. Among them, the decoding module 401 is used to perform video decoding on the acquired video file to obtain the decoded image frames; the identification module 402 is used to identify the decoded image frames to determine the objects of interest and their location information; the determination module 403 is used to determine the area to be subjected to blurring processing according to the location information; the blurring processing module 404 is used to perform blurring processing on the area to be subjected to blurring processing to obtain the blurred image frames; the packaging module 405 is used to package the blurred image frames.
[0063] In some alternative embodiments, the determination module 403 can be used to:
[0064] Determine the area other than the location information in the decoded image frame as the area to be subjected to blurring processing; or
[0065] Determine the area corresponding to the location information as the area to be subjected to blurring processing.
[0066] Figure 5 It is a schematic diagram of a video image processing device according to another embodiment of the present application. Compared with Figure 4 Figure 5The modules 502 to 506 are the same as Figure 4 the modules 401 to 405, except that Figure 5 the device shown further includes:
[0067] A training module 501, configured to train the original YOLO model according to samples related to pet scenarios to obtain a trained YOLO model.
[0068] In some alternative embodiments, the training module 501 may be configured to:
[0069] Collect a sample data set related to pet scenarios, where the sample data set includes a training data set, and each sample in the sample data set includes image data, an object of interest, and location information; and
[0070] Input the training data set into the original YOLO model for training to obtain the trained YOLO model.
[0071] Figure 6 is a schematic diagram of a video image processing device according to another embodiment of the present application. Compared with Figure 5 that, Figure 6 the modules 601, 603 to 607 are the same as Figure 5 the modules 501 to 506, except that Figure 6 the device shown further includes:
[0072] A verification module 602, configured to verify the trained YOLO model using a verification data set to determine the recognition accuracy rate of the trained YOLO model.
[0073] In some alternative embodiments, the verification module 602 may be configured to:
[0074] Predict the verification data set through the trained YOLO model to determine the location information corresponding to the object of interest in each image of the verification data set; and
[0075] Determine the recognition accuracy rate of the trained YOLO model according to the location information determined in each image and the true location information in each image.
[0076] According to the video image processing method and device provided by the present application, it is possible to identify the object of interest to be displayed and its location, and at the same time perform blurring processing on the information to be protected to avoid the risk of privacy leakage.
[0077] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0078] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.
[0079] In the several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical connection or other forms.
[0080] See also Figure 7 , Figure 7 An electronic device is provided, comprising a processor and a memory. The memory stores computer instructions, and when the computer instructions are executed by the processor, the processor executes the computer instructions to achieve the following Figures 1 to 3 The method and refinement scheme shown.
[0081] It should be understood that the above device embodiments are only illustrative, and the device disclosed in the present invention can also be implemented in other ways. For example, the division of units / modules described in the above embodiments is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.
[0082] In addition, unless otherwise specified, each functional unit / module in each embodiment of the present invention may be integrated into one unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The above-mentioned integrated unit / module may be implemented in the form of hardware or in the form of a software program module.
[0083] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor or chip can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the on-chip cache, off-chip memory, and memory can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0084] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer electronic device (which can be a personal computer, a server, or a network electronic device, etc.) to execute all or part of the steps of the methods described in various embodiments of this disclosure. And the aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.
[0085] The embodiments of this application also provide a non-transitory computer storage medium storing a computer program, which, when executed by multiple processors, causes the processors to execute the Figures 1 to 3 method and refinement solutions as shown.
[0086] References in this specification to features, advantages, or similar language do not imply that all features and advantages that can be realized by the solution should be included in or included in any single implementation thereof. On the contrary, language referring to features and advantages is understood to mean that a particular feature, advantage, or characteristic described in connection with an embodiment is included in at least one embodiment of the solution. Thus, the discussion of features, advantages, and similar language throughout the specification may, but does not necessarily, refer to the same embodiment.
[0087] In addition, the features, advantages, and characteristics of the solution may be combined in any suitable manner in one or more embodiments. Based on the description herein, those of ordinary skill in the relevant art will recognize that the solution can be implemented without one or more specific features or advantages of a particular embodiment. In other cases, additional features and advantages can be realized in a particular embodiment that is not presented in all embodiments of the solution.
[0088] The embodiments of the present application have been introduced in detail above. Specific examples are used herein to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, any changes or deformations made by those skilled in the art based on the idea of the present application, within the specific implementation manners and application scope of the present application, fall within the scope of protection of the present application. In summary, the content of this specification should not be construed as a limitation on the present application.
Claims
1. A method for processing a video image, characterized in that: include: Performing video decoding on the acquired video file to obtain decoded image frames, wherein the video file is a pet scene video file; Identifying the decoded image frame to determine the object of interest and its location information; Determine the area to be fuzzy processed according to the position information; Performing blur processing on the area to be blurred to obtain a blurred image frame; and The blurred image frame is encapsulated.
2. The method according to claim 1, characterized in that Also includes: The original YOLO model is trained according to samples related to pet scenes to obtain a trained YOLO model; Wherein, identifying the decoded image frame includes: The trained YOLO model is used to recognize the decoded image frame.
3. The method according to claim 2, characterized in that The original YOLO model is trained according to samples related to pet scenes, including: Collecting a sample data set related to a pet scene, wherein the sample data set includes a training data set, and each sample in the sample data set includes image data, an object of interest, and location information; and The training data set is input into the original YOLO model for training to obtain the trained YOLO model.
4. The method according to claim 3, characterized in that The sample data set also includes a validation data set, and the method further includes: The trained YOLO model is verified using the verification data set to determine the recognition accuracy of the trained YOLO model.
5. The method according to claim 4, characterized in that The using the verification data set to verify the trained YOLO model includes: Predicting the verification data set by using the trained YOLO model to determine the position information corresponding to the object of interest in each image of the verification data set; and The recognition accuracy of the trained YOLO model is determined based on the determined position information in each image and the actual position information in each image.
6. The method according to any one of claims 1 to 5, characterized in that: The step of determining the area to be fuzzy processed according to the position information includes: Determine the area outside the position information in the decoded image frame as the area to be blurred; or The area corresponding to the position information is determined as the area to be fuzzy processed.
7. The method according to any one of claims 1 to 5, characterized in that The object of interest includes at least one of a pet and a person.
8. A video image processing device, characterized in that: include: A decoding module, used for decoding the acquired video file to obtain a decoded image frame, wherein the video file is a pet scene video file; An identification module, used to identify the decoded image frame and determine the object of interest and its location information; A determination module, used to determine an area to be fuzzy processed according to the location information; A fuzzy processing module, used for performing fuzzy processing on the area to be fuzzy processed to obtain a fuzzy processed image frame; and The encapsulation module is used to encapsulate the image frame after the blurring process.
9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program in the memory.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.