Video content protection method, system, electronic device and storage medium

By dynamically adding complex random watermarks on the backend server, the problem of private recording of video content is solved, and effective copyright protection and user experience optimization are achieved.

CN120475228BActive Publication Date: 2025-09-16HEBEI YINGYAN INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510957091.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-16
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Video creators face serious unauthorized recording attacks, which lead to copyright damage and threats to economic benefits. Existing technologies make it difficult to effectively protect video content.

Method used

Watermarks are added dynamically frame by frame on the backend server. The FFmpeg processing engine and multi-layer frequency superposition watermarking strategy are used, combined with user behavior and device characteristics, to generate complex and random watermark motion trajectories, which are embedded in the video stream to ensure that the watermark is difficult to tamper with or remove.

Benefits of technology

It improves the copyright protection of video content, reduces the risk of private recording, ensures the robustness and integrity of watermarks, and optimizes the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475228B_ABST
    Figure CN120475228B_ABST
Patent Text Reader

Abstract

The present application relates to a method, system, electronic device and storage medium for protecting video content, and relates to the technical field of video image processing. The method includes: receiving a video access request sent by a client, and obtaining a video identifier corresponding to the video access request; if target video data corresponding to the video identifier is stored in a local storage unit, then reading the target video data in the local storage unit; based on a preset watermark strategy, dynamically adding a watermark to the target video data frame by frame to generate target video data with a watermark; based on a streaming media server, pushing the target video data with a watermark to the client for user viewing. The watermark of the present application is directly embedded in the video stream and cannot be removed or tampered with by the front-end code, thereby ensuring the integrity and security of the watermark. In addition, the position and content of the watermark are constantly changing, which increases the difficulty of removing the watermark.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video image processing, and in particular to a method, system, electronic device, and storage medium for protecting video content. Background Art

[0002] With the widespread adoption of internet technology and mobile devices, video content has become a crucial way for people to access information, entertainment, and communication. From early traditional television broadcasts to online streaming platforms and today's short video sharing on social media, the video industry has undergone rapid development and transformation. Throughout this process, the ease of use of video creation tools and the diversification of distribution channels have made it easier for individual creators to participate in content production, significantly enriching the video resources available on the market.

[0003] Despite the booming video industry, video creators face serious copyright infringement issues, particularly unauthorized private recordings. These recordings lack any information indicating their source and copyright, severely damaging the original creators' copyrights and threatening the economic benefits of copyright holders. Therefore, a technical solution for video content protection is urgently needed. Summary of the Invention

[0004] In order to protect video content and reduce the risk of private recording of video content, the present application provides a video content protection method, system, electronic device and storage medium.

[0005] In a first aspect, the present application provides a method for protecting video content, which adopts the following technical solution:

[0006] A video content protection method, applied to a back-end server, comprising:

[0007] Receive a video access request sent by a client, and obtain a video identifier corresponding to the video access request;

[0008] If the target video data corresponding to the video identifier is stored in the local storage unit, the target video data in the local storage unit is read; if the target video data is not stored in the local storage unit, a request for obtaining the target video data is sent to a video storage object source, and the target video data returned by the video storage object source is received;

[0009] Based on a preset watermark strategy, dynamically adding watermarks to the target video data frame by frame to generate target video data with watermarks;

[0010] Based on the streaming media server, the target video data with the watermark is pushed to the client for the user to watch.

[0011] By adopting this technical solution, the watermark is added on the backend server, embedded directly into the video stream. It cannot be removed or tampered with through front-end code, ensuring the integrity and security of the watermark and effectively protecting the video content. Using dynamic watermarking technology, the position and content of the watermark constantly change, making it more difficult to remove. Even if an attacker can remove the watermark in one frame, it is difficult to completely remove the watermark from the entire video. This improves the robustness of the watermark and reduces the risk of private recording of video content.

[0012] Optionally, dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy includes:

[0013] For each video frame of the target video data, obtaining a horizontal position in a horizontal direction based on a preset X-axis motion function; the X-axis motion function includes a first function, a second function, and a third function of different motion periods;

[0014] For each video frame of the target video data, obtaining a vertical position in a vertical direction based on a preset Y-axis motion function; the Y-axis motion function includes a fourth function, a fifth function, and a sixth function of different motion periods; and the watermark strategy includes the X-axis motion function and the Y-axis motion function;

[0015] For each video frame of the target video data, obtaining a target watermark position corresponding to the video frame based on the horizontal position and the vertical position, and adding a watermark to the video frame at the target watermark position based on an FFmpeg processing engine;

[0016] The X-axis motion function is expressed as:

[0017] X(t)=min(max(k,w / 2+(w / 2-tw-d)×[λ1×sin(2πt / T1)+λ2×sin(2πt / T2)+λ3×sin(2πt / T3)]),w-tw-k)

[0018] Wherein, k represents a safety margin, w represents the horizontal size of the video frame, tw represents the horizontal size of the watermark image, d represents an offset, sin(2πt / T1) represents the first function, λ1 represents the weight of the first function, T1 represents the period of the first function, sin(2πt / T2) represents the second function, λ2 represents the weight of the second function, T2 represents the period of the second function, sin(2πt / T3) represents the third function, λ3 represents the weight of the third function, T3 represents the period of the third function; λ1 is greater than λ2, λ2 is greater than λ3, T1 is greater than T2, T2 is greater than T3, and T1, T2, and T3 are coprime numbers;

[0019] The Y-axis motion function is expressed as:

[0020] Y(t)=min(max(k,h / 2+(h / 2-th-d)×[λ4×cos(2πt / T4)+λ5×cos(2πt / T5)+λ6×cos(2πt / T6)]),h-th-k)

[0021] Wherein, h represents the vertical size of the video frame, th represents the vertical size of the watermark image, cos(2πt / T4) represents the fourth function, λ4 represents the weight of the fourth function, T4 represents the period of the fourth function, cos(2πt / T5) represents the fifth function, λ5 represents the weight of the fifth function, T5 represents the period of the fifth function, cos(2πt / T6) represents the sixth function, λ6 represents the weight of the sixth function, and T6 represents the period of the sixth function; the λ4 is greater than λ5, λ5 is greater than λ6, the T4 is greater than T5, T5 is greater than T6, and T4, T5 and T6 are coprime numbers.

[0022] By adopting the above technical solution, a complex and random watermark motion trajectory is formed through multi-layer frequency superposition and different weight distribution, which greatly enhances the anti-tampering ability of video content and provides strong technical support for video content protection.

[0023] Optionally, the adding a watermark to the video frame at the target watermark position based on the FFmpeg processing engine includes:

[0024] For each video frame of the target video data, if the complexity of the watermark is higher than a preset threshold, the video frame is divided into sub-regions according to the real-time video content in the video frame, and the watermark is split into its component elements;

[0025] For each video frame of the target video data, dynamically adjusting the relative positions of the component elements of the watermark according to content characteristics of the plurality of sub-regions in the video frame, and synchronously adding each component element to the sub-region at a corresponding position based on the FFmpeg processing engine and the target watermark position;

[0026] For each video frame of the target video data, after all the component elements are added, the sub-regions are merged to form a video frame with a complete watermark.

[0027] By employing this technical solution, complex watermarks are broken down into multiple components and added to appropriate sub-regions of the video frame. This allows for more precise control over the watermark's location and embedding quality, improving its visibility and robustness, making it more difficult to remove or tamper with. Dividing the video frame into sub-regions and processing the addition of these components in parallel leverages the performance advantages of multi-core processors, improving watermark embedding efficiency and reducing processing time, ensuring the real-time and smooth flow of video streams.

[0028] Optionally, for each video frame of the target video data, before obtaining the horizontal position in the horizontal direction based on the preset X-axis motion function and obtaining the vertical position in the vertical direction based on the preset Y-axis motion function, the method further includes:

[0029] Obtaining a trust level of the user based on historical viewing data of the user corresponding to the client, and obtaining weights of the first function and the fourth function based on the trust level; the historical viewing data includes historical behavior, including whether there has been any private recording behavior, viewing duration, or interaction frequency;

[0030] For each video frame of the target video data, obtaining a weight of the second function and a weight of the fifth function based on real-time video content in the video frame;

[0031] For each video frame of the target video data, the weight of the third function is obtained based on the real-time network status between the client and the target video data; and the weight of the sixth function is obtained based on the device screen characteristics corresponding to the target video data, wherein the device screen characteristics include the screen aspect ratio.

[0032] By adopting this technical solution, the weight of the watermark motion function can be dynamically adjusted based on the user's historical behavior, real-time video content, real-time network conditions, and device screen characteristics. This not only improves the effectiveness of video copyright protection, but also optimizes the user's viewing experience.

[0033] Optionally, the dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes:

[0034] For each video frame of the target video data, emotion analysis and scene recognition are performed on the real-time video content of the video frame to extract emotion information and scene information corresponding to the video frame, and based on the emotion information and scene information, initial color information of the watermark is obtained; the emotion information includes joy, sadness or tension, and the scene information includes sky, ocean, forest or city;

[0035] For each video frame of the target video data, user color preference data is obtained based on the historical viewing data, and based on the user color preference data, the initial color information is adjusted to obtain target color information, so that the watermark at the target watermark position is displayed based on the target color information; the color information includes the watermark color corresponding to the video frame and the predicted color change trend of the watermark corresponding to subsequent video frames.

[0036] By adopting the above technical solution, the most appropriate target color information can be dynamically determined for the watermark of each video frame, thereby providing a more personalized and comfortable viewing experience while protecting the video copyright.

[0037] Optionally, the dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes:

[0038] Obtaining an initial transparency of the watermark based on a device type of a device screen corresponding to the client and metadata of the target video data;

[0039] Obtaining environmental display characteristics of the device screen corresponding to the client, wherein the environmental display characteristics include user viewing distance, screen brightness, and ambient light information;

[0040] For each video frame of the target video data, the initial transparency is adjusted based on the environmental display characteristics to obtain a target transparency, so that the watermark at the target watermark position is displayed based on the target transparency.

[0041] By adopting the above technical solution, by dynamically adjusting the transparency of the watermark according to the device characteristics and viewing environment, it can be ensured that the watermark can achieve the best display effect and copyright protection effect in different devices and environments, thereby optimizing the watermark display effect and user experience.

[0042] Optionally, obtaining environmental display characteristics of the device screen corresponding to the client, the environmental display characteristics including user viewing distance, screen brightness, and ambient light information, includes:

[0043] Analyzing the historical viewing data based on a machine learning algorithm to identify the user's current viewing behavior pattern; the historical viewing data includes viewing time, viewing location, viewing device type, type of video viewed, and whether the user has ever adjusted the watermark transparency or reported the watermark for obstructing content;

[0044] Based on the behavior prediction model and the current viewing behavior pattern, the device type and the ambient light information used by the user when currently watching the video are predicted; and the user viewing distance and the screen brightness sent by the client are received.

[0045] By adopting the above technical solution, by integrating the user's historical behavior and current viewing environment, the environmental display characteristics related to the user's viewing context can be predicted and obtained, providing a basis for the subsequent dynamic adjustment of the watermark transparency, thereby optimizing the user's viewing experience.

[0046] In a second aspect, the present application provides a protection system using the method described in any one of the first aspects, employing the following technical solutions:

[0047] A video content protection system includes a client and a back-end server; the client includes a video playback module, a user interaction module and a WebSocket client, and the back-end server includes a WebSocket server, an FFmpeg processing engine and a TCP server;

[0048] The video playing module is used to send a video access request to the back-end server and play the target video data with the watermark pushed by the back-end server;

[0049] The user interaction module is used to provide a user interface and receive user operation instructions;

[0050] The WebSocket client is used to establish a connection with the WebSocket server for real-time communication;

[0051] The WebSocket server is configured to receive a connection request from the WebSocket client and establish a real-time communication connection with the WebSocket client;

[0052] The FFmpeg processing engine is used to add a dynamic watermark to the target video data corresponding to the video access request;

[0053] The TCP server is used to transmit the target video data with the watermark to the WebSocket server.

[0054] In a third aspect, the present application provides an electronic device, which adopts the following technical solution:

[0055] An electronic device comprises a processor and a memory, wherein the processor is coupled to the memory;

[0056] The processor is configured to execute a computer program stored in the memory, so that the electronic device executes the method according to any one of the first aspects.

[0057] In a fourth aspect, the present application provides a computer-readable storage medium, which adopts the following technical solution:

[0058] A computer-readable storage medium comprises a computer program or instructions, which, when executed on a computer, causes the computer to execute the method according to any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flowchart of a method for protecting video content according to one embodiment of the present application.

[0060] Figure 2 This is a structural block diagram of a video content protection system according to one embodiment of the present application.

[0061] Figure 3 This is a structural block diagram of an electronic device according to one embodiment of the present application. DETAILED DESCRIPTION

[0062] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0063] The present application is further described in detail below with reference to the accompanying drawings.

[0064] An embodiment of the present application provides a method for protecting video content, which can be executed by a back-end server. The back-end server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0065] like Figure 1 As shown, a method for protecting video content is executed by a backend server. The main process of the method is described as follows (steps S101 to S104):

[0066] Step S101: receiving a video access request sent by a client, and obtaining a video identifier corresponding to the video access request.

[0067] The client communicates with the backend server, which receives video access requests from the client (e.g., a user device). The client can send requests through various methods, such as HTTP or WebSocket. The video access request can include a video identifier (e.g., a video ID) for the target video. The backend server uses the video identifier to determine the specific video content requested by the user, namely the target video data.

[0068] Step S102: If the target video data corresponding to the video identifier is stored in the local storage unit, the target video data in the local storage unit is read; if the target video data is not stored in the local storage unit, a request for obtaining the target video data is sent to the video storage object source, and the target video data returned by the video storage object source is received.

[0069] The backend server checks whether the requested target video data is already stored in a local storage unit (such as a local disk or cache system). If so, it reads the data directly from the local storage unit. If not, it sends a request to a remote video storage source (such as a cloud storage service or object storage system) and receives the returned data. This ensures that the target video data is quickly and reliably provided to subsequent processing steps, while reducing reliance on external storage and network transmission overhead.

[0070] In this embodiment, the backend server may choose to store specific video data in the local storage unit based on considerations such as access efficiency, cost optimization, or user experience improvement. Specifically:

[0071] The first type of storage: Storing frequently accessed video data based on play count and popularity. Dynamic thresholds, such as daily request volume, can be set to automatically cache video data to local storage when exceeded. Furthermore, the user's geographic location can be factored in to cache frequently accessed videos in that area, such as local news and footage of local events. This reduces the latency and bandwidth costs of repeatedly pulling data from remote video storage sources, improving response speed.

[0072] The second type of storage: Based on user behavior, interest tags are predicted through user profile models, personalized video preferences are analyzed, and related video data is cached in advance. Interest tags can be predicted based on data such as the user's historically frequently watched genres (such as science fiction movies), favorites, and subscription channel updates. For example, if user A frequently watches a certain series of videos, the unwatched episodes of the same series can be pre-stored in the local storage unit. If the user continuously requests similar videos (such as watching 5 episodes of a variety show within 30 minutes), the subsequent episodes can be pre-loaded into the local storage unit.

[0073] The third type of storage scenario: Storing multiple consecutive episodes of video content based on content relevance. For example, this includes TV series, course series, and anime series with strong continuity. When a user plays the first episode, the following two or three episodes can be automatically cached.

[0074] In this embodiment, videos are screened based on dimensions such as popularity, user behavior, and content continuity, and then stored in a local storage unit, thereby achieving "near-user-end acceleration."

[0075] Step S103: based on a preset watermark strategy, dynamically adding watermarks to the target video data frame by frame to generate target video data with watermarks.

[0076] The backend server reads the target video data and adds a watermark to each frame of the target video data according to the preset watermark strategy. After the watermark is added, the processed video frames are reassembled into a new video stream. The watermark strategy can define parameters such as the watermark content, style, and location, as well as how to dynamically adjust watermark properties such as transparency and color. Dynamic frame-by-frame watermarking means that the watermark of each video frame can change according to the watermark strategy, thereby improving copyright protection. By adding watermarks on the backend server, the possibility of video being privately recorded or watermarks being tampered with is reduced, thereby improving video security.

[0077] Step S104: Based on the streaming media server, the target video data with the watermark is pushed to the client for the user to watch.

[0078] In this embodiment, a streaming media server is used to transmit the processed watermarked video data to the client. The streaming media server is responsible for slicing and packaging the video data and sending it to the client according to the appropriate transmission protocol to ensure that the user can watch the video smoothly and ensure that the video data can be efficiently and reliably transmitted from the backend server to the user's device.

[0079] In this embodiment, the watermark is added on the backend server, embedded directly into the video stream. It cannot be removed or tampered with by front-end code, ensuring the integrity and security of the watermark and effectively protecting the video content. Using dynamic watermarking technology, the position and content of the watermark constantly change, making it more difficult to remove. Even if an attacker can remove the watermark in one frame, it is difficult to completely remove the watermark from the entire video. This improves the robustness of the watermark and reduces the risk of private recording of video content.

[0080] In this embodiment, the watermark is dynamically added frame by frame to the target video data based on a preset watermark strategy, which specifically includes the following processing:

[0081] For each video frame of the target video data, obtaining a horizontal position in a horizontal direction based on a preset X-axis motion function; the X-axis motion function includes a first function, a second function, and a third function of different motion periods;

[0082] For each video frame of the target video data, obtaining a vertical position in a vertical direction based on a preset Y-axis motion function; the Y-axis motion function includes a fourth function, a fifth function, and a sixth function of different motion periods; and the watermark strategy includes the X-axis motion function and the Y-axis motion function;

[0083] For each video frame of the target video data, obtaining a target watermark position corresponding to the video frame based on the horizontal position and the vertical position, and adding a watermark to the video frame at the target watermark position based on an FFmpeg processing engine;

[0084] The X-axis motion function is expressed as:

[0085] X(t)=min(max(k,w / 2+(w / 2-tw-d)×[λ1×sin(2πt / T1)+λ2×sin(2πt / T2)+λ3×sin(2πt / T3)]),w-tw-k)

[0086] Wherein, k represents a safety margin, w represents the horizontal size of the video frame, tw represents the horizontal size of the watermark image, d represents an offset, sin(2πt / T1) represents the first function, λ1 represents the weight of the first function, T1 represents the period of the first function, sin(2πt / T2) represents the second function, λ2 represents the weight of the second function, T2 represents the period of the second function, sin(2πt / T3) represents the third function, λ3 represents the weight of the third function, T3 represents the period of the third function; λ1 is greater than λ2, λ2 is greater than λ3, T1 is greater than T2, T2 is greater than T3, and T1, T2, and T3 are coprime numbers;

[0087] The Y-axis motion function is expressed as:

[0088] Y(t)=min(max(k,h / 2+(h / 2-th-d)×[λ4×cos(2πt / T4)+λ5×cos(2πt / T5)+λ6×cos(2πt / T6)]),h-th-k)

[0089] Wherein, h represents the vertical size of the video frame, th represents the vertical size of the watermark image, cos(2πt / T4) represents the fourth function, λ4 represents the weight of the fourth function, T4 represents the period of the fourth function, cos(2πt / T5) represents the fifth function, λ5 represents the weight of the fifth function, T5 represents the period of the fifth function, cos(2πt / T6) represents the sixth function, λ6 represents the weight of the sixth function, and T6 represents the period of the sixth function; the λ4 is greater than λ5, λ5 is greater than λ6, the T4 is greater than T5, T5 is greater than T6, and T4, T5 and T6 are coprime numbers.

[0090] The watermark strategy is based on the mathematical superposition of trigonometric functions (sine wave + cosine wave). The motion dimension of the watermark is two-dimensional plane motion (X axis + Y axis), and it is a combination of functions with multiple frequencies, multiple weights, and different periods, which improves the randomness of watermark generation.

[0091] In this embodiment, k can be 50 pixels, forming a boundary protection mechanism. Specifically, the X-axis boundary constraints include: left boundary: minimum 50 pixels min(...,w-tw-50); right boundary: maximum w-tw-50 pixels. Y-axis boundary constraints include: upper boundary: minimum 50 pixels min(...,h-th-50); lower boundary: maximum h-th-50 pixels.

[0092] d can be a 120-pixel safety margin, which implements an enhanced design and reduces the possibility of the watermark exceeding the screen boundary during complex movements.

[0093] The first function is the basic function in the X-axis motion function, λ1 can be 1.0, T1 can be 47, that is, the main period is 47 frames; the second function is the intermediate frequency function in the X-axis motion function, λ2 can be 0.7, T2 can be 23, that is, the secondary period is 23 frames; the third function is the high frequency function in the X-axis motion function, λ3 can be 0.3, T3 can be 13, that is, the short period is 13 frames.

[0094] The fourth function is the basic function in the Y-axis motion function, λ4 can be 1.0, T4 can be 31, that is, the main cycle is 31 frames; the fifth function is the intermediate frequency function in the Y-axis motion function, λ5 can be 0.6, T5 can be 19, that is, the main cycle is 19 frames; the sixth function is the high frequency function in the Y-axis motion function, λ6 can be 0.4, T6 can be 11, that is, the main cycle is 11 frames.

[0095] Through the X-axis motion function and the Y-axis motion function, a randomness enhancement mechanism of the watermark is formed. Specifically:

[0096] Multi-frequency superposition is achieved: frequency superposition of X-axis motion function = main wave (47 frames) + 0.7× medium wave (23 frames) + 0.3× short wave (13 frames); frequency superposition of Y-axis motion function = main wave (31 frames) + 0.6× medium wave (19 frames) + 0.4× short wave (11 frames).

[0097] A long period is achieved: the period combinations of the X-axis motion function include: 47, 23, 13 (coprime numbers); the period combinations of the Y-axis motion function include: 31, 19, 11 (coprime numbers); the least common multiple is: 47×23×13×31×19×11≈230 million; at 30fps, the actual frame period is approximately: 230 million ÷ 30≈7.7 million seconds≈89 days.

[0098] Multi-frequency is implemented: the frequencies of each layer of the X-axis motion function include: main frequency f1x = 1 / 47 ≈ 0.0213 Hz; intermediate frequency f2x = 1 / 23 ≈ 0.0435 Hz; high frequency f3x = 1 / 13 ≈ 0.0769 Hz. The frequencies of each layer of the Y-axis motion function include: main frequency f1y = 1 / 31 ≈ 0.0323 Hz; intermediate frequency f2y = 1 / 19 ≈ 0.0526 Hz; high frequency f3y = 1 / 11 ≈ 0.0909 Hz.

[0099] The motion trajectory characteristics of the watermark are formed through the X-axis motion function and the Y-axis motion function. Specifically:

[0100] Trajectory complexity characteristics: The basic motion trajectory formed by the main frequency is the basic ellipse; the trajectory deformation generated by the superposition of medium and high frequencies realizes detailed fluctuations; the interaction of different frequencies produces seemingly random motion.

[0101] Motion level features: The first level, based on the main frequency, realizes the basic elliptical motion of the watermark's motion trajectory; the second level, based on the intermediate frequency interference, realizes the elliptical deformation of the watermark's motion trajectory; the third level, based on high-frequency superposition, realizes the detailed jitter of the watermark's motion trajectory.

[0102] Visual effect prediction features: macroscopic motion is slow movement along an elliptical trajectory; mesoscopic fluctuation is deformation based on the ellipse; microscopic jitter is high-frequency detail changes.

[0103] Through the X-axis motion function and the Y-axis motion function, the watermark position appears random in a short period of time, and it is difficult to find a pattern in long-term observation, which improves the anti-tampering effect of the watermark.

[0104] Short-term randomness: The frame rate determines the refresh frequency. When the video frame rate is 30fps, 30 frames are calculated per second, which means that the watermark position changes about 30 times in 1 second. Each change is affected by six different frequencies, and six different periodic functions need to be considered simultaneously, making prediction difficult.

[0105] Long-term unpredictability: The complete repetition period is about 89 days. In practical applications, it is almost impossible to observe the complete cycle trajectory, and the position combination of the watermark at each time point is unique.

[0106] In this embodiment, a complex and random watermark motion trajectory is formed by multi-layer frequency superposition and different weight allocation, which greatly enhances the anti-tampering capability of the video content and provides strong technical support for video content protection.

[0107] The X-axis motion function and the Y-axis motion function can be preset in the FFmpeg processing engine, which is efficient in calculation and does not require additional storage. The FFmpeg processing engine can support multiple mainstream platform formats and is suitable for multiple output formats supported by the FFmpeg processing engine.

[0108] As an optional implementation of this embodiment, the process of adding a watermark to the video frame at the target watermark position based on the FFmpeg processing engine specifically includes the following processing:

[0109] For each video frame of the target video data, if the complexity of the watermark is higher than a preset threshold, the video frame is divided into sub-regions according to the real-time video content in the video frame, and the watermark is split into its component elements;

[0110] For each video frame of the target video data, dynamically adjusting the relative positions of the component elements of the watermark according to content characteristics of the plurality of sub-regions in the video frame, and synchronously adding each component element to the sub-region at a corresponding position based on the FFmpeg processing engine and the target watermark position;

[0111] For each video frame of the target video data, after all the component elements are added, the sub-regions are merged to form a video frame with a complete watermark.

[0112] For each frame of the target video data, the complexity of the watermark is first determined to be higher than a preset threshold. The threshold can be set based on actual needs and experience to distinguish the complexity of the watermark. For example, if the watermark contains rich details and complex patterns, then the watermark has a high complexity.

[0113] The process of calculating the complexity of a watermark may include: obtaining the constituent elements of the watermark, which may include Chinese characters, numbers, letters, and images; counting the number of Chinese characters, numbers, and letters in the watermark, as well as the area or number of pixels of the image portion, for example, by using image processing tools to measure the size of the image portion; and setting a corresponding weighting coefficient based on the relative contribution of each constituent element to the complexity. Generally speaking, the visual complexity of Chinese characters is higher than that of letters and numbers, while the complexity of images depends on the richness of their details. For example, the weighting coefficient of Chinese characters can be set to 0.4, the weighting coefficient of numbers to 0.2, the weighting coefficient of letters to 0.2, and the weighting coefficient of images to 0.2. Multiply the number or size of each component by its corresponding weighting coefficient to obtain the weighted complexity of each component. For example, if the watermark contains 5 Chinese characters, 3 numbers, and 7 letters, and the image accounts for 40% of the area, the weighted complexity of the Chinese characters is 5×0.4=2, the numbers is 3×0.2=0.6, the letters is 7×0.2=1.4, and the image is 40%×0.2=0.08. Adding the weighted complexities of each component gives the total complexity of the watermark. Continuing with the above example, the total complexity is 2+0.6+1.4+0.08=4.08.

[0114] If the complexity of the watermark exceeds a preset threshold, the video frame can be divided into sub-regions based on the real-time video content in the video frame. The purpose of sub-region division is to divide the video frame into multiple small blocks to better process and embed complex watermarks. At the same time, the watermark is split into component elements. Component element splitting is to divide the complete watermark into multiple components.

[0115] Specifically, based on the real-time video content, an image processing algorithm can be used to divide the video frame into multiple sub-regions. For example, a grid partitioning method can be used to divide the video frame into a regular rectangular grid, or an adaptive partitioning method can be used to adaptively divide irregularly shaped sub-regions based on characteristics such as color and texture.

[0116] The target watermark location can be the top left corner, center, or other key points of the watermark. Intelligent matching is performed based on the content characteristics of different sub-regions within the video frame (such as color, texture, object type), as well as the watermark's component elements (such as Chinese characters, numbers, letters, and images). The appropriate watermark element is placed at the optimal location to enhance watermark embedding effectiveness and copyright protection capabilities while minimizing the impact on the video viewing experience. This means that the watermark's components (such as specific text, letters, numbers, and images) remain unchanged. For example, if a watermark includes a company logo, copyright information, and a timestamp, the types of these components remain constant throughout the video. Based on the content characteristics of each sub-region within each video frame, the relative positions of the watermark's components can be dynamically adjusted, thereby improving the watermark's visibility and effectiveness while minimizing interference with the video content. Content characteristics can include texture and color.

[0117] For example, a watermark includes an image logo and text description. Based on the target watermark location, multiple sub-regions that need to contain the watermark are determined. Since image logos typically contain rich details, the image logo can be placed in a sub-region with complex textures among the multiple sub-regions that need to contain the watermark. This allows the image logo to better blend with the background, increasing robustness and making it more difficult to remove or tamper with. Since the text description requires high readability, it can be placed in a sub-region with rich colors among the multiple sub-regions that need to contain the watermark. This improves the contrast and readability of the text, ensuring that viewers can clearly see the copyright information.

[0118] Adjusting the relative positions of watermark elements based on live video content optimizes watermark embedding without affecting the overall composition, enhancing copyright protection and improving the user viewing experience. This ensures the watermark remains effective in all video scenarios, effectively preventing unauthorized recording.

[0119] Based on the FFmpeg processing engine and the target watermark location, each component element is added to the corresponding sub-region at the corresponding position. After all components are added, the sub-regions are merged to form a video frame with the complete watermark. The merging process reassembles all sub-regions into a complete video frame to ensure the watermark is displayed correctly in the video frame.

[0120] By breaking down complex watermarks into multiple components and adding them to appropriate sub-regions of the video frame, we can more precisely control the watermark's location and embedding quality, improving its visibility and robustness, making it more difficult to remove or tamper with. By dividing the video frame into sub-regions and processing the addition of component elements in parallel, we can fully leverage the performance advantages of multi-core processors, improve watermark embedding efficiency, reduce processing time, and ensure the real-time and smooth flow of video streams.

[0121] Specialized processing of complex watermarks further enhances copyright protection for video content. Even if an attacker can identify the watermark, it's difficult to completely remove or modify the complex watermark pattern, effectively preventing the unauthorized recording of video content. Flexible adjustments can be made based on the complexity of the video content and the watermark requirements to accommodate diverse application scenarios and needs. Whether it's a simple single-color watermark or a complex multi-color pattern, both can be effectively embedded and protected.

[0122] As an optional implementation manner of this embodiment, for each video frame of the target video data, before obtaining the horizontal position in the horizontal direction based on the preset X-axis motion function and obtaining the vertical position in the vertical direction based on the preset Y-axis motion function, the method further includes:

[0123] Obtaining a trust level of the user based on historical viewing data of the user corresponding to the client, and obtaining weights of the first function and the fourth function based on the trust level; the historical viewing data includes historical behavior, including whether there has been any private recording behavior, viewing duration, or interaction frequency;

[0124] For each video frame of the target video data, obtaining a weight of the second function and a weight of the fifth function based on real-time video content in the video frame;

[0125] For each video frame of the target video data, the weight of the third function is obtained based on the real-time network status between the client and the target video data; and the weight of the sixth function is obtained based on the device screen characteristics corresponding to the target video data, wherein the device screen characteristics include the screen aspect ratio.

[0126] The backend server can assess a user's trust level based on their viewing history. This includes whether they have engaged in unauthorized video recording (e.g., unauthorized video downloading, screen recording, etc.), viewing time (the average length of time a user views a video), and interaction frequency (e.g., the frequency of likes, comments, and shares). Based on this historical behavior, the user is assigned a trust level, for example, high, medium, or low. This level reflects the user's perceived riskiness of the video content. For example, if a user has never engaged in unauthorized video recording, has a long viewing time, and has a high interaction frequency, their trust level may be high.

[0127] The weights of the first and fourth functions are predefined based on different trust levels. Users with low trust levels have more noticeable watermarks, so their corresponding weights can be higher. For example, the weights for users with high trust levels can be set to λ1=0.8 and λ4=0.7; the weights for users with low trust levels can be set to λ1=1.0 and λ4=0.9.

[0128] For each video frame of the target video data, the real-time video content is analyzed to obtain the complexity of the real-time video content. For example, the complexity of the real-time video content can be calculated based on the scene complexity, the speed of the object movement, etc. The weight of the second function and the weight of the fifth function are dynamically adjusted according to the complexity of the real-time video content. In complex scenes, the weight can be increased to make the watermark more obvious; in simple scenes, the weight can be reduced to reduce interference. For example, when the complexity of the real-time video content is high, the corresponding weights can be adjusted to λ2=0.6 and λ5=0.5; when the complexity of the real-time video content is low, the corresponding weights can be adjusted to λ2=0.3 and λ5=0.2.

[0129] The weight of the third function is dynamically determined based on the real-time network conditions with the client, which may include bandwidth, packet loss rate, and other factors. When the real-time network conditions are poor, the weight can be lowered to reduce processing complexity; when the real-time network conditions are good, the weight can be increased to enhance watermark protection. For example, when the network bandwidth is low, the corresponding weight can be adjusted to w3 = 0.2; when the network bandwidth is sufficient, the corresponding weight can be adjusted to w3 = 0.5.

[0130] The weight of the sixth function is obtained based on the client's device screen characteristics, which may include screen aspect ratio, resolution, etc. Devices with different screen aspect ratios have significantly different visual effects, so the weights need to be adjusted accordingly. For example, for a device with an aspect ratio of 16:9, the corresponding weight can be adjusted to w6 = 0.4; and for a device with an aspect ratio of 4:3, the corresponding weight can be adjusted to w6 = 0.3.

[0131] Through the above steps, the weight of the watermark motion function can be dynamically adjusted based on the user's historical behavior, real-time video content, real-time network conditions, and device screen characteristics. This not only improves the effectiveness of video copyright protection, but also optimizes the user's viewing experience.

[0132] As another optional implementation in this embodiment, the weight of the first function to the weight of the sixth function may be dynamically adjusted according to time. For example:

[0133] The adjustment function of the weight of the first function can be expressed as:

[0134] weight1=1.0+0.1sin(2PIt / 100)

[0135] The adjustment function of the weight of the second function can be expressed as:

[0136] weight2=0.7+0.1cos(2PIt / 150)

[0137] The adjustment function of the weight of the third function can be expressed as:

[0138] weight3=0.3+0.1sin(2PIt / 200)

[0139] As an optional implementation manner of this embodiment, the dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes:

[0140] For each video frame of the target video data, emotion analysis and scene recognition are performed on the real-time video content of the video frame to extract emotion information and scene information corresponding to the video frame, and based on the emotion information and scene information, initial color information of the watermark is obtained; the emotion information includes joy, sadness or tension, and the scene information includes sky, ocean, forest or city;

[0141] For each video frame of the target video data, user color preference data is obtained based on the historical viewing data, and based on the user color preference data, the initial color information is adjusted to obtain target color information, so that the watermark at the target watermark position is displayed based on the target color information; the color information includes the watermark color corresponding to the video frame and the predicted color change trend of the watermark corresponding to subsequent video frames.

[0142] For each frame of the target video data, the system analyzes the emotional information (e.g., joy, sadness, tension) and scene information (e.g., sky, ocean, forest, city, etc.) contained in the frame. For example, by analyzing the color, lighting, and composition of the video frame, combined with the plot and background information of the video, it can determine whether the current scene is a cheerful indoor party or a tense chase scene, or a vast sky or a dense forest.

[0143] Based on the extracted emotion and scene information, the predefined color mapping table is searched to determine the initial color information that matches the current emotion and scene. For example, cheerful emotions can be mapped to bright yellow or orange, sad emotions can be mapped to soft blue or gray; sky scenes can be mapped to light blue, and ocean scenes can be mapped to dark blue, etc.

[0144] Based on the collected historical viewing data of users, the user's preferences for different colors are analyzed to obtain user color preference data. Historical viewing data may include records of users adjusting color settings while watching videos, viewing time of different videos, and interactive behaviors. For example, if a user frequently adjusts video settings to increase the saturation of red or prefers to watch videos with red as the primary color, then the user can be considered to have a high preference for red.

[0145] The backend server stores the relationship between the user's color preference information and the initial color information. By combining this data with the initial color information and adjusting the initial watermark color, the final target color information can be generated. For example, if the initial color information is light blue (suitable for sky scenes), but the user's preference data indicates that the user prefers warm colors, the warm color component can be appropriately increased to adjust the light blue to sky blue or a similar color.

[0146] When adding a watermark at the target watermark location, the watermark's color is set based on the adjusted target color information. For example, if the final target color information is a specific sky blue, the watermark will be displayed in that sky blue color. This ensures that the watermark is consistent with the emotion and scene of the video content while also meeting the user's color preferences.

[0147] The color information not only includes the watermark color corresponding to the current video frame, but also includes the color change trend of the watermark corresponding to the predicted subsequent video frames. In this optional embodiment, the entire target video data or a partial segment thereof can be pre-analyzed before the video is played, the emotion and scene information of each video frame can be extracted, and a time series containing emotion and scene changes can be constructed. Based on the emotion and scene time series obtained by the pre-analysis, combined with the predefined color mapping rules, the watermark color change curve of the subsequent video frames is generated. For example, if the emotion of the current video frame is cheerful and the scene is an indoor party, the initial color is bright yellow, and it is predicted that the next scene will change to outdoor natural scenery, then the color change trend can gradually transition from yellow to green or blue that is closer to nature. Based on the color change trend of the watermark corresponding to the predicted subsequent video frames, the gradient of the watermark color can be planned in advance, so that the change of the watermark color is more natural and smooth.

[0148] In this optional implementation, the most appropriate target color information can be dynamically determined for the watermark of each video frame, thereby providing a more personalized and comfortable viewing experience while protecting the video copyright.

[0149] As another optional implementation in this embodiment, the watermark color can be dynamically adjusted according to time. For example, the watermark color adjustment function can be expressed as:

[0150] fontcolor='gray@(0.2+0.1sin(2PIt / 60))'

[0151] Among them, gray@ represents the syntax for setting color and transparency in the FFmpeg processing engine. Gray indicates that the color of the watermark is gray. Through the watermark color adjustment function, the color transparency or brightness of the watermark can be dynamically adjusted to change periodically between 0.1 and 0.3 with a period of 60 seconds.

[0152] As an optional implementation manner of this embodiment, the dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes:

[0153] Obtaining an initial transparency of the watermark based on a device type of a device screen corresponding to the client and metadata of the target video data;

[0154] Obtaining environmental display characteristics of the device screen corresponding to the client, wherein the environmental display characteristics include user viewing distance, screen brightness, and ambient light information;

[0155] For each video frame of the target video data, the initial transparency is adjusted based on the environmental display characteristics to obtain a target transparency, so that the watermark at the target watermark position is displayed based on the target transparency.

[0156] Different device types and target video data metadata have different requirements for watermark visibility. For example, TV screens are typically larger and viewed from a distance, so higher transparency is required to prevent the watermark from being too prominent. Smartphone screens, on the other hand, are smaller and viewed from a closer distance, so lower transparency can be appropriate. Therefore, the initial watermark transparency can be determined based on the client device type (e.g., smartphone, tablet, TV) and the target video data metadata (e.g., video type, scene, etc.).

[0157] The metadata of the target video can help understand the characteristics of the video content. For example, action movies may contain many fast-moving scenes, so higher transparency can reduce the interference of the watermark on the dynamic image. Documentaries, on the other hand, may focus more on accurately conveying information, so lower transparency helps ensure the visibility of the watermark.

[0158] Collect information about the client device's ambient display characteristics, including user viewing distance, screen brightness, and ambient light. The viewing distance affects the watermark's visibility. At farther distances, the watermark may need to be more transparent to blend in; at closer distances, the transparency may be lower. Screen brightness and ambient light also affect the watermark's display quality. In bright environments, to ensure the watermark's visibility, the transparency may need to be reduced; in darker environments, the transparency can be increased appropriately.

[0159] Based on the collected environmental display characteristics, the initial transparency is dynamically adjusted to achieve the target transparency. The backend server presets the adjustment relationship between environmental display characteristics and initial transparency. For example, if the ambient light is bright and the screen brightness is high, the watermark transparency can be reduced to ensure that the watermark is clearly displayed on the screen. Conversely, in darker environments, the transparency can be increased to make the watermark blend more naturally into the picture and avoid disrupting the user's viewing experience.

[0160] When adding a watermark at the target watermark location, the watermark's transparency is set based on the calculated target transparency. By dynamically adjusting the watermark's transparency based on device characteristics and viewing environment, we ensure optimal display and copyright protection across different devices and environments, optimizing both the watermark's display quality and user experience.

[0161] In this optional implementation, obtaining the environmental display characteristics of the device screen corresponding to the client, wherein the environmental display characteristics include user viewing distance, screen brightness, and ambient light information, includes:

[0162] Analyzing the historical viewing data based on a machine learning algorithm to identify the user's current viewing behavior pattern; the historical viewing data includes viewing time, viewing location, viewing device type, type of video viewed, and whether the user has ever adjusted the watermark transparency or reported the watermark for obstructing content;

[0163] Based on the behavior prediction model and the current viewing behavior pattern, the device type and the ambient light information used by the user when currently watching the video are predicted; and the user viewing distance and the screen brightness sent by the client are received.

[0164] A user's viewing history includes viewing time, viewing location, viewing device type, video type, and whether the user has adjusted the watermark transparency or reported watermarks for obscuring content. For example, a user might watch an action movie on a smartphone at night and frequently adjust the watermark transparency for a better viewing experience. Backend servers can use machine learning algorithms to analyze this historical data and identify the user's current viewing behavior patterns. For example, cluster analysis revealed that users tend to use larger screen devices when watching movies at night and are more sensitive to the presence of watermarks.

[0165] Based on the behavior prediction model and the current viewing behavior pattern, the device type and ambient light information of the user are predicted. For example, if the user usually uses a TV to watch movies at night, it can be predicted that the current viewing device is a TV and the ambient light is likely dim.

[0166] The client can obtain the user's viewing distance and screen brightness information through sensors and send this information to the backend server. For example, a smartphone can estimate the user's viewing distance and current screen brightness using the front camera and light sensor respectively.

[0167] By integrating the user's historical behavior and current viewing environment, we can predict and obtain the environmental display characteristics related to the user's viewing context, providing a basis for the subsequent dynamic adjustment of the watermark transparency, thereby optimizing the user's viewing experience.

[0168] As another optional implementation in this embodiment, the transparency of the watermark can be dynamically adjusted according to time. For example, the adjustment function of the transparency of the watermark can be expressed as:

[0169] fontopacity='gray@(0.2+0.1cos(2PIt / 45))'

[0170] Among them, gray@ represents the syntax for setting color and transparency in the FFmpeg processing engine. Gray indicates that the color of the watermark is gray. Through the transparency adjustment function of the watermark, the transparency of the watermark can be dynamically adjusted to change periodically between 0.1 and 0.3 with a period of 45 seconds.

[0171] Based on the same technical concept, the present application also provides a protection system applying the above-mentioned video content protection method, such as Figure 2 As shown, the video content protection system 200 mainly includes a client 201 and a back-end server 202; the client 201 includes a video playback module 2011, a user interaction module 2012 and a WebSocket client 2013, and the back-end server 202 includes a WebSocket server 2021, an FFmpeg processing engine 2022 and a TCP server 2023.

[0172] The video playing module 2011 is used to send a video access request to the back-end server 202 and play the target video data with watermark pushed by the back-end server 202.

[0173] The user interaction module 2012 is used to provide a user interface and receive user operation instructions.

[0174] The WebSocket client 2013 is used to establish a connection with the WebSocket server 2021 for real-time communication.

[0175] The WebSocket server 2021 is used to receive the connection request of the WebSocket client 2013 and establish a real-time communication connection for the WebSocket client 2013.

[0176] The FFmpeg processing engine 2022 is used to add a dynamic watermark to the target video data corresponding to the video access request.

[0177] The TCP server 2023 is used to transmit the target video data with the watermark to the WebSocket server 2021.

[0178] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0179] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0180] Based on the same technical concept, the present application also provides an electronic device, such as Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302 , and may further include an information input / information output (I / O) interface 303 , one or more communication components 304 , and a communication bus 305 .

[0181] The processor 301 is used to control the overall operation of the electronic device 300 to complete all or part of the steps in the above-mentioned video content protection method. The memory 302 is used to store various types of data to support the operation of the electronic device 300. Such data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as one or more of static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0182] The I / O interface 303 provides an interface between the processor 301 and other interface modules, which may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 304 is used to test wired or wireless communication between the electronic device 300 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more thereof, may include: a Wi-Fi component, a Bluetooth component, and an NFC component.

[0183] Communication bus 305 may include a path for transmitting information between the aforementioned components. Communication bus 305 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, for example. Communication bus 305 may be divided into an address bus, a data bus, a control bus, and the like.

[0184] The electronic device 300 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the video content protection method provided in the above embodiment.

[0185] The electronic device 300 may include, but is not limited to, mobile terminals such as digital broadcast receivers, PDAs (Personal Digital Assistants), and PMPs (Portable Multimedia Players), and fixed terminals such as digital TVs and desktop computers, and may also be servers.

[0186] Based on the same technical concept, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned video content protection method are implemented.

[0187] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program codes.

[0188] The terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0189] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0190] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0191] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for protecting video content, characterized in that: Applicable to backend servers, including: Receive a video access request sent by a client, and obtain a video identifier corresponding to the video access request; If the target video data corresponding to the video identifier is stored in the local storage unit, the target video data in the local storage unit is read; if the target video data is not stored in the local storage unit, a request for obtaining the target video data is sent to a video storage object source, and the target video data returned by the video storage object source is received; Based on a preset watermark strategy, dynamically adding watermarks to the target video data frame by frame to generate target video data with watermarks; Based on the streaming media server, the target video data with the watermark is pushed to the client for the user to watch; The method of dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy includes: For each video frame of the target video data, obtaining a horizontal position in a horizontal direction based on a preset X-axis motion function; the X-axis motion function includes a first function, a second function, and a third function of different motion periods; For each video frame of the target video data, obtaining a vertical position in a vertical direction based on a preset Y-axis motion function; the Y-axis motion function includes a fourth function, a fifth function, and a sixth function of different motion periods; and the watermark strategy includes the X-axis motion function and the Y-axis motion function; For each video frame of the target video data, obtaining a target watermark position corresponding to the video frame based on the horizontal position and the vertical position, and adding a watermark to the video frame at the target watermark position based on an FFmpeg processing engine; The X-axis motion function is expressed as: X(t)=min(max(k,w / 2+(w / 2-tw-d)×[λ1×sin(2πt / T1)+λ2×sin(2π t / T2)+λ3×sin(2πt / T3)]),w-tw-k) Wherein, k represents a safety margin, w represents the horizontal size of the video frame, tw represents the horizontal size of the watermark image, d represents an offset, sin(2πt / T1) represents the first function, λ1 represents the weight of the first function, T1 represents the period of the first function, sin(2πt / T2) represents the second function, λ2 represents the weight of the second function, T2 represents the period of the second function, sin(2πt / T3) represents the third function, λ3 represents the weight of the third function, T3 represents the period of the third function; λ1 is greater than λ2, λ2 is greater than λ3, T1 is greater than T2, T2 is greater than T3, and T1, T2, and T3 are coprime numbers; The Y-axis motion function is expressed as: Y(t)=min(max(k,h / 2+(h / 2-th-d)×[λ4×cos(2πt / T4)+λ5×cos(2π t / T5)+λ6×cos(2πt / T6)]),h-th-k) Wherein, h represents the vertical size of the video frame, th represents the vertical size of the watermark image, cos(2πt / T4) represents the fourth function, λ4 represents the weight of the fourth function, T4 represents the period of the fourth function, cos(2πt / T5) represents the fifth function, λ5 represents the weight of the fifth function, T5 represents the period of the fifth function, cos(2πt / T6) represents the sixth function, λ6 represents the weight of the sixth function, and T6 represents the period of the sixth function; the λ4 is greater than λ5, λ5 is greater than λ6, the T4 is greater than T5, T5 is greater than T6, and T4, T5 and T6 are coprime numbers.

2. The protection method according to claim 1, characterized in that: The method of adding a watermark to the video frame at the target watermark position based on the FFmpeg processing engine includes: For each video frame of the target video data, if the complexity of the watermark is higher than a preset threshold, the video frame is divided into sub-regions according to the real-time video content in the video frame, and the watermark is split into its component elements; For each video frame of the target video data, dynamically adjusting the relative positions of the component elements of the watermark according to content characteristics of the plurality of sub-regions in the video frame, and synchronously adding each component element to the sub-region at a corresponding position based on the FFmpeg processing engine and the target watermark position; For each video frame of the target video data, after all the component elements are added, the sub-regions are merged to form a video frame with a complete watermark.

3. The protection method according to claim 1, characterized in that: For each video frame of the target video data, before obtaining the horizontal position in the horizontal direction based on the preset X-axis motion function and obtaining the vertical position in the vertical direction based on the preset Y-axis motion function, the method further includes: Obtaining a trust level of the user based on historical viewing data of the user corresponding to the client, and obtaining weights of the first function and the fourth function based on the trust level; the historical viewing data includes historical behavior, including whether there has been any private recording behavior, viewing duration, or interaction frequency; For each video frame of the target video data, obtaining a weight of the second function and a weight of the fifth function based on real-time video content in the video frame; For each video frame of the target video data, the weight of the third function is obtained based on the real-time network status between the client and the target video data; and the weight of the sixth function is obtained based on the device screen characteristics corresponding to the target video data, wherein the device screen characteristics include the screen aspect ratio.

4. The protection method according to claim 3, characterized in that: The method of dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes: For each video frame of the target video data, emotion analysis and scene recognition are performed on the real-time video content of the video frame to extract emotion information and scene information corresponding to the video frame, and based on the emotion information and scene information, initial color information of the watermark is obtained; the emotion information includes joy, sadness or tension, and the scene information includes sky, ocean, forest or city; For each video frame of the target video data, user color preference data is obtained based on the historical viewing data, and based on the user color preference data, the initial color information is adjusted to obtain target color information, so that the watermark at the target watermark position is displayed based on the target color information; the color information includes the watermark color corresponding to the video frame and the predicted color change trend of the watermark corresponding to subsequent video frames.

5. The protection method according to claim 4, characterized in that: The method of dynamically adding a watermark to the target video data frame by frame based on a preset watermark strategy further includes: Obtaining an initial transparency of the watermark based on a device type of a device screen corresponding to the client and metadata of the target video data; Obtaining environmental display characteristics of the device screen corresponding to the client, wherein the environmental display characteristics include user viewing distance, screen brightness, and ambient light information; For each video frame of the target video data, the initial transparency is adjusted based on the environmental display characteristics to obtain a target transparency, so that the watermark at the target watermark position is displayed based on the target transparency.

6. The protection method according to claim 5, characterized in that: The acquiring of the environmental display characteristics of the device screen corresponding to the client, wherein the environmental display characteristics include user viewing distance, screen brightness, and ambient light information, includes: Analyzing the historical viewing data based on a machine learning algorithm to identify the user's current viewing behavior pattern; the historical viewing data includes viewing time, viewing location, viewing device type, type of video viewed, and whether the user has ever adjusted the watermark transparency or reported the watermark for obstructing content; Based on the behavior prediction model and the current viewing behavior pattern, the device type and the ambient light information used by the user when currently watching the video are predicted; and the user viewing distance and the screen brightness sent by the client are received.

7. A protection system using the video content protection method according to any one of claims 1 to 6, characterized in that: It includes a client and a back-end server; the client includes a video playback module, a user interaction module and a WebSocket client, and the back-end server includes a WebSocket server, an FFmpeg processing engine and a TCP server; The video playing module is used to send a video access request to the back-end server and play the target video data with the watermark pushed by the back-end server; The user interaction module is used to provide a user interface and receive user operation instructions; The WebSocket client is used to establish a connection with the WebSocket server for real-time communication; The WebSocket server is configured to receive a connection request from the WebSocket client and establish a real-time communication connection with the WebSocket client; The FFmpeg processing engine is used to add a dynamic watermark to the target video data corresponding to the video access request; The TCP server is used to transmit the target video data with the watermark to the WebSocket server.

8. An electronic device, characterized in that: comprising a processor and a memory, wherein the processor is coupled to the memory; The processor is configured to execute the computer program stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The method comprises a computer program or an instruction, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video watermark processing method and device, electronic equipment and storage medium

    CN115314734A

  • Video watermark embedding method and device, electronic equipment and storage medium

    CN118317165A