Video watermark adding method and device and storage medium

By recognizing video clips with voice and image, automatically generating watermarks that fit the video style, the problem of single watermark addition method in the existing technology is solved, and the visual effect and user experience of the video are improved.

CN120343169APending Publication Date: 2025-07-18JUZHANG INTERACTIVE TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510486414.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video watermark addition method cannot be customized according to the video content and style, which affects the visual effect and user experience of the video.

Method used

By performing speech recognition and image recognition on video clips, voice feature information and image feature information are extracted, and corresponding watermark materials are determined in the case of matching, and watermarks that fit the video style are automatically generated.

Benefits of technology

It improves the overall visual effect and viewing of the video, meets the personalized needs of users, and improves the professionalism of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343169A_ABST
    Figure CN120343169A_ABST
Patent Text Reader

Abstract

The invention discloses a video watermark adding method and device and a storage medium, and the method comprises the steps: carrying out the voice recognition of a current video clip, obtaining the corresponding current voice feature information, and carrying out the image recognition of the current video clip, and obtaining the corresponding current image feature information; under the condition of determining that the current voice feature information is matched with the current image feature information, obtaining matched and consistent current feature information, and determining a current watermark material matched with the current feature information; and adding a corresponding watermark to the current video clip according to the current watermark material. Therefore, the watermark which is more suitable for the video style is automatically generated, the overall visual effect of the video is improved, certainly, the personalized requirements of users are met, and the ornamental value and the speciality of the video are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, and storage medium for video watermark addition. Background Art

[0002] With the development of digital image technology, various digital videos have emerged in people's lives. Among them, in some scenarios such as copyright protection, leakage traceability, document authenticity verification, and enhancing the interest of videos, watermarks can be added to videos, that is, digital information (i.e., digital watermarks) can be hiddenly embedded into carrier files such as audio - videos and pictures without affecting the visual quality and integrity of video pictures, images, etc.

[0003] Currently, some fixed - unchanged image or text information, such as copyright information, time stamps, brand logos, artistic patterns, etc., can be placed at set positions in the video to form corresponding video watermarks. However, this way of adding video watermarks is relatively single and cannot be customized according to the specific content and style of the video. This results in the watermark may not match the video content, affecting the visual effect of the video and the user experience. Summary of the Invention

[0004] To solve the above - mentioned problems existing in the prior art, the present invention provides a method, device, and storage medium for video watermark addition. The technical problems to be solved by the present invention are achieved through the following technical solutions:

[0005] In the first aspect of the embodiments of the present invention, a method for video watermark addition is provided, including:

[0006] Performing speech recognition on the current video segment to obtain corresponding current speech feature information, and performing image recognition on the current video segment to obtain corresponding current image feature information;

[0007] When it is determined that the current speech feature information matches the current image feature information, obtaining the current feature information that is matched and consistent, and determining the current watermark material that matches the current feature information;

[0008] Performing corresponding watermark addition to the current video segment according to the current watermark material.

[0009] In the second aspect of the embodiments of the present invention, a device for video watermark addition is provided. The device includes a processor and a memory storing program instructions. The processor is configured to execute the above - mentioned method for video watermark addition when executing the program instructions.

[0010] In the third aspect of the embodiments of the present invention, a storage medium is provided, storing program instructions, and the program instructions, when running, execute the above - mentioned method for video watermark addition.

[0011] Advantages of the present invention:

[0012] Perform speech and image recognition on the current video segment respectively to obtain the corresponding current speech feature information and current image feature information, and when the two match, determine the corresponding current watermark material and add it to the current video. In this way, a watermark that is more in line with the video style is automatically generated, improving the overall visual effect of the video. Of course, it also meets the personalized needs of users, enhancing the viewing and professionalism of the video.

[0013] Other features and advantages of the present invention will be described in the following specification, and some of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings.

[0014] The following will further describe the technical solutions of the present invention in detail through the drawings and embodiments. Brief Description of the Drawings

[0015] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0016] Figure 1 It is a schematic flow chart of a method for adding a video watermark provided by an embodiment of the present invention;

[0017] Figure 2 It is a schematic flow chart of a method for adding a video watermark provided by an embodiment of the present invention;

[0018] Figure 3 It is a schematic flow chart of a method for adding a video watermark provided by an embodiment of the present invention;

[0019] Figure 4 It is a schematic structural diagram of a device for adding a video watermark provided by an embodiment of the present invention;

[0020] Figure 5 It is a schematic structural diagram of a device for adding a video watermark provided by an embodiment of the present invention;

[0021] Figure 6 It is a schematic structural diagram of a device for adding a video watermark provided by an embodiment of the present invention. Detailed Embodiments

[0022] The following will further describe the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0023] Without affecting the visual quality and integrity of video images, pictures, etc., digital information (i.e., digital watermark) is embedded in carrier files such as audio - videos and pictures in a hidden manner, which is applicable to scenarios such as copyright protection, leakage traceability, file authenticity verification, and enhancing video interestingness. In the embodiments of the present invention, watermarks that are more in line with the video style can be automatically generated according to various features of the video, thereby meeting the personalized needs of users and improving the visual effect and user experience of the video.

[0024] As Figure 1 shown, the first aspect of the embodiments of the present invention provides a method for adding video watermarks, including the following steps:

[0025] Step 101, perform speech recognition on the current video segment to obtain the corresponding current speech feature information, and perform image recognition on the current video segment to obtain the corresponding current image feature information.

[0026] After devices such as mobile phones, tablets, and terminals obtain a video file, the video file can be segmented into one segment, two segments, or multiple segments. In some embodiments, the video file is obtained and segmented according to the set scene information or shots to obtain each video segment. For example: segment the video file according to indoor scenes and outdoor scenes. Or, segment the video file according to the set number of shots, or segment the video file according to the set duration, etc., and specific examples will not be listed one by one.

[0027] After the video file is segmented to obtain each video segment, various features of each video segment can be analyzed. Taking any video segment as the current video segment as an example, speech recognition can be performed on the current video segment to obtain the corresponding current speech feature information, and image recognition can be performed on the current video segment to obtain the corresponding current image feature information.

[0028] Among them, performing speech recognition on the current video segment to obtain the corresponding current speech feature information may include: extracting the corresponding audio track information from the current video segment, then performing corresponding pre - processing and speech recognition, and extracting keywords from the recognition results, so as to obtain the corresponding current speech feature information. Among them, natural language processing (NLP) tools such as NLTK, spaCy, or BERT can be used to extract keywords from the recognition results to obtain the corresponding current speech feature information.

[0029] And performing image recognition on the current video segment to obtain the corresponding current image feature information may include: determining the key frames in the current video segment and performing image recognition on the key frames to obtain the corresponding current image feature information.

[0030] Key frames refer to the representative and important frames in a video, usually containing important visual information or scene changes. Therefore, based on scene changes, content analysis, etc., the key frames in the current video segment can be determined. In some embodiments, the pixel features of each frame in the current video segment can be obtained, and the pixel feature difference values between adjacent frames can be obtained; according to the pixel feature difference values, the key frame images in the current video segment can be determined, and image recognition can be performed on the key frame images to obtain the corresponding current image feature information. For example: if the pixel feature difference value between the current frame image and the previous frame image is greater than the set value, the current frame image can be determined as the key frame image in the current video segment.

[0031] There are also many ways to perform image recognition on key frame images. First, perform image preprocessing and image recognition on the key frame images. Then, a suitable image recognition model can be used, such as a pre-trained convolutional neural network (CNN), such as ResNet, VGG, Inception, etc., to extract the image features of the key frames. Then, a pre-trained model such as CLIP can be used for image-to-text mapping to extract keywords from the image features, that is, the corresponding current image feature information can be obtained.

[0032] Step 102, in the case where it is determined that the current speech feature information matches the current image feature information, obtain the current feature information that matches, and determine the current watermark material that matches the current feature information.

[0033] The watermark material library can be saved in the device or the cloud device. The watermark material library can be indexed according to the corresponding relationship between the feature information and the watermark material. Thus, various feature extractions are performed on the current video segment, and then matching is performed. In this way, if it is obtained that the current speech feature information matches the current image feature information, the current feature information that matches can be obtained. For example: if the first keyword extracted by speech recognition of the current video segment matches the second keyword extracted by image recognition of the current video segment, it can be determined that the current speech feature information matches the current image feature information, and then the first keyword and the second keyword can be unified to obtain the matching keyword, that is, the current feature information.

[0034] If the watermark material library is saved locally in the device, the current watermark material that matches the current feature information can be determined in the saved watermark material library; if the cloud device, such as a cloud platform, saves the watermark material library, the current feature information can be sent to the cloud device. Thus, the cloud device determines the current watermark material that matches the current feature information in the saved watermark material library and sends it to the device. Thereby, the device determines the current watermark material that matches the current feature information.

[0035] Step 103, perform corresponding watermark addition to the current video segment according to the current watermark material.

[0036] In some embodiments, the current watermark material is added as a watermark to each frame image of the current video segment, wherein the position of the watermark for each frame image can be determined according to the position coordinate information input by the user, or according to the watermark style template selected by the user, or according to the recommendation after image analysis of the current video segment by the watermark data model.

[0037] Alternatively, in some embodiments, the current watermark material and the current video segment can be converted into the current prompt word information conforming to natural language through regular expressions, and then, based on the watermark data model according to the current watermark material, the current video segment, and the current prompt word information, the current watermark material is added as a watermark to a set position in each frame of the current video segment.

[0038] Since the current video segment not only has voice features and image features, but also can have other features, such as: emotional features, visual features, scene features, motion pattern features, etc., therefore, in some embodiments, adding the corresponding watermark to the current video segment according to the current watermark material includes: determining the current scene information corresponding to the current video segment; generating the corresponding first watermark material based on the training-generated material model according to the current watermark material and the current scene information; adding the first watermark material as a watermark to each frame of the current video segment.

[0039] Among them, the scene information can include one or more of: emotional information, image style type information, scene information, motion pattern information, etc. Therefore, in some embodiments, determining the current scene information corresponding to the current video segment includes: after performing sound preprocessing on the voice in the current video segment, extracting the corresponding voice signal features; classifying the voice signal features based on machine learning or deep learning algorithms to obtain the corresponding current emotional information. For example: performing sound preprocessing on the voice in the current video segment, including noise reduction, echo cancellation, etc., and then extracting the corresponding voice signal features, such as pitch, intensity, speech rate, rhythm, etc. These features can reflect the emotional state of the speaker. Based on machine learning or deep learning algorithms, such as support vector machines, convolutional neural networks, etc., classifying the extracted voice signal features to identify the corresponding current emotional information, which can include: happy, sad, angry, etc.

[0040] In some embodiments, determining the current scene information corresponding to the current video segment includes: after performing image preprocessing on the key frame images in the current video segment, extracting the corresponding visual features; classifying the visual features based on machine learning or deep learning algorithms to obtain the corresponding current style type information.

[0041] For example, perform image preprocessing on the key frame images in the current video segment, including cropping, scaling, normalization, etc., to improve the accuracy of recognition. Then, extract the visual features of the processed images, such as color distribution, texture, shape, etc. These features can reflect the style type of the images. Thus, based on machine learning or deep learning algorithms, such as convolutional neural networks, the extracted visual features can be classified to identify the corresponding current style type information, which may include: abstract style, realistic style, cartoon style, etc.

[0042] Similarly, in some embodiments, other current scene information corresponding to the current video segment, such as current scene information, or current motion state information, etc., can also be identified based on machine learning or deep learning algorithms, and will not be elaborated one by one here.

[0043] With the development of image processing technology, there are many large models good at image processing. These models are generated through multiple trainings, that is, through multiple watermark material samples, video samples, and scene information samples, and are trained and learned based on machine learning or deep learning algorithms to obtain the corresponding large models. Here, the large model can be a material model. Thus, in this embodiment, the current watermark material and the current scene information can be input into the material model to obtain the first watermark material obtained by the material model.

[0044] Then, add the first watermark material as a watermark to each frame of the current video segment. In some embodiments, it may include: converting the first watermark material and the current video segment into the first prompt word information that conforms to natural language through regular expressions; based on the first watermark material, the current video segment, and the first prompt word information, and based on the trained watermark data model, add the first watermark material as a watermark to the set positions of each frame of the current video segment.

[0045] Use regular expressions to convert the features corresponding to the first watermark material and the current video segment into words in natural language, and then generate a paragraph that conforms to natural language in combination with regular expressions. For example: Help me draw: The picture style is abstract painting style, a crane floating in the background of water and trees, traditional Chinese scenery, light white and light green, ultra-high definition image, abstract impressionism, motion blur. This paragraph is the corresponding first prompt word information.

[0046] Thus, the first prompt word information can be input into the watermark data model. The watermark data model can generate the corresponding watermark according to the first prompt word information and add the watermark to each frame of the current video segment. Of course, the watermark data model can also be obtained through training and learning based on watermark material samples and video segment samples using machine learning or deep learning algorithms.

[0047] In this embodiment, there are various processes for determining the position where the watermark is added to each frame of the current video clip, that is, the set position. Therefore, in some embodiments, adding the first watermark material as a watermark to the set position of each frame of the current video clip further includes: determining the set position in each frame according to the position coordinate information input by the user; or, determining the set position in each frame according to the watermark style template selected by the user; or, determining the recommended position obtained after the watermark data model performs image analysis on the current video clip as the set position in each frame.

[0048] For example: The user can directly input specific position coordinate information to specify the position of the watermark in the video frame. Or, the device can provide multiple preset watermark style templates, and each template contains the coordinates and style of the watermark. In this way, the user can select a suitable template according to their own needs. Thus, when the material model adds the watermark, it can add the watermark according to the coordinates and style in the template.

[0049] Or, the material model can perform image feature analysis on the key frames in the current video clip, automatically select a suitable watermark position, and determine it as the recommended position. For example: After the material model performs image feature analysis, it can select relatively blank or unobtrusive positions in each frame image and determine them as the recommended positions, so as to avoid blocking important video content. Or, the material model determines the image feature positions corresponding to other set keywords in the current video clip and determines them as the recommended positions. In this way, the visual relevance can be further increased.

[0050] Therefore, the device can perform corresponding watermark addition to the current video clip according to the current watermark material, or perform corresponding watermark addition to the current video clip according to the first watermark material generated based on the current watermark material and the current scene information. Thus, the automatic addition of watermarks that fit the video style is realized.

[0051] It can be seen that in the embodiments of the present invention, speech and image recognition are respectively performed on the current video clip to obtain the corresponding current speech feature information and current image feature information. When the two match, the corresponding current watermark material is determined and added to the current video. In this way, watermarks that are more in line with the video style are automatically generated, improving the overall visual effect of the video. Of course, it also meets the personalized needs of users, improving the viewing and professionalism of the video. And, multiple watermark addition strategies can be provided, such as large model analysis and position determination, increasing the flexibility and diversity of watermark addition.

[0052] Next, the operation process will be integrated into specific embodiments to illustrate the video watermark addition process provided by the embodiments of the present invention.

[0053] In one embodiment of the present invention, the device can be a user terminal such as a mobile phone, a computer, etc. The pre-trained watermark data model has been saved in the device, and a watermark material library has also been saved. As Figure 2 shown, the method for adding video watermark includes the following steps:

[0054] Step 201, the device obtains the video text, and segments the video file according to the set scene information to obtain each video segment.

[0055] Step 202, the device determines a video segment as the current video segment.

[0056] Step 203, after the device extracts the corresponding audio track information from the current video segment, it performs corresponding preprocessing and speech recognition, and extracts the first keyword from the recognition result.

[0057] Here, the first keyword is the current speech feature information.

[0058] Step 204, the device determines the key frame image in the current video segment, performs image preprocessing and image recognition on the key frame image, and extracts the second keyword from the recognition result.

[0059] Similarly, the second keyword is the current image feature information.

[0060] Step 205, determine whether the first keyword matches the second keyword? If so, execute step 206, otherwise, the process ends.

[0061] If the first keyword is the same as the second keyword, it can be determined that the first keyword matches the second keyword. Or, if the semantic similarity rate between the first keyword and the second keyword is greater than the set value, such as 70%, 80%, etc., it can also be determined that the first keyword matches the second keyword. Of course, other methods for determining the consistency of the two keywords can also be applied here, and will not be listed one by one.

[0062] Step 206, the device obtains the current keyword that matches and determines the current watermark material that matches the current keyword from the saved watermark material library.

[0063] Of course, the current keyword is the current feature information. And there are multiple ways for the device to obtain the current keyword that matches. The first keyword or the second keyword can be determined as the current keyword, or the current keyword that matches the same part of the semantics of the first keyword and the second keyword can be determined, etc.

[0064] Step 207, the device determines the set position in each frame according to the position coordinate information input by the user.

[0065] In step 208, the device adds the current watermark material as a watermark to a set position in each frame of the current video segment based on the current watermark material, the current video segment, and the watermark data model.

[0066] Step 209: Is each video segment the current video segment? If so, the process ends, otherwise, returns to step 202.

[0067] It can be seen that in this embodiment, the device performs voice and image recognition on the current video clip respectively, obtains the corresponding first keyword and second keyword, and when the two match, determines the corresponding current watermark material and adds it to the current video. In this way, a watermark that is more in line with the style of the video is automatically generated, which improves the overall visual effect of the video. Of course, it also meets the personalized needs of users and improves the viewing and professionalism of the video.

[0068] In one embodiment of the present invention, the device may be a user terminal or a cloud device, which not only stores the pre-trained watermark data model, but also stores the pre-trained material model and, of course, the watermark material library. Figure 3 As shown, the method for adding a video watermark comprises the following steps:

[0069] Step 301: The device obtains a video text and segments the video file according to a set duration to obtain each video segment.

[0070] Step 302: The device determines a video segment as a current video segment.

[0071] Step 303: After extracting the corresponding audio track information from the current video clip, the device performs corresponding preprocessing, performs speech recognition to extract the first keyword, and extracts the corresponding speech signal features.

[0072] Step 304: The device classifies the voice signal features based on a machine learning algorithm to obtain corresponding current emotion information.

[0073] Step 305: The device determines a key frame image in the current video clip, and performs image preprocessing and image recognition on the key frame image to extract a second keyword and a corresponding visual feature.

[0074] In step 306, the device uses a convolutional neural network algorithm to classify the extracted visual features and identify the corresponding current style type information.

[0075] Step 307, determine whether the first keyword matches the second keyword. If so, execute step 308, otherwise, the process ends.

[0076] Step 308: The device obtains the currently matched keyword and determines the current watermark material that matches the current keyword from the saved watermark material library.

[0077] Step 309, the device generates a corresponding first watermark material based on the current watermark material, the current emotion information, and the current style type information, and based on the trained material model.

[0078] Step 310, the device obtains the recommended position obtained by the watermark data model after performing image analysis on the current video segment and determines it as the set position in each frame.

[0079] Step 311, the device converts the first watermark material and the current video segment into the first prompt word information that conforms to natural language through regular expressions.

[0080] Step 312, the device adds the first watermark material as a watermark to the set position in each frame of the current video segment based on the set position, the first watermark material, the current video segment, and the first prompt word information, based on the watermark data model.

[0081] Step 313, is each video segment the current video segment? If so, the process ends; otherwise, return to Step 302.

[0082] It can be seen that in this embodiment, the device performs speech and image recognition on the current video segment respectively, not only obtains the corresponding first keyword and second keyword, but also can determine the corresponding current emotion information and current style type information. In this way. And in the case where the first keyword and the second keyword match, the corresponding current watermark material is determined, and the first watermark material that matches the current watermark material, the current emotion information, and the current style type information is obtained through the large model, and thus, added to the current video, so that a watermark that is more in line with the video style is automatically generated, improving the overall visual effect of the video. Of course, it also meets the personalized needs of users, improving the ornamental value and professionalism of the video. And various watermark addition strategies can be provided, for example: large model analysis, position determination, increasing the flexibility and diversity of watermark addition.

[0083] According to the above process of adding video watermarks, a device for adding video watermarks can be constructed. Figure 4 A device for adding video watermarks provided by an embodiment of the present invention can be applied to devices such as mobile phones, computers, cloud platforms, etc. As Figure 4 shown, the device 400 includes: an identification and extraction module 410, a matching and determination module 420, and a watermark addition module 430.

[0084] The recognition and extraction module 410 is configured to perform speech recognition on the current video segment to obtain corresponding current speech feature information, and perform image recognition on the current video segment to obtain corresponding current image feature information.

[0085] The matching and determination module 420 is configured to obtain the current feature information that matches when it is determined that the current speech feature information matches the current image feature information, and determine the current watermark material that matches the current feature information.

[0086] The watermark addition module 430 is configured to perform corresponding watermark addition to the current video segment according to the current watermark material.

[0087] In some embodiments, it further includes:

[0088] The video segmentation module is configured to obtain a video file and segment the video file according to set scene information or shots to obtain each video segment.

[0089] In some embodiments, the recognition and extraction module 410 includes:

[0090] The key frame determination unit is configured to obtain the pixel features of each frame in the current video segment and obtain the pixel feature difference value between adjacent frames; according to the pixel feature difference value, determine the key frame image in the current video segment.

[0091] The first image recognition unit is configured to perform image recognition on the key frame image to obtain corresponding current image feature information.

[0092] In some embodiments, the watermark addition module 430 includes:

[0093] The information determination unit is configured to determine the current scene information corresponding to the current video segment.

[0094] The material generation unit is configured to generate a corresponding first watermark material based on the current watermark material, the current scene information, and a trained material model.

[0095] The watermark addition unit is configured to add the first watermark material as a watermark to each frame of the current video segment.

[0096] In some embodiments, the information determination unit is specifically configured to perform sound preprocessing on the speech in the current video segment, extract corresponding speech signal features; classify the speech signal features based on machine learning or deep learning algorithms to obtain corresponding current emotion information.

[0097] In some embodiments, the information determination unit is specifically configured to perform image preprocessing on the key frame images in the current video segment, extract the corresponding visual features, and classify the visual features based on machine learning or deep learning algorithms to obtain the corresponding current style type information.

[0098] In some embodiments, the watermark addition unit is specifically configured to convert the first watermark material and the current video segment into the first prompt word information conforming to natural language through regular expressions, and based on the trained watermark data model according to the first watermark material, the current video segment, and the first prompt word information, add the first watermark material as a watermark to the set positions in each frame of the current video segment.

[0099] In some embodiments, the watermark addition unit is further configured to determine the set position in each frame according to the position coordinate information input by the user; or determine the set position in each frame according to the selected watermark style template; or determine the recommended position after the watermark data model performs image analysis on the current video segment as the set position in each frame.

[0100] The following is an example to illustrate the video watermark addition process of the video watermark addition device provided by the embodiments of the present invention.

[0101] The video watermark addition device can be applied to devices such as user terminals or cloud platforms. The device not only stores the pre-trained watermark data model, but also stores the pre-trained material model. Of course, it also stores the watermark material library. As Figure 5 shown, the device 400 includes: an identification and extraction module 410, a matching and determination module 420, a watermark addition module 430, and a video segmentation module 440. Among them, the identification and extraction module 410 includes: a key frame determination unit 411 and a first image recognition unit 412, and the watermark addition module 430 includes: an information determination unit 431, a material generation unit 432, and a watermark addition unit 433.

[0102] After obtaining the video file, the video segmentation module 440 segments the video file according to the set shots to obtain each video segment. Among them, any video segment can be the current video segment.

[0103] After the identification and extraction module 410 extracts the corresponding audio track information from the current video segment, performs corresponding preprocessing, and performs speech recognition to extract the first keyword, and the information determination unit 431 in the watermark addition module 430 can perform speech signal feature extraction and classify the speech signal features based on machine learning algorithms to obtain the corresponding current emotion information.

[0104] The key frame determination unit 411 in the recognition and extraction module 410 can obtain the pixel features of each frame in the current video segment and obtain the pixel feature difference values between adjacent frames; based on the pixel feature difference values, determine the key frame images in the current video segment. In this way, the first image recognition unit 412 performs image preprocessing and image recognition on the key frame images to extract the second keywords. At the same time, the information determination unit 431 in the watermark addition module 430 can also extract the corresponding visual features, and based on the convolutional neural network algorithm, classify the extracted visual features to identify the corresponding current style type information.

[0105] When it is determined that the first keyword and the second keyword match, the matching determination module 420 obtains the current keyword that matches and determines the current watermark material that matches the current keyword from the saved watermark material library. Then, the material generation unit 432 in the watermark addition module 430 can generate the corresponding first watermark material based on the current watermark material, the current emotion information, and the current style type information, based on the trained material model. Thus, the watermark addition unit 433 can, through regular expressions, convert the first watermark material and the current video segment into the first prompt word information that conforms to natural language; based on the first watermark material, the current video segment, and the first prompt word information, and according to the watermark style template selected by the user, determine the set positions in each frame, and based on the watermark data model, add the first watermark material as a watermark to the set positions in each frame of the current video segment.

[0106] It can be seen that in the embodiment of the present invention, the device for video watermark addition can perform speech and image recognition on the current video segment respectively, not only obtain the corresponding first keyword and second keyword, but also determine the corresponding current emotion information and current style type information. In this way. And when the first keyword and the second keyword match, determine the corresponding current watermark material, and obtain the first watermark material that matches the current watermark material, the current emotion information, and the current style type information through the large model, so as to add it to the current video, thus automatically generating a watermark that is more in line with the video style, improving the overall visual effect of the video. Of course, it also meets the personalized needs of users, improving the viewing and professionalism of the video. And multiple watermark addition strategies can be provided, for example: large model analysis, position determination, increasing the flexibility and diversity of watermark addition.

[0107] Combined Figure 6 , the embodiment of the present invention provides a device 600 for video watermark addition, including:

[0108] A processor 1000 and a memory 1001, and may also include a communication interface 1002 and a bus 1003. Among them, the processor 1000, the communication interface 1002, and the memory 1001 can communicate with each other through the bus 1003. The communication interface 1002 can be used for information transmission. The processor 1000 can call the logical instructions in the memory 1001 to execute the method for video watermark addition in the above embodiments.

[0109] In addition, when the logical instructions in the above-mentioned memory 1001 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0110] The memory 1001, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the method in the embodiments of the present invention. The processor 1000 executes functional applications and data processing by running the program instructions / modules stored in the memory 1001, that is, implements the method for video watermark addition in the above method embodiments.

[0111] The memory 1001 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 1001 can include high-speed random access memory and can also include non-volatile memory.

[0112] The embodiments of the present invention provide a device for video watermark addition, including: a processor and a memory storing program instructions, and the processor is configured to execute the method for video watermark addition when executing the program instructions.

[0113] The embodiments of the present invention provide a storage medium storing program instructions, and the program instructions, when running, execute the method for video watermark addition as described above.

[0114] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate for implementation in the process Figure 1means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes

[0115] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions in the process Figure 1 means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes

[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, such that a series of operational steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions in the process Figure 1 means for the functions specified in one or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes

[0117] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications

Claims

1. A method for video watermark addition, characterized in that It includes: Performing speech recognition on the current video segment to obtain corresponding current speech feature information, and performing image recognition on the current video segment to obtain corresponding current image feature information; When it is determined that the current speech feature information matches the current image feature information, obtaining the current feature information that matches consistently, and determining the current watermark material that matches the current feature information; Performing corresponding watermark addition to the current video segment according to the current watermark material.

2. The method according to claim 1, wherein It also includes: Obtaining a video file, and segmenting the video file according to the set scene information or shots to obtain each video segment.

3. The method according to claim 1, characterized in that, The obtaining of the corresponding current image feature information includes: Obtaining the pixel features of each frame in the current video segment, and obtaining the pixel feature difference value between adjacent frames; Determining the key frame images in the current video segment according to the pixel feature difference value; Performing image recognition on the key frame images to obtain corresponding current image feature information.

4. The method according to any one of claims 1 to 3, characterized in that The performing of corresponding watermark addition to the current video segment according to the current watermark material includes: Determining the current scene information corresponding to the current video segment; Generating a corresponding first watermark material based on the current watermark material, the current scene information, and the material model generated by training; Adding the first watermark material as a watermark to each frame of the current video segment.

5. The method according to claim 4, characterized in that The determining of the current scene information corresponding to the current video segment includes: Performing sound preprocessing on the speech in the current video segment, and extracting the corresponding speech signal features; Classifying the speech signal features based on machine learning or deep learning algorithms to obtain the corresponding current emotion information.

6. The method according to claim 4, wherein The determining of the current scene information corresponding to the current video segment includes: Performing image preprocessing on the key frame images in the current video segment, and extracting the corresponding visual features; Classifying the visual features based on machine learning or deep learning algorithms to obtain the corresponding current style type information.

7. The method according to claim 4, wherein The adding of the first watermark material as a watermark to each frame of the current video segment includes: Converting the first watermark material and the current video segment into the first prompt word information that conforms to natural language through regular expressions; Based on the first watermark material, the current video segment, and the first prompt word information, adding the first watermark material as a watermark to the set positions of each frame of the current video segment based on the watermark data model generated by training.

8. The method according to claim 7, wherein The adding of the first watermark material as a watermark to the set positions of each frame of the current video segment further includes: Determining the set position in each frame according to the position coordinate information input by the user; or, Determining the set position in each frame according to the watermark style template selected by the user; or, Determining the recommended position obtained by the watermark data model through image analysis of the current video segment as the set position in each frame.

9. An apparatus for video watermark addition, the apparatus comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to execute the method for video watermark addition according to any one of claims 1 to 8 when executing the program instructions.

10. A storage medium stores program instructions, characterized in that, When the program instructions are running, they execute the method for video watermark addition according to any one of claims 1 to 8.