Video processing method and device, readable storage medium and program product

By identifying the target text in video subtitles and performing text semantic analysis, video processing instructions for visual elements and time parameter information are generated, and inefficiency problems caused by relying on manual experience in traditional methods are solved, and automated video processing is achieved.

CN120529031APending Publication Date: 2025-08-22XIAMEN MEITUZHIJIA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510697549.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Traditional video processing methods rely on designer experience, resulting in inefficient video processing.

Method used

By obtaining the subtitle information of the video, identifying the target text and performing text semantic analysis, determining the visual element information and time parameter information, and generating video processing instructions to automatically add visual content.

Benefits of technology

It realizes automatic processing of automatically generating visual content, improving video processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529031A_ABST
    Figure CN120529031A_ABST
Patent Text Reader

Abstract

The invention relates to a video processing method and device, a readable storage medium and a program product. The method comprises the following steps: acquiring subtitle information of a video, identifying a target text in the subtitle information, and determining a speaking time range corresponding to the target text in the video; determining visual element information suitable for the target text according to a result obtained by performing text semantic analysis processing on the target text; the visual element information is used for indicating generation of visual content; according to the speaking time range corresponding to the target text, time parameter information adopted when the visual content is presented is determined; and generating a video processing instruction based on the visual element information and the time parameter information, the video processing instruction being used for instructing to add the visual content generated according to the visual element information into the video according to the time parameter information. By adopting the method, the video processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video technology, and in particular to a video processing method, device, readable storage medium, and program product. Background Art

[0002] With the development of computer technology, video has become an important medium for information dissemination, and the demand for video processing technologies is increasing. One example of video processing technology is video editing. Video editing uses video editing software to perform nonlinear editing of videos, adding images, special effects, scenes, and other materials to the video and remixing them to generate new videos with different expressive qualities. For example, in some scenarios, when playing spoken content in a video, visual information needs to be added to enhance the video's expressiveness. Traditionally, designers have used video editing software to manually design visual information related to the spoken content and then manually add it to the video.

[0003] However, traditional methods rely on designers' processing experience and suffer from low video processing efficiency. Summary of the Invention

[0004] Based on this, the present application provides a video processing method, device, readable storage medium and program product, which can improve video processing efficiency.

[0005] In one aspect, the present application provides a video processing method, comprising:

[0006] Obtaining subtitle information of a video, identifying a target text in the subtitle information, and determining a speaking time range corresponding to the target text in the video;

[0007] Determining visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate the generation of visual content;

[0008] Determining time parameter information used when presenting the visual content according to a speaking time range corresponding to the target text;

[0009] A video processing instruction is generated based on the visual element information and the time parameter information, where the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0010] In one aspect, the present application further provides a video processing device, comprising:

[0011] An information processing module is configured to obtain subtitle information of a video, identify a target text in the subtitle information, and determine a speech time range corresponding to the target text in the video; determine visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate generation of visual content; and determine time parameter information used when presenting the visual content based on the speech time range corresponding to the target text;

[0012] An instruction generation module is used to generate a video processing instruction based on the visual element information and the time parameter information, wherein the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0013] In one aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0014] Obtaining subtitle information of a video, identifying a target text in the subtitle information, and determining a speaking time range corresponding to the target text in the video;

[0015] Determining visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate the generation of visual content;

[0016] Determining time parameter information used when presenting the visual content according to a speaking time range corresponding to the target text;

[0017] A video processing instruction is generated based on the visual element information and the time parameter information, where the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0018] In one aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0019] Obtaining subtitle information of a video, identifying a target text in the subtitle information, and determining a speaking time range corresponding to the target text in the video;

[0020] Determining visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate the generation of visual content;

[0021] Determining time parameter information used when presenting the visual content according to a speaking time range corresponding to the target text;

[0022] A video processing instruction is generated based on the visual element information and the time parameter information, where the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0023] The above-mentioned video processing method, device, readable storage medium and program product obtain subtitle information of the video, identify the target text in the subtitle information, perform text semantic analysis on the target text, and then determine the visual element information applicable to the target text based on the processing results. The visual element information is used to indicate the generation of visual content, creating conditions for automatically generating visual content applicable to the target text; then, based on the speaking time range corresponding to the target text, the time parameter information used when presenting the visual content is determined, and a video processing instruction is generated based on the visual element information and the time parameter information. Since the video processing instruction is used to indicate that the visual content generated according to the visual element information is added to the video according to the time parameter information, when the video processing instruction is executed, the visual content applicable to the target text can be automatically added to the video according to the time parameter information applicable to the target text, thereby improving the video processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A diagram showing an application environment of a video processing method in one embodiment;

[0026] Figure 2 1 is a flow chart of a video processing method according to an embodiment;

[0027] Figure 3 A schematic diagram of a video processing flow in one embodiment;

[0028] Figure 4 is a schematic diagram of an example of visual content in one embodiment;

[0029] Figure 5 is a schematic diagram of another example of visual content in one embodiment;

[0030] Figure 6This is a schematic diagram of another example of visual content in one embodiment;

[0031] Figure 7 is a structural block diagram of a video processing device in one embodiment;

[0032] Figure 8 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0034] The video processing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers. The terminal 102 can import a video through a running client, and the server running on the server 104 can obtain the subtitle information of the video, identify the target text in the subtitle information, determine the speaking time range corresponding to the target text in the video, and then determine the visual element information applicable to the target text for indicating the generation of visual content. The time parameter information used when presenting the visual content is determined based on the speaking time range corresponding to the target text, and generate video processing instructions based on the visual element information and time parameter information. The terminal 102 can be various personal computers, laptops, smartphones, tablets, or other. The server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. It is understandable that the above-mentioned video processing method can also be executed independently by the terminal 102 or by the server 104.

[0035] In an exemplary embodiment, Figure 2 As shown, a video processing method is provided, which is applied to Figure 1 The server 104 in the example is used as an example to illustrate the process, including the following steps 202 to 208.

[0036] Step 202: Acquire subtitle information of the video, identify the target text in the subtitle information, and determine the speaking time range corresponding to the target text in the video.

[0037] Subtitle information may include subtitle text and subtitle timing information. Subtitle text is the spoken content in a video recorded in text form. The spoken content in a video may be, for example, dialogue, narration, or other content. Subtitle text may include multiple sentences. Subtitle timing information may include the speaking time range corresponding to each sentence in the subtitle text in the video. Subtitle timing information may also include the speaking time range corresponding to each word in each sentence.

[0038] The speech time range for a sentence is the time span from the start of the sentence to the end of the sentence. The speech time range for a word is the time span from the start of the word to the end of the word. Here, a word can be a phrase or a single word.

[0039] The target text may be a text that meets the preset conditions. The preset conditions may be a sentence or phrase containing preset keywords. Preset keywords may include rise, fall, proportion, opinion, or others. The preset conditions may also be a sentence or phrase that represents a preset semantic type. The preset semantic type may be, for example, a quantitative statistical type or an opinion analysis type. The quantitative statistical type may involve specific numerical content, that is, it involves numerical expression. The opinion analysis type may involve opinion content, that is, it involves opinion expression. The preset conditions may also be a sentence or phrase that indicates a preset data type. Preset data types may include rise type, fall type, proportion type, opinion comparison type, opinion display type, or others.

[0040] The target text can be part of a sentence in the subtitle text, such as a phrase within a sentence. The target text can also be a sentence or multiple sentences in the subtitle text. The speech time range corresponding to the target text is the time span from the start of utterance of the target text to the end of utterance in the video.

[0041] Exemplarily, the server may receive audio data from the terminal, where the audio data is separated from the video; identify subtitle information based on the audio data, and obtain the subtitle information of the video.

[0042] In one embodiment, a server receives a video from a terminal and extracts subtitle information from the video. Where subtitle text is already annotated in the video frame, the server can extract the subtitle text in the frame and extract audio data from the video to identify subtitle time information of the subtitle text based on the audio data.

[0043] In one embodiment, the server can identify the target text from the subtitle text of the subtitle information according to preset conditions; and determine the speaking time range corresponding to the target text in the video based on at least one of the speaking time range corresponding to each sentence and the speaking time range corresponding to each word in the subtitle time information of the subtitle information.

[0044] Specifically, when the target text is a sentence in the subtitle text, the speech time range corresponding to the sentence can be used as the speech time range corresponding to the target text. For another example, when the target text is a plurality of consecutive sentences in the subtitle text, the start time of the speech time range corresponding to the first sentence in the plurality of sentences can be obtained, and the end time of the speech time range corresponding to the last sentence in the plurality of sentences can be obtained, and the time range formed by the start time to the end time can be determined as the speech time range corresponding to the target text. For another example, when the target text is part of a sentence in the subtitle text, the speech time range corresponding to the target text in the video can be determined based on the speech time ranges corresponding to the respective words in the sentence in which the target text is located.

[0045] In one embodiment, the server can identify the target text in the subtitle information and determine the speaking time range corresponding to the target text in the video using a trained large language model based on the subtitle information and pre-configured prompt information. The prompt information may include first prompt information indicating that the target text has been identified. The first prompt information may be information indicating the conditions that the target text must meet, such as information indicating that the target text must meet preset conditions. The prompt information may also include second prompt information indicating that the speaking time range corresponding to the target text in the video has been determined. The large language model may use DeepSeek (a deep search tool based on artificial intelligence technology), ChatGPT (a chatbot model released by OpenAI), or other methods. The prompt information can be input in the form of a prompt in the large language model.

[0046] Step 204 : Determine the visual element information applicable to the target text based on the result of the text semantic analysis processing on the target text; the visual element information is used to indicate the generation of visual content.

[0047] The result of performing text semantic analysis on the target text may include a data type and a data content text. The data content text may be the text of the data content to be presented in the form of graphic text in the video. The graphic text form refers to presenting the text in the form of graphics.

[0048] The data type describes the type of data content text. The data type can vary depending on the semantic type represented by the target text. For example, when the semantic type represented by the target text is a quantitative statistical type, the data type can include a statistical type and a business type. In this case, the data content text can be a quantitative numerical text. When the semantic type represented by the target text is an opinion analysis type, the data type can include both an opinion type and a business type, or it can include only an opinion type. In this case, the data content text can be a conceptual content text.

[0049] Statistical types can be used to describe the statistical dimension or direction of the data text content. For example, statistical types can include rising types, falling types, percentage types, or other types. Business types are used to describe the business involved in the target text. Business types can include market types, personnel types, exchange rate types, cost types, or other types. Opinion types can include opinion comparison types and opinion display types. Opinion comparison types can involve comparing two opinions. Opinion display types can involve displaying three or more opinions.

[0050] Visual element information may include data content text and visual element address. The visual element address may be determined based on the data type. The visual element address may be used to obtain the visual element. The data content text and the visual element may be used to generate visual content. The visual element may be, for example, a static or dynamic graphic, or other. The graphic may be a pattern or a chart. The visual content may be formed by combining the data content text and the visual element. For example, the data content text and the visual element may be combined to form visual content according to a specified relative position relationship. Specifically, for example, the visual element may be an icon representing an increase, and the data content text may be presented on top of the visual element in a preset text style. The preset text style may include a preset font, a preset text color, a preset spacing, etc. The preset font may be Songti, the preset text color may be black, and the preset spacing may be single line spacing.

[0051] Exemplarily, the server may perform text semantic analysis on the target text, obtain the data type and data content text indicated by the target text, determine the target visual element address that matches the data type, and generate visual element information applicable to the target text based on the target visual element address and data content text.

[0052] In an exemplary embodiment, the server can perform text semantic analysis on the target text through a trained large language model to obtain the data type and data content text indicated by the target text. The large language model can be instructed to identify the data type and data content text by means of a prompt. For example, the large language model can be instructed to identify the data type and data content text, and indicate that the data type can include a statistical type and a business type, or can include an opinion type and a business type, or can only include an opinion type; the specific values ​​of the statistical type, business type, and opinion type can be further indicated, for example, the statistical type can be an upward type, a downward type, and a proportion type. The large language model can be the same as the large language model used to identify the target text in the subtitle information.

[0053] Step 206 : Determine the time parameter information used when presenting the visual content according to the speaking time range corresponding to the target text.

[0054] Among them, the time parameter information is used to control the presentation rhythm of the visual content in the video. For example, the visual content can be presented in the form of animation, and the time parameter information can include the values ​​of the animation start time parameter, the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter. For another example, the visual content can include multiple sub-visual contents, each of which can be formed based on partial text and visual elements in the data content text. In this case, the time parameter information can include the values ​​of the start time parameter and the duration parameter respectively used when presenting the multiple sub-visual contents in sequence.

[0055] Exemplarily, the server may determine the type of time parameter to be used when presenting visual content based on the data type indicated by the target text; when the time parameter type is the first type, the server may determine the respective values ​​of the animation start time parameter, animation entry duration parameter, animation hold duration parameter, and animation exit duration parameter to be used when presenting visual content based on the speaking time range corresponding to the target text.

[0056] Among them, the first type is a parameter type related to animation time. For example, when the semantic type of the target text representation is a quantitative statistical type, or the semantic type of the target text representation is a viewpoint analysis type, the data type includes a viewpoint type, and the viewpoint type is a viewpoint comparison type, the time parameter type can be the first type. The value of the animation start time parameter can be the time when the visual content begins to be presented in the video. The value of the animation entry duration parameter can be the length of time required for the visual content to be completely presented from empty content when the visual content is presented in the form of animation. The value of the animation hold duration parameter can be the length of time to keep the visual content in a completely presented state. The value of the animation exit duration parameter can be the length of time required for the visual content to be completely presented from complete presentation to empty content when the visual content is presented in the form of animation.

[0057] In one embodiment, when the time parameter type is the second type, the values ​​of the start time parameter and the duration parameter to be used when sequentially presenting multiple sub-visual contents included in the visual content are determined according to the speaking time range corresponding to the target text.

[0058] Among them, when the data type includes a viewpoint type, and the viewpoint type is a viewpoint display type, the time parameter type can be the second type. The second type involves presenting viewpoints in sequence. The value of the start time parameter can be the time when the corresponding sub-visual content begins to be presented in the video. The value of the duration parameter can be the length of time the corresponding sub-visual content continues to be presented in the video, that is, the length of time required from the start to the end of presentation.

[0059] Step 208 : Generate a video processing instruction based on the visual element information and the time parameter information. The video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0060] The video processing instructions are computer instructions and can be written in JSON (JavaScript Object Notation, a lightweight data exchange format), YAML (YAML Ain't Markup Language, a human-readable data serialization format), DSL or other languages.

[0061] Exemplarily, the server can generate a video processing instruction based on the visual element information and time parameter information, and send the video processing instruction to the client running on the terminal. The client can respond to the video processing instruction, parse the video processing instruction to obtain the visual element information and time parameter information, generate visual content according to the visual element information, and add the visual content to the video according to the time parameter information.

[0062] The visual element information may include the target visual element address and data content text. The target visual element address can be used to index the visual element. The client can then obtain the visual element according to the target visual element address and generate visual content based on the visual element and data content text. The client can add visual content to the video and set the values ​​of various time parameters corresponding to the visual content in the video according to the time parameter information. Time parameters may include animation start time parameters, animation entry duration parameters, animation hold duration parameters, and animation exit duration parameters, or start time parameters and duration parameters.

[0063] In one embodiment, the visual element information may include the target visual element address and data content text. The server may generate a video processing instruction based on the data type, target visual element address, data content text, and time parameter information. In response to the video processing instruction, the client may parse the video processing instruction to obtain the data type, target visual element address, data content text, and time parameter information, obtain the visual element according to the target visual element address, and generate visual content based on the visual element and data content text. In response to the video processing trigger operation, the client may add the visual content to the video and set the values ​​of the various time parameters corresponding to the visual content in the video according to the time parameter information.

[0064] The video processing trigger operation may be a trigger operation on a video processing control displayed in the client. The client may display the parsed time parameter information, data type, and generated visual content. The data type may be displayed as a triggerable data type control. In response to the trigger operation on the data type control, the client may display various preset visual elements under the data type. In response to the selection operation of any preset visual element, the client may regenerate visual content based on the data content text and the selected visual element. In response to the video processing trigger operation, the client may add the regenerated visual content to the video and set the values ​​of various time parameters corresponding to the visual content in the video according to the time parameter information.

[0065] In the above-mentioned video processing method, the subtitle information of the video is obtained, the target text in the subtitle information is identified, and the target text is subjected to text semantic analysis processing. Then, based on the processing results, the visual element information applicable to the target text is determined, and the visual element information is used to indicate the generation of visual content, creating conditions for automatically generating visual content applicable to the target text; then, based on the speaking time range corresponding to the target text, the time parameter information used when presenting the visual content is determined, and a video processing instruction is generated based on the visual element information and the time parameter information. Since the video processing instruction is used to indicate that the visual content generated according to the visual element information is added to the video according to the time parameter information, when the video processing instruction is executed, the visual content applicable to the target text can be automatically added to the video according to the time parameter information applicable to the target text, thereby improving the video processing efficiency.

[0066] In an exemplary embodiment, step 204 may include: performing text semantic analysis on the target text to obtain the data type and data content text indicated by the target text; obtaining a target visual element address that matches the data type from a preconfigured visual element address set; generating visual element information applicable to the target text based on the target visual element address and data content text; the target visual element address is used to obtain the visual element, and the visual element and data content text are used to generate visual content.

[0067] Among them, the preconfigured visual element address set is a set formed by multiple preconfigured visual element addresses. The visual element address set can also record the data type that each visual element address matches. The target visual element address is the visual element address in the preconfigured visual element address set that matches the data type indicated by the target text. For example, the data type may include a statistical type and a business type. The statistical type may be, for example, an upward type, and the business type may be, for example, a market type. Then the target visual element address may be a visual element address that matches both the upward type and the market type. The target visual element address and the data content text are combined to form visual element information. For example, the visual element information may be: target visual element address: xxx; data content text: "30%".

[0068] In this embodiment, the data type and data content text are obtained by performing text semantic analysis on the target text, and then the target visual element address that matches the data type is obtained from the preconfigured visual element address set. In this way, the visual elements subsequently obtained based on the target visual element address are accurately adapted to the target text, and the visual element information applicable to the target text is generated based on the target visual element address and the data content text. Subsequently, the visual content adapted to the target text can be automatically and accurately generated, thereby improving the video processing efficiency.

[0069] In an exemplary embodiment, the visual content is presented in the form of animation, and the time parameter information includes the values ​​of the animation start time parameter, the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter. Step 206 may include: determining the start time in the speech time range corresponding to the target text as the value of the animation start time parameter used when presenting the visual content; and determining the values ​​of the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter used when presenting the visual content according to the duration of the speech time range corresponding to the target text and the preset weights corresponding to the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter.

[0070] The start time of the target text's speech time range is the time when the target text begins to be spoken in the video. The duration of the target text's speech time range is the length of time from the start time to the end time of the speech time range. The end time of the speech time range is the time when the target text stops being spoken in the video.

[0071] The preset weights can be used to control the values ​​of the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter, and their proportion in the duration of the speech time range corresponding to the target text. For example, the preset weights corresponding to the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter can be 0.1, 0.5, and 0.4, respectively. The duration of the speech time range corresponding to the target text can be multiplied by the preset weights corresponding to the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter, respectively, to obtain the values ​​of the animation entry duration parameter, the animation hold duration parameter, and the animation exit duration parameter.

[0072] In this embodiment, through the animation start time parameter, animation entry duration parameter, animation hold duration parameter and animation exit duration parameter, the animation presentation rhythm of the visual content can be clearly controlled later, and the start time in the speaking time range corresponding to the target text is determined as the value of the animation start time parameter used when presenting the visual content. This can ensure that the presentation of the visual content is synchronized with the time when the target text starts to be spoken in the video, and the duration of the speaking time range is allocated to the other three time parameters according to preset weights. The presentation of the visual content can be accurately controlled within the speaking time range corresponding to the target text, thereby improving the accuracy of video processing.

[0073] In an exemplary embodiment, the data type includes an opinion type, and the data content text is an opinion content text; when the opinion type is an opinion display type, the opinion content text includes multiple opinion fragments respectively representing different opinion points, and the visual content includes sub-visual contents corresponding to the multiple opinion fragments, and the multiple opinion fragments respectively form corresponding sub-visual contents with the visual elements.

[0074] Among them, an opinion fragment is a portion of text within the opinion content that represents a point of view. For example, the multiple opinion fragments can be three opinion fragments, which can be "agree," "disagree," and "neither agree nor disagree." Each opinion fragment can form sub-visual content corresponding to the opinion fragment with a visual element. For example, the visual element can be a circular pattern, and the sub-visual content can be the corresponding opinion fragment drawn into the circular pattern according to a preset text style.

[0075] In this embodiment, the data type includes an opinion type, and the data content text is an opinion content text. When the opinion type is an opinion display type, the opinion content text can be divided into multiple opinion fragments of different opinion points, and the visual content can include sub-visual content formed by each of the multiple opinion fragments and visual elements. In this way, when the visual content is subsequently presented, each opinion can be presented separately to improve the efficiency of information presentation.

[0076] In an exemplary embodiment, the time parameter information includes respective values ​​of a start time parameter and a duration parameter; step 206 may include: determining the speaking time range corresponding to each of the multiple viewpoint segments in the video within the speaking time range corresponding to the target text; determining the start time in the speaking time range corresponding to each of the multiple viewpoint segments as the value of the start time parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments; and determining the value of the duration parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments based on the duration of the speaking time range corresponding to each of the multiple viewpoint segments.

[0077] The speaking time ranges corresponding to the multiple opinion segments are all within the speaking time range corresponding to the target text. The target text may include multiple opinion segments and their corresponding detailed content. The speaking time range corresponding to each opinion segment can include the time range in the video when the opinion segment and its corresponding detailed content are expressed. For example, the target text may be "Regarding this phenomenon, there are three viewpoints: agree with his behavior because... (specific reasons, details omitted); oppose his behavior because... (specific reasons, details omitted); neither agree nor oppose because... (specific reasons, details omitted)." The multiple opinion segments can be "agree," "oppose," or "neither agree nor oppose." The detailed content corresponding to each opinion segment can be the specific reasons supporting the view expressed in the opinion segment. Taking "agree" as an example, the speaking time range corresponding to the opinion segment can be the time range in the video when "agree with his behavior because..." is said.

[0078] The value of the duration parameter used when presenting the sub-visual content corresponding to each viewpoint segment can be the duration of the speaking time range corresponding to the viewpoint segment (in this case, the sub-visual content corresponding to different viewpoint segments can be presented independently in sequence), it can be the duration of the target time range formed from the start time in the speaking time range corresponding to the viewpoint segment to the end time in the speaking time range corresponding to the target text (in this case, the sub-visual content corresponding to different viewpoint segments can start to be presented in sequence, and the sub-visual content presented first will remain displayed until the end time in the speaking time range corresponding to the target text is reached), or it can be a value that is not less than the duration of the speaking time range corresponding to the viewpoint segment and not greater than the duration of the target time range. In some scenarios, the value of the duration parameter used when presenting the sub-visual content corresponding to each viewpoint segment can also be pre-set, for example, 3000 milliseconds.

[0079] In this embodiment, the starting time in the speaking time range corresponding to each of the multiple viewpoint segments is respectively determined as the value of the starting time parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments, so that the sub-visual content can start synchronously with the speaking time of the corresponding viewpoint segment in the video, and based on the duration of the speaking time range corresponding to each of the multiple viewpoint segments, the value of the duration parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments is respectively determined. The duration of the sub-visual content can be flexibly set on a certain basis, thereby improving the flexibility of video processing.

[0080] In an exemplary embodiment, the step of obtaining subtitle information of the video in step 202 may include: receiving audio data separated from the video; performing speech recognition processing on the audio data to obtain subtitle information of the video; the subtitle information includes subtitle text and subtitle time information, and the subtitle time information includes the speaking time range corresponding to each sentence of the subtitle text in the video, and the speaking time range corresponding to each word of each sentence.

[0081] Speech recognition can be performed using pre-configured speech recognition tools, such as Whisper (OpenAI's automatic speech recognition system), Buzz (an offline speech-to-text tool built on Whisper), or others. Subtitle text can include multiple sentences, and the speech recognition tool can segment each sentence to obtain the individual words within each sentence.

[0082] In this embodiment, compared to receiving the video, separating the audio data from the video in advance and then receiving the audio data can reduce the amount of data transmitted and improve processing efficiency. Moreover, the audio data can reflect the speaking content in the video and the time information corresponding to the speaking content, thereby performing speech recognition processing on the audio data and obtaining the subtitle information of the video. The subtitle information includes subtitle text and subtitle time information. The subtitle time information includes the speaking time range corresponding to each sentence and each word, which facilitates the subsequent accurate determination of the speaking time range corresponding to the target text of various text lengths, thereby improving the efficiency and accuracy of video processing.

[0083] In an exemplary embodiment, the step of determining the speaking time range corresponding to the target text in the video in step 202 may include: when the target text is part of a sentence in the subtitle text, determining the first word ranked first and the second word ranked last among the words of the target text; determining the speaking time range corresponding to the first word and the speaking time range corresponding to the second word based on the subtitle time information; and determining the speaking time range corresponding to the target text in the video based on the speaking time range corresponding to the first word and the speaking time range corresponding to the second word.

[0084] The speaking time range corresponding to the first word and the speaking time range corresponding to the second word can be obtained from the subtitle time information. The speaking time range corresponding to the target text can be formed by the start time of the speaking time range corresponding to the first word and the end time of the speaking time range corresponding to the second word.

[0085] In this embodiment, when the target text is part of a sentence in the subtitle text, since the subtitle time information includes the speaking time range corresponding to each word in each sentence, by determining the speaking time range corresponding to the first word ranked first among the words of the target text, and the speaking time range corresponding to the second word ranked last, the speaking time range corresponding to the target text can be quickly and accurately determined.

[0086] In a specific application scenario, see Figure 3 As shown in the video processing step flow diagram, the above video processing method may specifically include the following steps.

[0087] The client running on the terminal can receive the video imported by the user, separate the audio data from the video, and send the audio data to the server running on the server.

[0088] The server running on the server can receive audio data, perform speech recognition processing on the audio data, obtain subtitle information of the video, identify the target text in the subtitle information, determine the speaking time range corresponding to the target text in the video, and determine the visual element information applicable to the target text based on the results of text semantic analysis processing on the target text; the visual element information is used to indicate the generation of visual content; based on the speaking time range corresponding to the target text, the time parameter information used when presenting the visual content is determined; video processing instructions are generated based on the visual element information and the time parameter information, and the video processing instructions are sent to the client.

[0089] The client can receive a video processing instruction, respond to the video processing instruction, parse the video processing instruction to obtain visual element information and time parameter information, generate visual content according to the visual element information, and add the visual content to the video according to the time parameter information.

[0090] For example, the target text may be a text (sentence or phrase) representing a preset semantic type, and the preset semantic type may be a quantitative statistical type or an opinion analysis type. Specifically, for example, the target text may be "market size increased by 1.27%", and the semantic type represented by the target text may be a quantitative statistical type, wherein the data type may include a statistical type and a business type, the statistical type may be an increase, the business type may be a market, the data content text may be a quantitative numerical text, specifically "1.27%", the target visual element address may match the market and the increase, the visual content may include the visual element obtained according to the target visual element address, and the graphic text formed by drawing "1.27%" according to the preset text style, and the visual content may be, for example, Figure 4 In the diagram below, a video processing instruction might read: {Statistical type: rising; Target visual element address: xxx; Data content text: "1.27%"; Animation start time parameter: 3400 milliseconds; Animation entry duration parameter: 30 milliseconds; Animation hold duration parameter: 300 milliseconds; Animation exit duration parameter: 300 milliseconds}. The target visual element address "xxx" is merely an example and does not specify a specific address. It could be a URL, file path, or other URL.

[0091] The target text may be, for example, "Enterprise costs dropped by 2 million." The semantic type represented by the target text may be a quantitative statistical type, wherein the data type may include a statistical type and a business type, the statistical type may be a decline, the business type may be a cost, the data content text may be "2 million," the target visual element address may match the cost and decline, the visual content may include visual elements obtained according to the target visual element address, and graphic text formed by drawing "2 million" according to a preset text style. The video processing instruction may be, for example, {statistical type: decline; target visual element address: xxx; data content text: "2 million"; animation start time parameter: 3400 milliseconds; animation entry duration parameter: 30 milliseconds; animation hold duration parameter: 300 milliseconds; animation exit duration parameter: 300 milliseconds}.

[0092] The target text may be, for example, "Is an old, dilapidated, and small house in the city center better or a new house in the suburbs better?" The semantic type represented by the target text may be an opinion analysis type, wherein the data type may include an opinion type, and the opinion type may be an opinion comparison type (PK). The data content text may be an opinion content text, specifically "an old, dilapidated, and small house in the city center" and "a new house in the suburbs." The target visual element address may match the opinion comparison type, and the visual content may include visual elements obtained according to the target visual element address, as well as graphic text formed by drawing "an old, dilapidated, and small house in the city center" and "a new house in the suburbs" according to a preset text style. The visual content may be, for example, Figure 5Another example diagram of visual content is shown, in which the video processing instruction may be, for example, {viewpoint type: PK; target visual element address: xxx; data content text: ["old and dilapidated house in the city center", "new house in the suburbs"]; animation start time parameter: 3400 milliseconds; animation entry duration parameter: 30 milliseconds; animation hold duration parameter: 300 milliseconds; animation exit duration parameter: 300 milliseconds}.

[0093] For example, the target text may be "There are three opinions on this phenomenon: agree with his behavior because...; oppose his behavior because...; neither agree nor disagree because..." The semantic type represented by the target text may be an opinion analysis type, wherein the data type may include an opinion type, and the opinion type may be an opinion display type (parallel). The data content text may specifically be "agree" (opinion 1), "oppose" (opinion 2), and "neither agree nor disagree" (opinion 3). The target visual element address may match the opinion display type. The visual content may include the visual element obtained according to the target visual element address, and the graphic text formed by drawing "agree", "oppose" and "neither agree nor disagree" according to a preset text style. The video processing instruction may be, for example, {opinion type: parallel; target visual element address: xxx; data content text: ["agree", "oppose", "neither agree nor disagree"]; [{start time parameter of opinion 1: 3400 milliseconds, duration parameter of opinion 1: 3000 milliseconds}, {start time parameter of opinion 2: 5400 milliseconds, duration parameter of opinion 3: 2000 milliseconds} milliseconds}, {start time parameter of view 3: 7400th millisecond, duration parameter of view 3: 2000 milliseconds}]}.

[0094] The target text may be, for example, "The market share of our company is 90%." The semantic type represented by the target text may be a quantitative statistical type, wherein the data type may include a statistical type and a business type. The statistical type may be a percentage, and the business type may be a market. The data content text may be "90%." The target visual element address may match the percentage and the market. The visual content may include visual elements obtained according to the target visual element address, and graphic text formed by drawing "90%" according to a preset text style. The visual content may be, for example, Figure 6 Another example diagram of visual content is shown, in which the video processing instruction may be, for example, {statistical type: percentage; target visual element address: xxx; data content text: "90%"; animation start time parameter: 3400 milliseconds; animation entry duration parameter: 30 milliseconds; animation hold duration parameter: 300 milliseconds; animation exit duration parameter: 300 milliseconds}.

[0095] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0096] Based on the same inventive concept, the embodiments of the present application further provide a video processing device for implementing the aforementioned video processing method. The implementation solution provided by the device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in one or more video processing device embodiments provided below can be found in the above-mentioned limitations on the video processing method and will not be repeated here.

[0097] In an exemplary embodiment, Figure 7 As shown, a video processing device 700 is provided, comprising: an information processing module 710 and an instruction generation module 720, wherein:

[0098] Information processing module 710 is used to obtain subtitle information of the video, identify the target text in the subtitle information, and determine the speaking time range corresponding to the target text in the video; determine the visual element information applicable to the target text based on the results of text semantic analysis processing on the target text; the visual element information is used to indicate the generation of visual content; and determine the time parameter information used when presenting the visual content based on the speaking time range corresponding to the target text.

[0099] The instruction generation module 720 is used to generate a video processing instruction based on the visual element information and the time parameter information, where the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

[0100] In an exemplary embodiment, the information processing module 710 is also used to perform text semantic analysis on the target text to obtain the data type and data content text indicated by the target text; obtain the target visual element address that matches the data type from a preconfigured visual element address set; generate visual element information applicable to the target text based on the target visual element address and data content text; the target visual element address is used to obtain the visual element, and the visual element and data content text are used to generate visual content.

[0101] In an exemplary embodiment, the visual content is presented in the form of animation, and the time parameter information includes the values ​​of the animation start time parameter, the animation entry duration parameter, the animation hold duration parameter and the animation exit duration parameter. The information processing module 710 is also used to determine the start time in the speech time range corresponding to the target text as the value of the animation start time parameter used when presenting the visual content; according to the duration of the speech time range corresponding to the target text, according to the preset weights corresponding to the animation entry duration parameter, the animation hold duration parameter and the animation exit duration parameter, the values ​​of the animation entry duration parameter, the animation hold duration parameter and the animation exit duration parameter used when presenting the visual content are determined.

[0102] In an exemplary embodiment, the data type includes an opinion type, and the data content text is an opinion content text; when the opinion type is an opinion display type, the opinion content text includes multiple opinion fragments respectively representing different opinion points, and the visual content includes sub-visual contents corresponding to the multiple opinion fragments, and the multiple opinion fragments respectively form corresponding sub-visual contents with the visual elements.

[0103] In an exemplary embodiment, the time parameter information includes the values ​​of the start time parameter and the duration parameter; the information processing module 710 is also used to determine the speaking time range corresponding to each of the multiple viewpoint segments in the video within the speaking time range corresponding to the target text; the start time in the speaking time range corresponding to each of the multiple viewpoint segments is determined as the value of the start time parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments; based on the duration of the speaking time range corresponding to each of the multiple viewpoint segments, the value of the duration parameter used when presenting the sub-visual content corresponding to each of the multiple viewpoint segments is determined.

[0104] In an exemplary embodiment, the information processing module 710 is also used to receive audio data separated from the video; perform speech recognition processing on the audio data to obtain subtitle information of the video; the subtitle information includes subtitle text and subtitle time information, and the subtitle time information includes the speaking time range corresponding to each sentence of the subtitle text in the video, and the speaking time range corresponding to each word of each sentence.

[0105] In an exemplary embodiment, the information processing module 710 is also used to determine the first word ranked first and the second word ranked last in the target text when the target text is part of a sentence in the subtitle text; determine the speaking time range corresponding to the first word and the speaking time range corresponding to the second word based on the subtitle time information; and determine the speaking time range corresponding to the target text in the video based on the speaking time range corresponding to the first word and the speaking time range corresponding to the second word.

[0106] Each module in the above-mentioned video processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0107] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data that needs to be stored when executing the above-mentioned video processing method. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a video processing method is implemented.

[0108] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0109] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0110] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0111] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0112] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0113] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0114] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A video processing method, characterized in that: The method comprises: Obtaining subtitle information of a video, identifying a target text in the subtitle information, and determining a speaking time range corresponding to the target text in the video; Determining visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate the generation of visual content; Determining time parameter information used when presenting the visual content according to a speaking time range corresponding to the target text; A video processing instruction is generated based on the visual element information and the time parameter information, where the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

2. The method according to claim 1, characterized in that Determining the applicable visual element information of the target text based on the result of performing text semantic analysis on the target text includes: Performing text semantic analysis on the target text to obtain the data type and data content text indicated by the target text; Acquire a target visual element address that matches the data type from a preconfigured visual element address set; According to the target visual element address and the data content text, visual element information applicable to the target text is generated; the target visual element address is used to obtain a visual element, and the visual element and the data content text are used to generate visual content.

3. The method according to claim 2, characterized in that The visual content is presented in an animation form, the time parameter information includes respective values ​​of an animation start time parameter, an animation entry duration parameter, an animation hold duration parameter, and an animation exit duration parameter, and determining the time parameter information used when presenting the visual content based on a speech time range corresponding to the target text includes: Determining the start time in the speaking time range corresponding to the target text as the value of the animation start time parameter used when presenting the visual content; According to the duration of the speaking time range corresponding to the target text, and in accordance with the preset weights corresponding to the animation entry duration parameter, the animation hold duration parameter and the animation exit duration parameter, the respective values ​​of the animation entry duration parameter, the animation hold duration parameter and the animation exit duration parameter to be used when presenting the visual content are determined.

4. The method according to claim 2, characterized in that The data type includes an opinion type, and the data content text is an opinion content text; when the opinion type is an opinion display type, the opinion content text includes multiple opinion fragments respectively representing different opinion points, and the visual content includes sub-visual content corresponding to each of the multiple opinion fragments, and the multiple opinion fragments respectively form corresponding sub-visual content with the visual elements.

5. The method according to claim 4, characterized in that The time parameter information includes respective values ​​of a start time parameter and a duration parameter; and determining the time parameter information used when presenting the visual content based on the speech time range corresponding to the target text includes: Determining, within a speaking time range corresponding to the target text, a speaking time range corresponding to each of the plurality of viewpoint segments in the video; determining the start time in the speech time range corresponding to each of the plurality of viewpoint segments as the value of the start time parameter used when presenting the sub-visual content corresponding to each of the plurality of viewpoint segments; Based on the duration of the speech time ranges corresponding to the multiple viewpoint segments, the values ​​of the duration parameters used when presenting the sub-visual contents corresponding to the multiple viewpoint segments are determined respectively.

6. The method according to any one of claims 1 to 5, characterized in that The acquisition of subtitle information of a video includes: Receiving audio data separated from the video; The audio data is subjected to speech recognition processing to obtain subtitle information of the video; the subtitle information includes subtitle text and subtitle time information, and the subtitle time information includes a speaking time range corresponding to each sentence of the subtitle text in the video, and a speaking time range corresponding to each word of each sentence.

7. The method according to claim 6, characterized in that Determining the speaking time range corresponding to the target text in the video includes: When the target text is part of a sentence in the subtitle text, determining a first word that ranks first and a second word that ranks last among the words of the target text; determining, based on the subtitle time information, a speaking time range corresponding to the first word and a speaking time range corresponding to the second word; The speaking time range corresponding to the target text in the video is determined according to the speaking time range corresponding to the first word and the speaking time range corresponding to the second word.

8. A video processing device, characterized in that: The device comprises: An information processing module is configured to obtain subtitle information of a video, identify a target text in the subtitle information, and determine a speech time range corresponding to the target text in the video; determine visual element information applicable to the target text based on a result of performing text semantic analysis on the target text; the visual element information is used to indicate generation of visual content; and determine time parameter information used when presenting the visual content based on the speech time range corresponding to the target text; An instruction generation module is used to generate a video processing instruction based on the visual element information and the time parameter information, wherein the video processing instruction is used to instruct to add the visual content generated according to the visual element information to the video according to the time parameter information.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.