A background system for video playback of video clips

By using a backend system for video playback and editing, speech audio data is acquired and processed in real time to generate associated images and subtitle videos, solving the problem of lack of visual coordination in speeches and improving the overall presentation effect.

CN114218413BActive Publication Date: 2026-03-24IIE STAR (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing speeches often lack necessary visual accompaniment in large settings, making it difficult for audiences to effectively grasp the content and thus affecting the overall impact of the speech.

Method used

A backend system for video playback and editing is used to acquire speech audio data from amplification devices in real time, extract keywords and match them with a pre-built library of related images for speech topics, generate related images and subtitle videos, and play them synchronously to enhance the combination of visual and auditory experiences.

Benefits of technology

It enables the timely provision of visual cues during the speech, enhancing the audience's speaking experience and improving the audiovisual effects of the speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114218413B_ABST
    Figure CN114218413B_ABST
Patent Text Reader

Abstract

The application is suitable for the field of computers and provides a background system for playing video clips of videos, which is applied to a speech scene and comprises a conversion and extraction module, a matching module and a classification module.The conversion and extraction module is used to acquire speech voice data transmitted in real time by a sound amplification device and extract speech content keywords in the speech voice data.The matching module is used to match the extracted speech content keywords in a pre-constructed speech theme association database to obtain preselected speech theme association pictures.The classification module is used to associate the preselected speech theme association pictures with a pre-established picture classification model to obtain speech association pictures.A synthesis and cutting module is used to synthesize the obtained speech association pictures to obtain speech association videos.The application has the beneficial effect that it can link the vision and hearing of the audience together, provide a better audio-visual scene for speeches and ensure the effect of speeches.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computers, and particularly relates to a background system for playing a video clip in a video. BACKGROUND

[0002] The video clip is nonlinear editing of a video source by using software, and the essence is to re-mix pictures, background music, special effects, scenes and other materials added with the video, cut and merge the video source, generate a new video with different performance through secondary coding.

[0003] Speech, simply speaking, is a language communication activity in which one's own opinions and views on a specific issue are clearly and completely expressed in a public place by using voice as the main means and auxiliary body language.

[0004] The existing speech lacks necessary picture cooperation when being performed, especially in some large speech occasions, some audience have poor content acquisition ability due to the reasons such as loudspeakers or distance, and the speech cannot achieve ideal speech effect. SUMMARY

[0005] The purpose of the embodiment of the application is to provide a background system for playing a video clip in a video, which aims to solve the problems proposed in the background.

[0006] The embodiment of the application is implemented as follows: a background system for playing a video clip in a video is applied to a speech scene, comprising:

[0007] A conversion extraction module is configured to acquire speech voice data transmitted in real time in a loudspeaker device, and extract speech content keywords in the speech voice data;

[0008] A matching module is configured to match the extracted speech content keywords in a pre-constructed speech theme association picture library, and obtain pre-selected speech theme association pictures;

[0009] A classification module is configured to associate the pre-selected speech theme association pictures by using a pre-established picture classification model, and obtain speech association pictures;

[0010] A synthesis and cutting module is configured to synthesize the obtained speech association pictures, obtain speech association videos, and cut the speech association videos exceeding a preset playing time length, so that the time length of the cut videos does not exceed the preset playing time length;

[0011] An adding and playing module is configured to add speech text information as display subtitles of the cut speech association videos, and synchronously play the speech association videos with subtitles when a speaker is speaking.

[0012] As a further scheme of the present application, the conversion and extraction module comprises:

[0013] a recording unit configured to record the speech voice data of the speaker in real time;

[0014] a conversion unit configured to convert the speech voice data into speech text information;

[0015] an extraction unit configured to extract the speech content keywords from the speech text information.

[0016] As a further scheme of the present application, the step of constructing the speech theme associated image library in advance comprises:

[0017] obtaining the speech theme information before the speech starts, the speech theme information comprising the speech theme and the speech name;

[0018] decomposing the speech theme information according to the keyword attribute to obtain a plurality of speech theme keywords;

[0019] performing picture search on the speech theme, the speech name and each decomposed speech theme keyword as a composite search item in sequence, and constructing the speech theme associated image library by using the pictures in the search results that meet the corresponding first preset threshold as the construction elements;

[0020] respectively classifying and marking the speech theme keywords and the pictures obtained by the composite search using the speech theme keywords in the speech theme associated image library.

[0021] As a further scheme of the present application, the matching module comprises:

[0022] a keyword matching unit configured to classify and match the speech content keywords and the speech theme keywords, and divide the speech content keywords into a plurality of classified and marked keywords and non-classified and marked keywords according to a preset classified and matched condition;

[0023] a classified and selected unit configured to select the corresponding classified and marked pictures in the speech theme associated image library according to the classified and marked keywords to obtain the preselected speech theme associated pictures.

[0024] As a further scheme of the present application, the classification module comprises:

[0025] a classification unit configured to classify the speech content keywords according to the keyword attribute, count the number of classification items in the classified results of the speech content keywords, and obtain the speech content keyword classification data set;

[0026] an arrangement unit configured to arrange the speech content keyword classification data set in descending order to obtain the speech content descending keyword classification data set;

[0027] The determining unit is configured to determine the speech theme associated picture according to the keywords in descending order of the speech content.

[0028] As a further scheme of the present application, the determining unit comprises:

[0029] The data set determining subunit is configured to set the data set in which the keyword classification data of the speech content in descending order is greater than the second preset threshold as a first data set, and set the data set in which the keyword classification data of the speech content in descending order is not greater than the second preset threshold as a second data set.

[0030] The picture determining subunit is configured to associate the speech content keyword corresponding to the first data set with the preselected speech theme associated picture, and display the speech content keyword corresponding to the second data set as a text picture, wherein the speech associated picture comprises the preselected speech theme associated picture corresponding to the first data set and the text picture corresponding to the second data set.

[0031] As a further scheme of the present application, the determining unit further comprises a text picture rendering unit, which is configured to render the speech content keyword displayed by the text picture.

[0032] As a further scheme of the present application, the synthesis cutting module comprises:

[0033] The synthesizing unit is configured to synthesize the preselected speech theme associated picture in the speech associated picture in real time according to the speech time progress.

[0034] The dynamic display unit is configured to insert the text picture in the speech associated picture as a floating frame into the video synthesized by the synthesizing unit, wherein the speech associated video is a video obtained by inserting the floating frame into the video synthesized by the synthesizing unit.

[0035] The cutting unit is configured to cut the speech associated video that exceeds the preset playing time length.

[0036] The embodiment of the present application provides a background system for video playing video clips, which is applied to a speech scene, can obtain speech voice data transmitted in real time in a sound amplification device through a conversion and extraction module, and extracts speech content keywords in the speech voice data, so that the speech voice data can be obtained in advance before sound amplification, and a preselected speech theme associated picture is obtained by matching the extracted speech content keywords in a preconstructed speech theme associated picture library through a matching module, the preselected speech theme associated picture is associated through a picture classification model preestablished by a classification module to obtain a speech associated picture, so that the working efficiency of a synthesis and cutting module can be ensured, the preconstructed speech theme associated picture library and the picture classification model preestablished can simplify the step of directly obtaining a picture at the beginning of the speech, so that the timeliness of synthesizing and matching a video for the speech is ensured, the adding and playing module can add speech text information as display subtitles of the cut speech associated video, and the speech associated video with subtitles is played synchronously when the speaker speaks, the background system is applied to the speech scene, the visual and auditory of the audience can be connected, a better speech audio-visual scene is provided, and the speech effect is ensured. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 It is a main structure schematic diagram of a background system for video playing video clips.

[0038] Figure 2 It is a structure schematic diagram of a conversion and extraction module in a background system for video playing video clips.

[0039] Figure 3 It is a step flow chart of preconstructing a speech theme associated picture library in a background system for video playing video clips.

[0040] Figure 4 It is a structure schematic diagram of a matching module in a background system for video playing video clips.

[0041] Figure 5 It is a structure schematic diagram of a classification module in a background system for video playing video clips.

[0042] Figure 6 It is a structure schematic diagram of a determination unit in a background system for video playing video clips. DETAILED DESCRIPTION

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0044] The specific implementation of the present application is described in detail below in combination with specific embodiments.

[0045] The present application provides a background system for playing a video clip of a video, which solves the problem that the existing speech lacks necessary picture cooperation when being performed, especially in some large speech occasions, some audiences have poor ability to obtain the content of the speech due to the loudspeaker or distance, and the speech is difficult to achieve the ideal speech effect.

[0046] As shown in Figure 1 FIG. 1 is a schematic diagram of the main structure of a background system for playing a video clip of a video according to an embodiment of the present application, the background system for playing a video clip of a video is applied to a speech scene, and includes:

[0047] The conversion extraction module 100 is configured to obtain speech voice data transmitted in real time in a sound amplification device and extract speech content keywords from the speech voice data.

[0048] The matching module 200 is configured to match the extracted speech content keywords in a pre-constructed speech theme association picture library to obtain pre-selected speech theme association pictures.

[0049] The classification module 300 is configured to associate the pre-selected speech theme association pictures with a pre-established picture classification model to obtain speech association pictures.

[0050] The synthesis and cutting module 400 is configured to synthesize the obtained speech association pictures to obtain speech association videos, and cut the speech association videos that exceed a preset playing time length, so that the time length of the cut videos does not exceed the preset playing time length.

[0051] The adding and playing module 500 is configured to add speech text information as display subtitles for the cut speech association videos, and synchronously play the speech association videos with the added subtitles when the speaker is speaking.

[0052] The embodiment of the application is applied to a speech scene, speech voice data transmitted in real time in a sound amplification device can be acquired by the conversion and extraction module 100, and speech content keywords in the speech voice data are extracted, so that the speech voice data can be obtained in advance before sound amplification, and the preselected speech theme associated picture is obtained by matching the extracted speech content keywords in the preconstructed speech theme associated picture library according to the matching module 200, the speech associated picture is obtained by associating the preselected speech theme associated picture with the preestablished picture classification model according to the classification module 300, the working efficiency of the synthesis and cutting module 400 is ensured, the step of directly obtaining a picture at the beginning of the speech is simplified by the preconstructed speech theme associated picture library and the preestablished picture classification model, the timeliness of synthesizing and matching a video for the speech is ensured, the added playing module 500 can add speech text information as display subtitles of the cut speech associated video, the speech associated video with subtitles is played synchronously when the speaker is speaking, the visual and auditory of the audience are connected by applying the background system in the speech scene, a better speech audio-visual scene is provided, and the speech effect is ensured.

[0053] As shown in Figure 2 FIG. 1 is a structural schematic diagram of the conversion and extraction module 100 in the background system for video playing video clips provided by the embodiment of the application, the conversion and extraction module 100 comprises:

[0054] The recording unit 101 is used for recording speech voice data of a speaker in real time.

[0055] The conversion unit 102 is used for converting the speech voice data into speech text information.

[0056] The extraction unit 103 is used for extracting speech content keywords in the speech text information.

[0057] The recording unit 101, the conversion unit 102 and the extraction unit 103 can acquire speech voice data transmitted in real time in a sound amplification device, and extract speech content keywords in the speech voice data, and it should be noted that the recording and conversion are performed before sound amplification of a speech sound amplification device or synchronously with sound amplification, which provides time convenience for subsequent video synthesis and editing.

[0058] As shown in Figure 3 FIG. 2 is a step flowchart of preconstructing a speech theme associated picture library provided by the embodiment of the application, the steps of preconstructing the speech theme associated picture library comprise:

[0059] S10: Obtain speech theme information before a speech starts, the speech theme information comprises a speech theme and a speech name.

[0060] S11: Decompose the speech topic information according to keyword attributes to obtain several speech topic keywords;

[0061] S12: Use the speech topic, speech name and each decomposed speech topic keyword as composite search terms to perform image search, and use the images in the search results that meet the corresponding first preset threshold as building elements to construct a speech topic associated image library.

[0062] S13: Categorize and label the speech topic keywords and the images obtained by combining the speech topic keywords in the speech topic-related image library.

[0063] In application, this invention obtains speech topic information, including the speech topic and speech name, before the speech begins. The speech topic and speech name are then decomposed according to keyword attributes to obtain several speech topic keywords. The speech topic, speech name, and each decomposed speech topic keyword are sequentially used as composite search terms for image searches. Images that meet the corresponding first preset threshold in the search results are used as building elements to construct a speech topic-related image library. The speech topic keywords and the images obtained through composite searches using these keywords are categorized and labeled in the speech topic-related image library. For example, in a speech themed "Patriotism among Middle School Students," a speaker's speech title might be "I love you, China, always ready to contribute to the motherland." The images are then categorized and labeled according to keyword attributes. The search criteria can be categorized into "personal terms," ​​"purpose terms," ​​"theme terms," ​​and "making contributions," specifically: "middle school students," "making contributions," and "patriotism." Using these three keywords as a composite search term for image searches, the limited number of keywords allows for a wider range of images to be retrieved. The first preset threshold can be image pixels, image size, image format, etc., without specific limitations, and can be set at the beginning of the speech. A pre-built image library related to the speech theme greatly simplifies the online search process at the start of the speech, enabling timely matching of video and speech content. The pre-built image library facilitates the extraction and generation of related images for subsequent speeches, effectively reducing the time-consuming process of directly searching images using keywords from the speech content.

[0064] like Figure 4 The diagram shown is a structural schematic of a matching module 200 in a background system for video playback and video editing. As another preferred embodiment of the present invention, the matching module 200 includes:

[0065] Keyword matching unit 201 is used to classify and match speech content keywords and speech theme keywords, and divide speech content keywords into several classification mark keywords and non-classification mark keywords according to preset classification matching conditions;

[0066] The categorization selection unit 202 is configured to select a corresponding categorized marked picture from the speech theme associated picture library according to the categorization marked keyword, and obtain a preselected speech theme associated picture.

[0067] In the application, according to the aforementioned example of the speech, during the speech, the speech content keywords are decomposed, and the decomposed speech content keywords are categorized. Since the speech content is an interpretation and expansion of the speech theme and name, the decomposed speech content keywords can be classified in the speech theme keywords, that is, the “speech content” is classified into the “speech theme”. Thus, the categorization and matching of the speech content keywords and the speech theme keywords are completed, and the speech content keywords that are not classified are classified into the non-categorized marked keywords. Therefore, the preselected speech theme associated picture can be obtained by the categorization selection unit 202, and the preselected speech picture is extracted.

[0068] As shown in FIG. 3, it is a structural schematic diagram of a classification module 300 in a background system for video playing video clips, and the classification module 300 comprises: Figure 5 The statistical unit 301 is configured to count the number of attribute decomposition items of the speech theme information keywords, and obtain a speech theme keyword classification data set;

[0069] The classification unit 302 is configured to classify the speech content keywords according to the keyword attributes, count the number of classification items in the speech content keyword attribute classification result, and obtain a speech content keyword classification data set;

[0070] The arrangement unit 303 is configured to arrange the speech content keyword classification data set in descending order, and obtain a speech content descending keyword classification data set;

[0071] The determination unit 304 is configured to determine the speech theme associated picture according to the speech content descending keywords.

[0072] As shown in FIG. 4, it is a structural schematic diagram of the determination unit 304 in the background system for video playing video clips provided by the embodiment of the application, and the determination unit 304 comprises:

[0073] Figure 6 The data set determination subunit 3041 is configured to set the data set with a value greater than a second preset threshold in the speech content descending keyword classification data set as a first data set, and set the data set with a value not greater than the second preset threshold in the speech content descending keyword classification data set as a second data set;

[0074] The data set determination subunit 3041 is configured to set the data set with a value greater than a second preset threshold in the speech content descending keyword classification data set as a first data set, and set the data set with a value not greater than the second preset threshold in the speech content descending keyword classification data set as a second data set;

[0075] ​The picture determining sub-unit 3042 is configured to associate the speech content keywords corresponding to the first data set with the preselected speech theme associated pictures, and display the speech content keywords corresponding to the second data set as text pictures.

[0076] In another case of the embodiment, the determining unit 304 further comprises a text picture rendering unit configured to render the speech content keywords displayed as text pictures, so that the text pictures are more suitable for being combined with the preselected speech theme associated pictures corresponding to the first data set after the combination and adaptation, and the difference between the frames of the video is smaller.

[0077] In application of the embodiment, the picture classification model pre-established by the classification module 300 is used to associate the preselected speech theme associated pictures, so as to obtain the speech associated pictures. The second preset threshold is a condition for distinguishing the first data set and the second data set. For example, the second preset threshold is 5. Only when the classification result is greater than 5 after the classification according to the keyword attribute, the data set is determined as the first data set. When the classification result is less than or equal to 5, the data set is determined as the second data set. Through the first data set and the second data set, the preselected speech theme associated pictures corresponding to the first data set and the text pictures corresponding to the second data set can be obtained by the picture determining unit 304.

[0078] The synthesis cutting module 400 comprises:

[0079] The synthesis unit 401 is configured to synthesize the preselected speech theme associated pictures in the speech associated pictures in real time according to the speech time progress.

[0080] The dynamic display unit 402 is configured to insert the text pictures in the speech associated pictures as floating frames into the video synthesized by the synthesis unit 401, wherein the speech associated video is the video obtained by inserting the floating frames into the video synthesized by the synthesis unit 401.

[0081] The cutting unit 403 is configured to cut the speech associated video that exceeds the preset playing time.

[0082] In application of the embodiment, the synthesis unit 401 is configured to synthesize the preselected speech theme associated pictures in the speech associated pictures in real time according to the speech time progress, so that the preselected speech theme associated pictures are associated with each other (because the speech content is logically associated). The dynamic display unit 402 is configured to insert the text pictures as floating frames into the video synthesized by the synthesis unit 401, so that the final synthesized video is more complete.

[0083] The background system for playing video clips of a video provided by the above-mentioned embodiments of the present application can obtain the speech data transmitted in real time by the amplification device through the conversion and extraction module 100, and extract the speech content keywords in the speech data, so that the speech data can be obtained before amplification, and the preselected speech theme associated pictures can be obtained by matching the extracted speech content keywords with the preconstructed speech theme associated picture library through the matching module 200, and the speech associated pictures can be obtained by associating the preselected speech theme associated pictures with the preestablished picture classification model of the classification module 300, so that the working efficiency of the synthesis and cutting module 400 can be ensured, the step of directly obtaining the pictures at the beginning of the speech can be simplified through the preconstructed speech theme associated picture library and the preestablished picture classification model, so that the timeliness of synthesizing and matching the video for the speech can be ensured, the display subtitles of the cut speech associated video can be added by adding the playing module 500, the speech associated video with subtitles added can be played synchronously when the speaker speaks, and the visual and auditory of the audience can be connected by applying the background system in the speech scene, a better speech audio-visual scene can be provided, and the speech effect can be ensured.

[0084] In order to enable the above-mentioned method and system to run smoothly, the system can include more or fewer components than those described above, or combine certain components, or different components, for example, can include input and output devices, network access devices, buses, processors and memories, etc.

[0085] It should be understood that although each step in the flowchart of each embodiment of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in each embodiment can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or sub-steps or stages of other steps.

[0086] Each technical feature of the above-mentioned embodiments can be combined arbitrarily, and in order to make the description concise, not all possible combinations of each technical feature in the above-mentioned embodiments are described, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.

[0087] The above embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but should not be understood as a limitation on the patent scope of the present application. It should be noted that, for ordinary skilled persons in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

[0088] The above merely describes the preferred embodiments of the present application and should not be used to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A backend system for video playback and video editing, characterized in that, Applied to speech scenarios, including: The conversion and extraction module is used to acquire speech audio data transmitted in real time from the loudspeaker and extract keywords from the speech audio data. The matching module is used to match the extracted speech content keywords with a pre-built speech topic association image library to obtain pre-selected speech topic association images. The classification module is used to associate pre-built image classification models with images related to pre-selected speech topics to obtain speech-related images; The synthesis and cutting module is used to synthesize the obtained speech-related images to obtain speech-related videos, and to cut speech-related videos that exceed the preset playback duration so that the length of the cut video does not exceed the preset playback duration. Add a playback module to add text information to the speech and display subtitles for the cut speech-related video, and play the speech-related video with added subtitles synchronously while the speaker is speaking; The classification module includes: a determination unit, used to determine images associated with the speech theme based on descending keywords in the speech content; the determination unit includes: Dataset sub-units are defined as follows: the dataset with values ​​greater than the second preset threshold in the speech content descending keyword classification dataset is set as the first dataset, and the dataset with values ​​not greater than the second preset threshold in the speech content descending keyword classification dataset is set as the second dataset. The image determination subunit is used to associate the speech content keywords corresponding to the first dataset with images associated with pre-selected speech topics, and to display the speech content keywords corresponding to the second dataset as text images. The associated speech images include images associated with the pre-selected speech topics in the first dataset and text images corresponding to the second dataset. The determination unit also includes a text image rendering unit, which renders the speech content keywords displayed in the text images.

2. The backend system for video playback and video editing according to claim 1, characterized in that, The conversion and extraction module includes: Recording unit, used to record the speaker's speech audio data in real time; The conversion unit is used to convert speech audio data into speech text information. The extraction unit is used to extract keywords from the speech text.

3. The backend system for video playback and video editing according to claim 1, characterized in that, The steps of pre-building the speech topic association library include: Before the speech begins, obtain the speech topic information, which includes the speech topic and speech title; The speech topic information is broken down according to keyword attributes to obtain several speech topic keywords; The speech topic, speech name, and each decomposed speech topic keyword are used as composite search terms for image search. The speech topic-related image library is constructed using images in the search results that meet the corresponding first preset threshold as building elements. The keywords of the speech topic and the images obtained by combining these keywords in a speech topic-related image library will be categorized and labeled separately.

4. The backend system for video playback and video editing according to claim 3, characterized in that, The matching module includes: The keyword matching unit is used to categorize and match the keywords of the speech content and the keywords of the speech theme. The keywords of the speech content are divided into several categorized keywords and non-categorized keywords according to the preset categorization and matching conditions. The categorization selection unit is used to select corresponding categorized images from the speech topic-related image library based on categorization tag keywords, thus obtaining pre-selected speech topic-related images.

5. The backend system for video playback and video editing according to claim 4, characterized in that, The classification module includes: The classification unit is used to classify the keywords of the speech content according to the keyword attributes, count the number of classification items in the classification results of the keyword attributes of the speech content, and obtain the speech content keyword classification dataset. The sorting unit is used to sort the speech content keyword classification dataset in descending order, resulting in a speech content descending keyword classification dataset.

6. The backend system for video playback and video editing according to claim 1, characterized in that, The synthesis and shearing module includes: The compositing unit is used to synthesize images related to pre-selected speech topics in real time according to the speech timeline. A dynamic display unit is used to insert text images from speech-related images as floating frames into the video synthesized by the synthesis unit, wherein the speech-related video is a video obtained by inserting floating frames into the video synthesized by the synthesis unit. The cutting unit is used to cut out speech-related videos that exceed a preset playback duration.

Citation Information

Patent Citations

  • Lecture background matching method and apparatus

    CN101314081A

  • Video generation method and device, electronic equipment and storage medium

    CN110347869A