Video processing method and device, electronic equipment and storage medium

By separating the video and extracting the video, and generating detailed video content description information, the problem of insufficient depth and accuracy of video understanding in the prior art is solved, and more efficient video information extraction and user operation experience are achieved.

CN120017906APending Publication Date: 2025-05-16BEIJING DONGCHEZU TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510034175.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video understanding algorithms mainly focus on the mining of video picture information, resulting in limited depth and accuracy of video understanding, which makes it difficult to meet complex and changeable application needs.

Method used

By separating the original video with audio and video, audio information and images are extracted; audio information is voice recognition to obtain audio text; extracting text in the image to obtain subtitles; based on the image, audio text and subtitles, the content description information of the video is generated, including content overview and description of each content unit.

Benefits of technology

It improves the depth and accuracy of video understanding, and can more effectively extract key information in the video to meet users' needs for rapid positioning and operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017906A_ABST
    Figure CN120017906A_ABST
Patent Text Reader

Abstract

The invention relates to a video processing method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the audio-picture separation of an original video, and obtaining the audio information of the original video and an image of the original video; the original video comprises a plurality of content units; performing voice recognition on the audio information of the original video to obtain an audio text of the original video; extracting characters in the image of the original video to obtain subtitles of the original video; obtaining content description information of the original video based on the image, the audio text and the subtitles of the original video; the content description information of the original video comprises original video content summary description information and content description information of each content unit of the original video. The depth and accuracy of video understanding can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video processing method, device, electronic device and storage medium. Background Art

[0002] In today's era of rapid development of digital information, the importance of video understanding technology is self-evident. It is widely used in many fields such as intelligent security, media content management, video editing and creation, video recommendation systems, etc., providing strong support for people to extract key information from massive video data and realize intelligent interaction.

[0003] However, existing video understanding algorithms often focus on mining video image information. This processing mode has little reference information for understanding video content, which seriously limits the depth and accuracy of video understanding and is difficult to meet the current complex and changing application needs. Summary of the invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a video processing method, device, electronic device and storage medium.

[0005] In a first aspect, the present disclosure provides a video processing method, comprising:

[0006] Performing audio and video separation processing on the original video to obtain audio information of the original video and an image of the original video; the original video includes multiple content units;

[0007] Performing speech recognition on the audio information of the original video to obtain an audio text of the original video;

[0008] Extracting text from the image of the original video to obtain subtitles of the original video;

[0009] Based on the image, audio text and subtitle of the original video, content description information of the original video is obtained; the content description information of the original video includes summary description information of the original video content and content description information of each content unit of the original video.

[0010] In a second aspect, the present disclosure further provides a video processing device, including:

[0011] A separation module, used to perform audio and video separation processing on the original video to obtain the audio information of the original video and the image of the original video; the original video includes multiple content units;

[0012] A recognition module, used for performing speech recognition on the audio information of the original video to obtain the audio text of the original video;

[0013] An extraction module, used for extracting text from the image of the original video to obtain subtitles of the original video;

[0014] The output module is used to obtain content description information of the original video based on the image, audio text and subtitles of the original video; the content description information of the original video includes the original video content overview description information and the content description information of each content unit of the original video.

[0015] In a third aspect, the present disclosure further provides an electronic device, the electronic device comprising:

[0016] one or more processors;

[0017] A storage device for storing one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described above.

[0019] In a fourth aspect, the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the video processing method as described above when executed by a processor.

[0020] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0021] The technical solution provided by the embodiment of the present disclosure is to obtain the audio information and the image of the original video by setting the original video to perform audio and picture separation processing; the original video includes multiple content units; the audio information of the original video is voice recognized to obtain the audio text of the original video; the text in the image of the original video is extracted to obtain the subtitle of the original video; based on the image, audio text and subtitle of the original video, the content description information of the original video is obtained; the content description information of the original video includes the summary description information of the original video content and the content description information of each content unit of the original video. Its essence is that when understanding the original video, it is not based solely on the mining of the video picture information, but on the basis of the video picture information, combined with the audio text and subtitles, to obtain the video content understanding result, so that the setting can improve the depth and accuracy of video understanding. In addition, in practical applications, users often need to locate specific content units in the video in order to carry out subsequent viewing, editing, recommendation and other operations. If only the overall content of the video is summarized, it is not conducive to the user to quickly locate the content unit that the user wants to view. The present application sets the content description information of the original video to include the summary description information of the original video content and the content description information of each content unit of the original video, which helps to meet the needs of users to quickly locate a certain content unit of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0024] Figure 1 A flowchart of a video processing method provided by an embodiment of the present disclosure;

[0025] Figure 2 A flowchart of a method for implementing S140 provided in an embodiment of the present disclosure;

[0026] Figure 3 is a structural schematic diagram of a video processing device in an embodiment of the present disclosure;

[0027] Figure 4 It is a structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0029] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0030] Figure 1 This is a flowchart of a video processing method provided in an embodiment of the present disclosure. This embodiment is applicable to the case where video processing is performed in a client. The method can be executed by a video processing device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a terminal, specifically including but not limited to a smart phone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one machine, a smart home device, etc. Alternatively, this embodiment is applicable to the case where video processing is performed in a server. The method can be executed by a video processing device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a server.

[0031] like Figure 1 As shown, the method may specifically include:

[0032] S110, performing audio and video separation processing on the original video to obtain audio information and an image of the original video; the original video includes multiple content units.

[0033] The original video is a video that needs to be understood and the original video content description information needs to be obtained. In practice, the original video may be a video specified by a user, which may specifically be a video uploaded by the user or downloaded from the network.

[0034] The original video includes multiple content units. A content unit is an independent component of a video and is obtained by analyzing the original video. A content unit can be a specific event shown in the original video or an independent scene formed by time and space differences.

[0035] For example, taking the video in the automotive field as an example, if a video introduces a certain model of vehicle, it first introduces the appearance of the vehicle, then the interior of the vehicle, and finally the configuration of the vehicle. The appearance of the vehicle, the interior of the vehicle, and the configuration of the vehicle can be respectively regarded as a content unit, that is, the vehicle includes 3 content units. If another video introduces the driving experience of the vehicle by driving the vehicle in different road conditions, it first introduces the driving experience of the vehicle on urban roads, then introduces the driving experience of the vehicle on highways, then introduces the driving experience of the vehicle on mountain roads, and finally introduces the driving experience of the vehicle on rural roads. Urban roads, highways, mountain roads, and rural roads can be respectively regarded as a content unit, that is, the vehicle includes 4 content units.

[0036] Performing audio and video separation processing on the original video to obtain the audio information and the image of the original video means extracting the audio information and the image from the original video. The extracted image of the original video is specifically represented by a series of continuous image frames.

[0037] S120: Perform speech recognition on the audio information of the original video to obtain the audio text of the original video.

[0038] In practice, a speech recognition model can be used to recognize the audio information of the original video. The purpose of recognition is to convert the audio information in the original video into text.

[0039] S130: Extract text from the image of the original video to obtain subtitles of the original video.

[0040] In practice, the OCR (Optical Character Recognition) technology can be used to extract text from the image frame to obtain the subtitles of the original video.

[0041] S140, obtaining content description information of the original video based on the image, audio text and subtitle of the original video; the content description information of the original video includes summary description information of the original video content and content description information of each content unit of the original video.

[0042] The original video content summary description information may be, for example, information that generally introduces the original video content.

[0043] The content description information of a content unit may be, for example, description information that only introduces the content of the content unit. Exemplarily, a certain original video includes three content units, the first content unit introduces the appearance of the vehicle, the second content unit introduces the interior of the vehicle, and the third content unit introduces the power of the vehicle. The description information of the original video content overview is: This video mainly introduces the changes in the appearance, interior and power of the new model XX vehicle. Content description information of content unit one: The video shows the appearance of the XX model vehicle, including the overall design of the front and body. Content description information of content unit two: The video shows the interior details of the XX model vehicle, especially the imitation carbon fiber panel under the central control screen and the newly added small screen. Content description information of content unit three: The video points out that the power of the XX model vehicle is basically the same as the current model, but a new motor may be used in the high-performance version.

[0044] The above technical solution is to obtain the audio information and the image of the original video by setting the original video to perform audio and picture separation processing; the original video includes multiple content units; the audio information of the original video is voice recognized to obtain the audio text of the original video; the text in the image of the original video is extracted to obtain the subtitle of the original video; based on the image, audio text and subtitle of the original video, the content description information of the original video is obtained; the content description information of the original video includes the summary description information of the original video content and the content description information of each content unit of the original video. Its essence is that when understanding the original video, it is not based solely on the mining of the video picture information, but on the basis of the video picture information, combined with the audio text and subtitles, the video content understanding result is obtained. Such a setting can improve the depth and accuracy of video understanding. In addition, in practical applications, users often need to locate specific content units in the video in order to carry out subsequent viewing, editing, recommendation and other operations. If only the overall content of the video is summarized, it is not conducive to the user to quickly locate the content unit that the user wants to view. The present application sets the content description information of the original video to include the summary description information of the original video content and the content description information of each content unit of the original video, which helps to meet the needs of users to quickly locate a certain content unit of the video.

[0045] Based on the above technical solution, optionally, there are multiple implementation methods of S140, which are not limited in this application. Exemplarily, the implementation method of this step may include:

[0046] S141. According to the content units, the images, audio texts and subtitles of the original video are segmented to obtain segmented image frame sets, audio text segments and subtitle segments; each segmented image frame set, audio text segment and subtitle segment corresponds to a content unit.

[0047] The segmented image frame set includes a series of continuous video frames. All video frames in the segmented image frame set are image frames used to reflect the corresponding content units.

[0048] Optionally, the images, audio text and subtitles of the original video are segmented according to the content units to obtain a segmented set of image frames, audio text segments and subtitle segments, including: based on the original video, determining the time period corresponding to each content unit in the original video; based on the time period corresponding to each content unit in the original video, segmenting the images, audio text and subtitles of the original video to obtain a segmented set of image frames, audio text segments and subtitle segments.

[0049] The time period corresponding to each content unit in the original video refers to determining the time period occupied by each content unit on the time axis of the original video.

[0050] There are many implementation methods for "determining the time period corresponding to each content unit in the original video based on the original video", and the present application does not limit this. Exemplarily, in one embodiment, the implementation method of this step includes: if the audio information of the original video includes voice, based on the content of the voice, determining the time period corresponding to each content unit in the original video. In some scenarios, in the original video, when different content units are converted, iconic words or sentences are used. Based on these words or sentences used to indicate the conversion of content units, the time period corresponding to each content unit in the original video can be accurately determined. In other scenarios, in the original video, the use of words or sentences in different content units is quite different. For example, in the content unit showing the appearance of the vehicle, words such as "taillight" and "streamlined" are used, but the probability of mentioning "central control screen" and "seat ventilation" is low. In the content unit showing the interior of the vehicle, words such as "central control screen" and "seat ventilation" appear more frequently, but the probability of mentioning "taillight" and "streamlined" is low. Based on the difference in the use of words or sentences, the time period corresponding to each content unit in the original video can be accurately determined.

[0051] In another embodiment, the implementation method of this step may include: determining the time period corresponding to each content unit in the original video based on the similarity of the image frame pictures of the original video. In some scenarios, the image frames of different content units are quite different, and these differences can be manifested as differences in image features, for example. For example, in a video introducing the driving experience of a vehicle in different road conditions, in the content unit of an urban road, the image frame pictures usually include various man-made building facilities, traffic lights for traffic management, and pedestrians, etc.; while in the content unit of a highway, the image frame pictures usually include signs indicating direction and distance, isolation guardrails for separating lanes and ensuring safety, etc. It can be set that when there are image frames including the same or similar elements in the original video, the similarity between them is high. These same or similar elements are typical features of the same content unit. On the contrary, when the image frames contain different elements, the similarity between them is low, which indicates that these pictures are likely to belong to different content units. By analyzing and comparing the elements of the image frames, the high and low changes in the similarity of the image frames can be obtained, and then the time period corresponding to each content unit in the original video can be determined based on the law of similarity change. Specifically, when image frames with high similarity appear continuously, they can be divided into a content unit, and the start and end periods of the content unit can be determined based on the timestamps of the first and last appearances of these images; and when the elements of the image frame change significantly, that is, the similarity changes from high to low, it means entering the next content unit.

[0052] In practice, the time period corresponding to each content unit in the original video can be determined based only on the language content in the original video audio information, or based only on the similarity of the image frames of the original video. It can also be set to determine the time period corresponding to each content unit in the original video based on both the language content in the original video audio information and the similarity of the image frames of the original video.

[0053] S142, aligning the segmented image frame sets, audio text segments, and subtitle segments to obtain multiple data groups; the data groups correspond to the content units one by one.

[0054] The essence of alignment is to determine which segmented image frame sets, audio text segments and subtitle segments correspond to the same time period, and to group the segmented image frame sets, audio text segments and subtitle segments corresponding to the same time period together as a data group.

[0055] Exemplarily, it is assumed that by analyzing an original video, it is found that the original video includes 3 content units. On the timeline of the original video, the time periods corresponding to these three content units are the 0-t1 period, the t1-t2 period, and the t2-t3 period. The image, audio text, and subtitles of the original video are segmented, and 3 image frame sets, 3 audio text segments, and 3 subtitle segments are obtained after segmentation. Different image frame sets correspond to different time periods, different audio text segments correspond to different time periods, and different subtitle segments correspond to different time periods. The image frame sets, audio text segments, and subtitle segments corresponding to the 0-t1 period are collected together as data group 1; the image frame sets, audio text segments, and subtitle segments corresponding to the t1-t2 period are collected together as data group 2; the image frame sets, audio text segments, and subtitle segments corresponding to the t2-t3 period are collected together as data group 3.

[0056] S143: Determine content description information of each content unit of the original video and summary description information of the original video content based on the data in each data group.

[0057] There are many methods for implementing this step, which are not limited in this application. Exemplarily, the method for implementing this step may include: inputting the data in each data group into a content description extraction model to obtain content description information of each content unit of the original video and the original video content summary description information. The content description extraction model may be, for example, a large speech model.

[0058] By setting the image, audio text and subtitle of the original video according to the content unit, the segmented image frame set, audio text segment and subtitle segment are obtained; each segmented image frame set, audio text segment and subtitle segment corresponds to a content unit respectively; the segmented image frame set, audio text segment and subtitle segment are aligned to obtain multiple data groups; the data groups correspond to the content units one by one; based on the data in each data group, the content description information of each content unit of the original video and the original video content overview description information are determined. The essence of this setting is to split and sort out the content of the original video before determining the content description information of the original video. In this way, the accuracy of the subsequent original video content description information can be guaranteed, the difficulty of determining the information can be reduced, and the rapid positioning of the content unit can be realized later.

[0059] On the basis of the above technical solution, optionally, after S141, the method may further include: performing frame extraction processing on the segmented image frame set according to a preset frame extraction rule to obtain a frame extracted image frame set; S142 may include: aligning the frame extracted image frame set, audio text segments and subtitle segments to obtain multiple data groups.

[0060] The preset frame extraction rule is a pre-set frame extraction rule, and the present application does not limit the specific rules of the preset frame extraction rule. Exemplarily, the preset frame extraction rule can be a frame extraction rule of uniform frame extraction type, such as extracting an image frame every N frames (N is a positive integer greater than or equal to 1), or extracting an image frame every preset time length. In addition, the preset frame extraction rule can also be a frame extraction rule of the frame extraction type based on actual needs. For example, according to the importance of the content introduced in different time periods in the content unit, the frame extraction time interval, frame number interval or total number of frames corresponding to the time period is determined. For example, if a content unit involves the introduction of a certain vehicle interior, covering the introduction of the center console, instrument panel and interior lighting system, and the difference between the new version and the old version of the vehicle is mainly reflected in the design of the center console and instrument panel, that is, the introduction of the center console and instrument panel is more important than the introduction of the interior lighting system. Set the time period for introducing the center console and instrument panel to extract an image frame every 3 frames, and the time period for introducing the interior lighting system to extract an image frame every 15 frames.

[0061] The purpose of "performing frame extraction processing on the segmented image frame set according to preset frame extraction rules" is to reduce the number of image frames used in the subsequent determination of the content description information of the original video, reduce the computational complexity of determining the content description information of the original video, and increase the rate of determining the content description information of the original video.

[0062] Furthermore, before S143, the method may also include: splicing the image frames in the image frame set in the data group to obtain a target image; and replacing the image frame set in the data group with the target image.

[0063] Taking the example of inputting a data set into a content description extraction model to obtain the content description information of the original video, if multiple image frames are directly input into the content description extraction model, the amount of data input into the model will be large, increasing the model operation complexity and the demand for hardware resources. However, splicing multiple image frames into a target image is equivalent to integrating and compressing the data that needs to be input into the model, which can reduce the amount of data input into the model, making the model processing more efficient and reducing the demand for hardware resources.

[0064] Further, it can be set that “joining the image frames in the image frame set in the data group to obtain the target image” includes: joining the image frames in the image frame set after frame extraction to obtain the target image;

[0065] In the above technical solution, there are many specific training methods for the content description extraction model, and this application does not limit this. For example, after obtaining the sample video, the method described in the previous article can be used to construct sample data. The sample video is a video similar to the original video, and the main difference is that the two are used in different stages. The use stage of the original video is the content description extraction model reasoning stage, while the use stage of the sample video is the content description extraction model training stage. Specifically, the sample video is subjected to audio and video separation processing to obtain the audio information of the sample video and the image of the sample video; the sample video includes multiple content units; the audio information of the sample video is subjected to speech recognition to obtain the audio text of the sample video; the text in the image of the sample video is extracted to obtain the subtitle of the sample video; the image, audio text and subtitle of the sample video are segmented according to the content unit to obtain the segmented image frame set, audio text segment and subtitle segment; each segmented image frame set, audio text segment and subtitle segment respectively corresponds to a content unit; the segmented image frame set, audio text segment and subtitle segment are aligned to obtain multiple sample data groups; the sample data group corresponds to the content unit one by one; the content description information of each content unit of the sample video is determined based on the sample data group. Based on the content description information of each content unit of the sample video, the sample video content overview description information is obtained. The content description information of each content unit of the sample video and the sample video content overview description information are integrated to obtain the content description information of the sample video. The sample data group and the content description information are further used as sample data. The content description extraction model is trained using the sample data.

[0066] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0067] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0068] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0069] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0070] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0071] Figure 3 FIG. 1 is a schematic diagram of the structure of a video processing device in an embodiment of the present disclosure. The video processing device provided in the embodiment of the present disclosure may be configured in a client or in a server. Figure 3 , the video processing device specifically comprises:

[0072] The separation module 310 is used to perform audio and video separation processing on the original video to obtain the audio information of the original video and the image of the original video; the original video includes multiple content units;

[0073] The recognition module 320 is used to perform speech recognition on the audio information of the original video to obtain the audio text of the original video;

[0074] An extraction module 330, configured to extract text from the image of the original video to obtain subtitles of the original video;

[0075] The output module 340 is used to obtain content description information of the original video based on the image, audio text and subtitles of the original video; the content description information of the original video includes the original video content summary description information and the content description information of each content unit of the original video.

[0076] Furthermore, the output module 340 is used to:

[0077] According to the content unit, the image, audio text and subtitle of the original video are segmented to obtain a segmented image frame set, an audio text segment and a subtitle segment; each segmented image frame set, audio text segment and subtitle segment corresponds to a content unit;

[0078] Aligning the segmented image frame sets, audio text segments, and subtitle segments to obtain a plurality of data groups; the data groups correspond one-to-one to the content units;

[0079] Based on the data in each of the data groups, content description information of each content unit of the original video and summary description information of the original video content are determined.

[0080] Furthermore, the output module 340 is used to:

[0081] Based on the original video, determining a time period corresponding to each content unit in the original video;

[0082] Based on the time periods corresponding to the various content units in the original video, the images, audio texts and subtitles of the original video are segmented to obtain segmented image frame sets, audio text segments and subtitle segments.

[0083] Furthermore, the output module 340 is used to:

[0084] If the audio information of the original video includes voice, the time period corresponding to each content unit in the original video is determined based on the content of the voice.

[0085] Furthermore, the output module 340 is used to:

[0086] Based on the similarity of the image frames of the original video, the time period corresponding to each content unit in the original video is determined.

[0087] Further, the output module 340 is used to: segment the image, audio text and subtitle of the original video according to the content unit to obtain the segmented image frame set, audio text segment and subtitle segment, and then perform frame extraction processing on the segmented image frame set according to a preset frame extraction rule to obtain a frame extracted image frame set;

[0088] The extracted image frame sets, audio text segments and subtitle segments are aligned to obtain multiple data groups.

[0089] Further, the output module 340 is used to: before determining the content description information of each content unit of the original video and the summary description information of the original video content based on the data in each of the data groups, splice the image frames in the image frame set in the data group to obtain the target image;

[0090] A set of image frames in the data set is replaced with the target image.

[0091] Furthermore, the output module 340 is used to:

[0092] The data in each of the data groups is input into a content description extraction model to obtain content description information of each content unit of the original video and summary description information of the original video content.

[0093] The video processing device provided in the embodiment of the present disclosure can execute the steps executed by the client or server in the video processing method provided in the embodiment of the method of the present disclosure, and has execution steps and beneficial effects, which will not be repeated here.

[0094] Figure 4 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 4 , which shows a schematic diagram of the structure of an electronic device 1000 suitable for implementing the embodiment of the present disclosure. The electronic device 1000 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0095] like Figure 4As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 to a random access memory (RAM) 1003 to implement the video processing method of the embodiment described in the present disclosure. In the RAM 1003, various programs and information required for the operation of the electronic device 1000 are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0096] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange information. Although Figure 4 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0097] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart, thereby implementing the video processing method as described above. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 1009, or installed from a storage device 1008, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0098] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include an information signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated information signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0099] In some embodiments, the client and the server may communicate using any known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital information communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any known or future developed network.

[0100] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0101] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0102] Performing audio and video separation processing on the original video to obtain audio information of the original video and an image of the original video; the original video includes multiple content units;

[0103] Performing speech recognition on the audio information of the original video to obtain an audio text of the original video;

[0104] Extracting text from the image of the original video to obtain subtitles of the original video;

[0105] Based on the image, audio text and subtitle of the original video, content description information of the original video is obtained; the content description information of the original video includes summary description information of the original video content and content description information of each content unit of the original video.

[0106] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0107] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0108] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0109] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.

[0110] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0111] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0112] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:

[0113] one or more processors;

[0114] A memory for storing one or more programs;

[0115] When the one or more programs are executed by the one or more processors, the one or more processors implement any video processing method provided in the present disclosure.

[0116] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the program implements any of the video processing methods provided by the present disclosure.

[0117] The embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the video processing method described above is implemented.

[0118] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0119] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video processing method, characterized in that: include: Performing audio and video separation processing on the original video to obtain audio information of the original video and an image of the original video; the original video includes multiple content units; Performing speech recognition on the audio information of the original video to obtain an audio text of the original video; Extracting text from the image of the original video to obtain subtitles of the original video; Based on the image, audio text and subtitle of the original video, content description information of the original video is obtained; the content description information of the original video includes summary description information of the original video content and content description information of each content unit of the original video.

2. The method according to claim 1, characterized in that The obtaining content description information of the original video based on the image, audio text and subtitle of the original video includes: According to the content unit, the image, audio text and subtitle of the original video are segmented to obtain a segmented image frame set, an audio text segment and a subtitle segment; each segmented image frame set, audio text segment and subtitle segment corresponds to a content unit; Aligning the segmented image frame sets, audio text segments, and subtitle segments to obtain a plurality of data groups; the data groups correspond one-to-one to the content units; Based on the data in each of the data groups, content description information of each content unit of the original video and summary description information of the original video content are determined.

3. The method according to claim 2, characterized in that The method of segmenting the image, audio text and subtitle of the original video according to the content unit to obtain a segmented image frame set, audio text segment and subtitle segment comprises: Based on the original video, determining a time period corresponding to each content unit in the original video; Based on the time periods corresponding to the various content units in the original video, the images, audio texts and subtitles of the original video are segmented to obtain segmented image frame sets, audio text segments and subtitle segments.

4. The method according to claim 3, characterized in that The determining, based on the original video, a time period corresponding to each content unit in the original video includes: If the audio information of the original video includes voice, the time period corresponding to each content unit in the original video is determined based on the content of the voice.

5. The method according to claim 3, characterized in that: The determining, based on the original video, a time period corresponding to each content unit in the original video includes: Based on the similarity of the image frames of the original video, the time period corresponding to each content unit in the original video is determined.

6. The method according to claim 2, characterized in that After segmenting the image, audio text and subtitle of the original video according to the content unit to obtain the segmented image frame set, audio text segment and subtitle segment, the method further includes: Performing frame extraction processing on the segmented image frame set according to a preset frame extraction rule to obtain a frame extracted image frame set; The segmented image frame set, audio text segment and subtitle segment are aligned to obtain multiple data groups, including: The extracted image frame sets, audio text segments and subtitle segments are aligned to obtain multiple data groups.

7. The method according to claim 2, characterized in that Before determining the content description information of each content unit of the original video and the summary description information of the original video content based on the data in each data group, the method includes: splicing the image frames in the image frame set in the data group to obtain a target image; A set of image frames in the data set is replaced with the target image.

8. The method according to claim 2, characterized in that: The step of determining the content description information of each content unit of the original video and the summary description information of the original video content based on the data in each data group includes: The data in each of the data groups is input into a content description extraction model to obtain content description information of each content unit of the original video and summary description information of the original video content.

9. A video processing device, characterized in that: include: A separation module is used to separate the audio and video of the original video to obtain the audio information of the original video and the image of the original video; The original video includes multiple content units; A recognition module, used for performing speech recognition on the audio information of the original video to obtain the audio text of the original video; An extraction module, used for extracting text from the image of the original video to obtain subtitles of the original video; The output module is used to obtain content description information of the original video based on the image, audio text and subtitles of the original video; the content description information of the original video includes the original video content overview description information and the content description information of each content unit of the original video.

10. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Video file processing method and server

    CN110493661A

  • Video generation method and device, electronic equipment and storage medium

    CN115955585A

  • Video processing method, electronic equipment and storage medium

    CN118175383A

  • Multi-modal video question answering method and device and computer equipment

    CN118194230A

  • Audio storage method and device, electronic equipment and computer readable medium

    CN118400563A