Video processing method and device, client device, server device and medium

By extracting and transmitting video frames in segments by the client device and processing them in batches by the server device, the problem of time-consuming video processing in the existing technology is solved, and more efficient video processing and a better user experience are achieved.

CN120751198APending Publication Date: 2025-10-03BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410346014.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the prior art, a client device extracts and uploads all audio and video frames of an original video to a server device for processing, resulting in a time-consuming video processing process and a poor user experience.

Method used

The client device extracts and transmits part of the video frames to the server device for image processing at target time intervals. The server device processes the video frames in batches, combines the parallel processing of audio and video, reduces the peak memory usage, and realizes the pipeline operation of extraction, transmission and processing at the same time.

Benefits of technology

It reduces the time consumption of video processing, improves overall efficiency and usability, enhances user experience, and reduces CPU utilization and memory peak of client devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751198A_ABST
    Figure CN120751198A_ABST
Patent Text Reader

Abstract

The invention relates to a video processing method and device, client equipment, server equipment, a computer readable storage medium and a computer program product, and relates to the technical field of video processing. The video processing method applied to the client device comprises the following steps: acquiring an original video; extracting a part of video frames of the original video from the video frames which are not extracted in the original video every target duration, and taking the extracted video frames as target video frames; in response to the completion of the operation of extracting the partial video frames of the original video, transmitting the target video frame to server-side equipment; receiving a video processing result obtained by the server-side device based on image processing of the target video frame; and displaying the video processing result. According to the invention, the memory peak value occupied by the client can be reduced, the time consumption is reduced, the overall efficiency and availability of video processing are improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video processing method and apparatus, a client device and a server device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Chapters can help users filter videos and improve completion rates during mid-length video consumption. Actively creating chapters on the client side is expensive. Intelligent chapter generation can significantly reduce the cost of creating chapters.

[0003] In related technologies, a client device extracts audio and video frames from an original video. After extracting all audio and video frames, the client device uploads all audio and video frames together to a server device, which processes all audio and video frames together, generates video chapters based on the processing results, and sends the video chapters to the client device. Summary of the Invention

[0004] According to a first aspect of some embodiments of the present disclosure, a video processing method is provided, which is applied to a client device, and includes: acquiring an original video; extracting, at target time intervals, partial video frames of the original video from video frames that have not been extracted from the original video, as target video frames; in response to completion of the operation of extracting partial video frames of the original video, transmitting the target video frames to a server device; receiving a video processing result obtained by the server device based on image processing of the target video frame; and displaying the video processing result.

[0005] According to a second aspect of some embodiments of the present disclosure, a video processing method is provided, which is applied to a server device, comprising: obtaining a target video frame transmitted by a client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from video frames that have not been extracted from the original video at target time intervals, and an operation of the client device transmitting the target video frame to the server device is performed in response to the client device completing an operation of extracting the partial video frame of the original video; in response to obtaining the target video frame, performing image processing on the target video frame; determining a video processing result based on a result of the image processing, and sending the video processing result to the client device.

[0006] According to a third aspect of some embodiments of the present disclosure, a client device is provided, including: an acquisition module configured to acquire an original video; an extraction module configured to extract, at target time intervals, partial video frames of the original video from video frames that have not been extracted from the original video, as target video frames; a transmission module configured to transmit the target video frames to a server device in response to completion of the operation of extracting partial video frames of the original video; a receiving module configured to receive a video processing result obtained by the server device based on image processing of the target video frame; and a display module configured to display the video processing result.

[0007] According to a fourth aspect of some embodiments of the present disclosure, a server device is provided, comprising: an acquisition module configured to acquire a target video frame transmitted by a client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from video frames that have not been extracted from the original video at target time intervals, and the operation of the client device transmitting the target video frame to the server device is performed in response to the client device completing the operation of extracting the partial video frame of the original video; an image processing module configured to perform image processing on the target video frame in response to acquiring the target video frame; a determination module configured to determine a video processing result based on the result of the image processing; and a sending module configured to send the video processing result to the client device.

[0008] According to a fifth aspect of some embodiments of the present disclosure, a video processing device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute any one of the aforementioned video processing methods based on instructions stored in the memory.

[0009] According to a sixth aspect of some embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, any one of the aforementioned video processing methods is implemented.

[0010] According to a seventh aspect of some embodiments of the present disclosure, a computer program product is provided. When the computer program product is run on a computer, the computer is enabled to implement any one of the aforementioned video processing methods.

[0011] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0012] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The drawings described herein are used to provide a further understanding of the present disclosure. Each of the drawings, together with the following detailed description, is included in this specification and forms a part of the specification to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation of the present disclosure. In the drawings:

[0014] Figure 1 A flowchart showing a video processing method according to some embodiments of the present disclosure is shown;

[0015] Figure 2 A flowchart showing a video processing method according to some other embodiments of the present disclosure is shown;

[0016] Figure 3 An interactive schematic diagram illustrating a video processing method according to some embodiments of the present disclosure is shown.

[0017] Figure 4 A schematic diagram showing a flow chart of a video processing method according to some embodiments of the present disclosure;

[0018] Figure 5 A block diagram illustrating a client device according to some embodiments of the present disclosure;

[0019] Figure 6 A block diagram showing a server device according to some embodiments of the present disclosure is shown;

[0020] Figure 7 A block diagram showing a video processing apparatus according to some embodiments of the present disclosure;

[0021] Figure 8 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0022] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION

[0023] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. However, it is obvious that the embodiments described are only some embodiments of the present disclosure, rather than all embodiments. The following description of the embodiments is actually only illustrative and is in no way intended to limit the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein.

[0024] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values ​​of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.

[0025] As used in this disclosure, the term "include" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Furthermore, the term "comprise" and its variations are intended to be open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to." Therefore, "include" and "include" are synonymous. The term "based on" means "based, at least in part, on."

[0026] Reference throughout this specification to "one embodiment," "some embodiments," or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Furthermore, the appearances of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may.

[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.

[0028] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0030] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0031] In the related art, the client device extracts all audio and video frames of the original video and transmits them together to the server device, which processes all audio and video frames together. The entire video processing process is time-consuming and the user experience is poor.

[0032] In response to the above technical problems, the present disclosure proposes a video processing method and apparatus, a client device and a server device, which can reduce the time consumption of the entire video processing process, improve the overall efficiency and usability of video processing, and enhance the user experience.

[0033] The following will be combined Figure 1 The video processing method of the present disclosure is described in detail from the perspective of a client device.

[0034] Figure 1 A flowchart of a video processing method according to some embodiments of the present disclosure is shown.

[0035] like Figure 1 As shown, the video processing method includes: step S110, obtaining an original video; step S120, extracting a portion of the original video frames from the unextracted video frames in the original video at target time intervals as target video frames; step S130, in response to completing the operation of extracting the portion of the original video frames, transmitting the target video frames to a server device; step S140, receiving a video processing result obtained by the server device based on image processing of the target video frames; and step S150, displaying the video processing result. The video processing method is applied to a client device.

[0036] In the above embodiment, the client device extracts part of the video frames of the original video at regular intervals, and in response to completing the operation of extracting part of the video frames of the original video, transmits the target video frame to the server device, so that the server device obtains the video processing result based on the image processing of the target video frame, and the client device receives and displays the video processing result, realizing the pipeline operation of the client device extracting and transmitting the video frame in segments and the server device processing the video frame in segments, so that the client device can extract the video frame once at regular intervals, reduce the CPU (Central Processing Unit) utilization of the client device, reduce the peak memory occupied by the client device, and also enable the server device to process all the video frames of the original video in batches, reduce the time consumption of the entire video processing process, improve the overall efficiency and availability of video processing, and enhance the user experience. The present disclosure can also reduce the CPU utilization of the client device by extracting and transmitting the video frame in segments, thereby reducing the overheating of the client device.

[0037] The video processing method in some embodiments of the present disclosure will be described in detail below.

[0038] In step S110, the original video is obtained. In some embodiments, the client device receives the original video uploaded by the user. For example, the original video can be a video created by the user.

[0039] In step S120 , at every target duration, some video frames of the original video are extracted from the video frames that have not been extracted from the original video as target video frames.

[0040] In some embodiments, the target duration is greater than or equal to the duration used to extract the portion of the original video frame, and the difference between the target duration and the duration used to extract the portion of the original video frame is less than a difference threshold. The target duration being greater than or equal to the duration used to extract the portion of the original video frame can further reduce peak memory usage and improve video processing efficiency. Furthermore, by controlling the difference threshold, it can fully utilize client device memory and improve resource utilization.

[0041] In some embodiments, the target duration may also be less than the duration used to extract a portion of the original video frame. In this case, the peak memory usage may be reduced to a certain extent, thereby improving video processing efficiency.

[0042] In some embodiments, the target duration can be determined based on the CPU processing capabilities of different client devices. The stronger the CPU processing capabilities of the client device, the shorter the target duration can be set. That is, client devices with different CPU processing capabilities can set different target durations. In this way, personalized customization can be achieved for different client devices.

[0043] In step S130 , in response to the completion of the operation of extracting a portion of the video frames of the original video, the target video frame is transmitted to the server device.

[0044] In some embodiments, in response to completing the operation of extracting a portion of the video frame from the original video, transmitting the target video frame to the server device includes: transmitting the target video frame to an intermediate device in response to completing the operation of extracting the portion of the video frame from the original video; and sending a second notification message to the server device in response to completing the transmission of the target video frame to the intermediate device. The second notification message is used to notify the server device to obtain the target video frame from the intermediate device. The client device sending the target video frame to the intermediate device can reduce storage pressure on the server device.

[0045] In some embodiments, the intermediate device storing the target video frame is a cloud device. The target video frame extracted from the original video is stored in the cloud with a larger storage space, which can reduce storage costs.

[0046] In step S140, a video processing result obtained by the server device based on image processing of the target video frame is received. In some embodiments, the image processing includes optical character recognition (OCR). The server device may be deployed with an OCR model, and then use the OCR model to perform optical character recognition on the target video frame.

[0047] In step S150 , the video processing result is displayed.

[0048] In some embodiments, the client device can also perform audio separation on the original video to obtain the target audio. In response to completing the audio separation operation on the original video, the target audio is transmitted to the server device, and the server device receives and displays the video processing results obtained based on the image processing of the target video frame and the audio processing of the target audio. The transmission of the target audio is parallel to the transmission of the target video frame. The audio separation operation is performed in parallel with the operation of extracting part of the video frame of the original video. Since the audio file is very small, parallel processing of the target audio and target video frames can further improve the overall efficiency of video processing and reduce time consumption without significantly affecting the peak memory usage of the client device.

[0049] In some embodiments, in response to completing the audio separation operation on the original video, transmitting the target audio to the server device includes: transmitting the target audio to an intermediate device in response to completing the audio separation operation on the original video; and sending a first notification message to the server device in response to completing the transmission of the target audio to the intermediate device. The first notification message is used to notify the server device to obtain the target audio from the intermediate device. The client device sending the target audio to the intermediate device can reduce storage pressure on the server device.

[0050] In some embodiments, the intermediate device storing the target audio is a cloud device. Storing the target audio in the cloud with larger storage space can reduce storage costs.

[0051] In some embodiments, audio processing may include automatic speech recognition (ASR). An ASR model may be deployed on the server side to perform automatic speech recognition on the target audio using the ASR model.

[0052] In some embodiments, taking optical character recognition as image processing and automatic speech recognition as audio processing as an example, the video processing result includes the server-side device's results based on optical character recognition and automatic speech recognition, and the chapter information of the original video generated using a generation model. For example, the generation model is an intelligent chapter model. The video processing method disclosed herein can be applied to video chapter generation scenarios, which can reduce the peak memory occupied by client devices and server devices, reduce the time consumption of video chapter generation, reduce user waiting time, and reduce the situation where users give up chapter creation due to lack of patience to wait for a long time for the video chapter generation process, thereby improving user experience and the usability of video chapter generation.

[0053] In some embodiments, the results of optical character recognition and automatic speech recognition are used by a server-side device to generate initial chapter information for the original video using a generative model. This initial chapter information is then desensitized using a security model to obtain the chapter information for the original video. For example, the security model can be a Transformer-based machine learning model. Desensitization involves removing information from the initial chapter information that does not meet preset criteria.

[0054] In some embodiments, taking the video chapter creation scenario as an example, the client device also receives and displays chapter information of the original video from the server device. For example, the client device also publishes the original video with chapter information in response to the user's publishing operation.

[0055] The following will be combined Figure 2 The video processing method of the present disclosure is described in detail from the perspective of the server device.

[0056] Figure 2 A flowchart of a video processing method according to some other embodiments of the present disclosure is shown.

[0057] like Figure 2As shown, the video processing method includes: step S210, obtaining a target video frame transmitted by a client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from the video frames not extracted from the original video at a target time interval, and the operation of the client device transmitting the target video frame to the server device is performed in response to the client device completing the operation of extracting the partial video frame of the original video; step S220, in response to obtaining the target video frame, performing image processing on the target video frame; step S230, determining a video processing result based on the result of the image processing; and step S240, sending the video processing result to the client device. The video processing method is applied to the server device.

[0058] In the above embodiment, the client device extracts partial video frames of the original video at target time intervals and transmits the extracted partial video frames to the server device. The server device can receive the original video from the client device multiple times, receive partial video frames of the original video each time, and perform image processing on these partial video frames, so that the server device processes all video frames of the original video in batches, reduces the peak memory usage of the server device, cooperates with the client device to realize the pipeline operation of extraction, transmission and processing at the same time, reduces the time consumption of the entire video processing process, improves the overall efficiency and availability of video processing, and enhances user experience.

[0059] In step S210, a target video frame transmitted by the client device is obtained, wherein the target video frame is a partial video frame of the original video extracted by the client device from video frames that have not been extracted from the original video at target time intervals. The operation of transmitting the target video frame from the client device to the server device is performed in response to the client device completing the operation of extracting the partial video frame of the original video.

[0060] In some embodiments, obtaining the target video frame transmitted by the client device includes: obtaining the target video frame from the intermediate device in response to receiving a second notification message from the client device, wherein the target video frame is transmitted to the intermediate device by the client device in response to completing the operation of extracting the partial video frame of the original video. Storing the target video frame on the intermediate device can reduce storage pressure on the server device.

[0061] In some embodiments, the intermediate device storing the target video frame is a cloud device. The target video frame extracted from the original video is stored in the cloud with a larger storage space, which can reduce storage costs.

[0062] In some embodiments, the target duration is greater than or equal to the duration used to extract the portion of the original video frame, and the difference between the target duration and the duration used to extract the portion of the original video frame is less than a difference threshold. The target duration being greater than or equal to the duration used to extract the portion of the original video frame can further reduce peak memory usage and improve video processing efficiency. Furthermore, by controlling the difference threshold, the client device memory can be fully utilized, improving resource utilization.

[0063] In some embodiments, the target duration may also be less than the duration used to extract a portion of the original video frame. In this case, the peak memory usage may be reduced to a certain extent, thereby improving video processing efficiency.

[0064] In some embodiments, the target duration can be determined based on the CPU processing capabilities of different client devices. The stronger the CPU processing capabilities of the client device, the shorter the target duration can be set. That is, client devices with different CPU processing capabilities can set different target durations. In this way, personalized customization can be achieved for different client devices.

[0065] In step S220 , in response to acquiring the target video frame, image processing is performed on the target video frame.

[0066] In some embodiments, the image processing includes optical character recognition. An OCR model may be deployed on the server side, so that the server side device uses the OCR model to perform optical character recognition on the target video frame.

[0067] In step S230 , a video processing result is determined based on the result of the image processing.

[0068] In step S240, the video processing result is sent to the client device, and the client device displays the video processing result.

[0069] In some embodiments, the server device also obtains target audio transmitted by the client device; in response to obtaining the target audio, performs audio processing on the target audio; and determines the video processing result based on the image processing result and the audio processing result. The target audio is obtained by the client device performing audio separation on the original video. The operation of the client device transmitting the target audio to the server device is performed in response to the client completing the audio separation operation on the original video. The transmission of the target audio is performed in parallel with the transmission of the target video frame. The audio separation operation is performed in parallel with the operation of extracting part of the video frame of the original video. The audio processing of the target audio is performed in parallel with the image processing of the target video frame. Since the audio file is very small, parallel processing of the target audio and target video frames can further improve the overall efficiency of video processing and reduce time consumption without significantly affecting the peak memory usage of the client device and the server device.

[0070] In some embodiments, obtaining the target audio transmitted by the client device includes: obtaining the target audio from the intermediate device in response to receiving a first notification message from the client device, wherein the target audio is transmitted to the intermediate device by the client device in response to completing the audio separation operation on the original video. Storing the target audio on the intermediate device can reduce storage pressure on the server device.

[0071] In some embodiments, the intermediate device storing the target audio is a cloud device. Storing the target audio in the cloud with larger storage space can reduce storage costs.

[0072] In some embodiments, the audio processing may include automatic speech recognition. An ASR model may be deployed on the server side, so that the server side device uses the ASR model to perform automatic speech recognition on the target audio.

[0073] In some embodiments, taking optical character recognition as image processing and automatic speech recognition as audio processing as an example, based on the result of the image processing, determining the video processing result includes generating chapter information of the original video as the video processing result using a generative model based on the result of the optical character recognition and the result of the automatic speech recognition. For example, the generative model is an intelligent chapter model. The video processing method disclosed herein can be applied to a video chapter generation scenario, which can reduce the peak memory occupied by client devices and server devices, reduce the time consumption of video chapter generation, reduce user waiting time, reduce the situation where users give up chapter creation due to lack of patience to wait for a long time for the video chapter generation process, improve user experience, and improve the usability of video chapter generation.

[0074] In some embodiments, the server-side device can use a generative model based on the results of optical character recognition and automatic speech recognition to generate initial chapter information for the original video, and then use the security model to desensitize the initial chapter information to obtain the chapter information for the original video. For example, the security model can be a Transformer-based machine learning model. In some embodiments, the security model can be deployed on the server-side device.

[0075] Figure 3 An interactive schematic diagram illustrating a video processing method according to some embodiments of the present disclosure is shown.

[0076] like Figure 3 As shown, the video processing method includes steps S300 to S303.

[0077] In step S300, the client device obtains an original video. For example, the client device receives an original video uploaded by a user. The original video can be a video created by the user.

[0078] In step S301 , the client device extracts some video frames of the original video from the video frames that have not been extracted from the original video at every target duration as target video frames.

[0079] In some embodiments, the target duration is greater than or equal to the duration used to extract the portion of the original video frame, and the difference between the target duration and the duration used to extract the portion of the original video frame is less than a difference threshold. The target duration being greater than or equal to the duration used to extract the portion of the original video frame can further reduce peak memory usage and improve video processing efficiency. Furthermore, by controlling the difference threshold, the client device memory can be fully utilized, improving resource utilization.

[0080] In some embodiments, the target duration may also be less than the duration used to extract a portion of the original video frame. In this case, the peak memory usage may be reduced to a certain extent, thereby improving video processing efficiency.

[0081] In some embodiments, the target duration can be determined based on the CPU processing capabilities of different client devices. The stronger the CPU processing capabilities of the client device, the shorter the target duration can be set. That is, client devices with different CPU processing capabilities can set different target durations. In this way, personalized customization can be achieved for different client devices.

[0082] In step S302 , in response to completing the operation of extracting a portion of the video frames of the original video, the client device transmits the target video frame to the server device.

[0083] In some embodiments, in response to completing the operation of extracting a portion of the video frames from the original video, the client device transmits the target video frame to the server device, including: in response to completing the operation of extracting the portion of the video frames from the original video, the client device transmits the target video frame to the intermediate device; and in response to completing the transmission of the target video frame to the intermediate device, the client device sends a second notification message to the server device. The second notification message is used to notify the server device to obtain the target video frame from the intermediate device. The client device sending the target video frame to the intermediate device can reduce storage pressure on the server device.

[0084] In some embodiments, the intermediate device storing the target video frame is a cloud device. The target video frame extracted from the original video is stored in the cloud with a larger storage space, which can reduce storage costs.

[0085] In step S303, the server device performs image processing on the target video frames transmitted from the client device. By cooperating with the client device to extract and transmit the original video frames in batches and performing image processing on the original video frames in batches, the server device can reduce the peak memory usage of the server device and cooperate with the client device to implement a pipeline operation of simultaneous extraction, transmission, and processing, thereby reducing the time consumption of the entire video processing process, improving the overall efficiency and usability of video processing, and enhancing the user experience.

[0086] In some embodiments, the image processing includes optical character recognition. An OCR model may be deployed on the server side, and the OCR model is used to perform optical character recognition on the target video frame.

[0087] In some embodiments, the video processing method further includes steps S304 to S306.

[0088] In step S304, the client device performs audio separation on the original video to obtain target audio. Audio separation is the process of extracting or extracting audio from the original video.

[0089] In step S305, in response to completing the audio separation operation on the original video, the client device transmits the target audio to the server device. The transmission of the target audio is parallel to the transmission of the target video frame. Since the audio file is very small, parallel processing of the target audio and target video frames can further improve the overall efficiency of video processing and reduce time consumption without having a significant impact on the peak memory usage of the client device. In some embodiments, the audio separation operation of the client device is also performed in parallel with the operation of extracting a portion of the video frame of the original video. That is, the audio separation operation and the operation of transmitting the target audio of the client device can be performed in parallel with the operation of extracting a portion of the video frame of the original video and the operation of transmitting the target video frame of the client device.

[0090] In some embodiments, in response to completing the audio separation operation on the original video, the client device transmits the target audio to the server device, including: in response to completing the audio separation operation on the original video, the client device transmits the target audio to the intermediate device; and in response to completing the transmission of the target audio to the intermediate device, the client device sends a first notification message to the server device. The first notification message is used to notify the server device to obtain the target audio from the intermediate device. The client device sending the target audio to the intermediate device can reduce storage pressure on the server device.

[0091] In some embodiments, the intermediate device storing the target audio is a cloud device. Storing the target audio in the cloud with larger storage space can reduce storage costs.

[0092] In step S306, the server device performs audio processing on the target audio transmitted from the client device. In some embodiments, the audio processing may include automatic speech recognition. An ASR model may be deployed on the server device to perform automatic speech recognition on the target audio.

[0093] In the above embodiment, steps S301 to S303 can be executed in parallel with steps S304 to S306, or partially in parallel, partially in series, or all in series. In the case of serial execution, the execution order of steps S301 to S306 can be reasonably set according to actual conditions. Figure 3 The illustration does not represent any fixed or exclusive order of execution.

[0094] In some embodiments, taking optical character recognition as image processing and automatic speech recognition as audio processing as an example, the video processing method further includes step S307. In step S307, the server-side device generates chapter information of the original video using a generation model based on the results of optical character recognition and the results of automatic speech recognition. For example, the generation model is an intelligent chapter model. The video processing method disclosed herein can be applied to video chapter generation scenarios, which can reduce the peak memory occupied by client devices, reduce the time consumption of video chapter generation, reduce user waiting time, and reduce the situation where users give up chapter creation due to lack of patience to wait for a long time for the video chapter generation process, thereby improving user experience and the usability of video chapter generation.

[0095] In some embodiments, the results of optical character recognition and automatic speech recognition are used by the server-side device to generate initial chapter information for the original video using a generative model. The initial chapter information is then desensitized using a security model to obtain the chapter information for the original video. For example, the security model can be a Transformer-based machine learning model.

[0096] In some embodiments, the video processing method further includes step S308. In step S308, the server device sends chapter information of the original video to the client device.

[0097] In some embodiments, the video processing method further includes step S309. In step S309, the client device publishes the original video carrying the chapter information. For example, the client device can publish the original video carrying the chapter information in response to a user's publishing operation.

[0098] Figure 4 A flowchart of a video processing method according to some embodiments of the present disclosure is shown.

[0099] like Figure 4As shown, the original video goes through the extraction and image processing pipeline operations of the client device extracting the target video frame, the client device transmitting the target video frame, and the server device performing image processing on the target video frame, thereby realizing a pipeline processing process of extraction, transmission (or uploading), and image processing (for example, OCR recognition) at the same time.

[0100] Through pipeline operations, the overall efficiency of video processing can be improved, and thus the efficiency of video chapter generation can be improved when applied in the field of video chapter generation. The client device implements fragment extraction and fragment delay, which can reduce the peak memory occupied by the client device, reduce the overheating of the client device, reduce the time consumed by video processing, and improve the overall efficiency and availability of video processing. The fragment extraction here refers to extracting video frames from the original video in batches, and the fragment delay refers to extracting video frames in batches for the original video, so that the client device can delay the extraction of video frames, thereby appropriately reducing the CPU utilization of the client device, thereby reducing the occurrence of overheating of the client device, reducing the peak memory occupied by the client device, and improving overall availability. The server device cooperates with the fragment extraction and transmission of the client device to implement fragment processing, which can reduce the peak memory occupied by the server device, reduce the time consumed by video processing, and improve the overall efficiency and availability of video processing.

[0101] In some embodiments, reference Figure 4 , the client device separates the audio from the original video, obtains the target audio, and transmits the target audio to the server device, and then the server device performs audio processing on the target audio.

[0102] In some embodiments, the server device generates initial chapter information of the original video using a generation model based on the image processing results and the audio processing results, and then processes the initial chapter information using a security model to obtain the chapter information of the original video.

[0103] For example, in a video chapter creation scenario, the user enters the chapter creation interface of the client device to automatically generate chapters. The client device initiates the process of extracting video frames and separating audio. The client device uses a segmented extraction method to extract a preset number of video frames every preset milliseconds. Each time a preset number of video frames are extracted, the extracted preset number of video frames are transmitted to the server device. After the server device obtains the preset number of video frames, it begins OCR recognition, thereby starting the pipeline process. After the last preset number of video frames are extracted, the server device completes OCR recognition of the previous frames. During this period, the audio obtained from audio separation is also transmitted to the server device, which performs automatic speech recognition. The server device finally summarizes the results of OCR recognition and automatic speech recognition into the generation model and security model, and returns the model calculation results to the client device.

[0104] Through the above embodiments, the efficiency of video chapter generation can be improved and the time of video chapter generation can be shortened.

[0105] For other embodiments described in the above embodiments, please refer to the above Figures 1 to 3 Different embodiments of the present invention will not be described in detail here.

[0106] Figure 5 A block diagram illustrating a client device according to some embodiments of the present disclosure is shown.

[0107] like Figure 5 As shown, the client device 51 includes an acquisition module 511 , an extraction module 512 , a transmission module 513 , a receiving module 514 and a display module 515 .

[0108] The acquisition module 511 is configured to acquire the original video, for example, by performing the following steps: Figure 1 In some embodiments, the acquisition module 511 may also execute Figure 1 The steps of any embodiment of step S110 are shown.

[0109] The extraction module 512 is configured to extract part of the video frames of the original video from the video frames that have not been extracted from the original video as the target video frames at every target duration, for example, by performing the following steps: Figure 1 In some embodiments, the extraction module 512 may also perform Figure 1 The steps of any embodiment of step S120 are shown.

[0110] The transmission module 513 is configured to transmit the target video frame to the server device in response to the completion of the operation of extracting the partial video frame of the original video, for example, performing the following steps: Figure 1 In some embodiments, the acquisition module 513 may also execute Figure 1 The steps of any embodiment of step S130 are shown.

[0111] The receiving module 514 is configured to receive the video processing result obtained by the server device based on image processing of the target video frame, for example, Figure 1 In some embodiments, the receiving module 514 may also perform Figure 1 The steps of any embodiment of step S140 are shown.

[0112] The display module 515 is configured to display the video processing results, for example, Figure 1 In some embodiments, the display module 515 can also perform step S150. Figure 1 The steps of any embodiment of step S150 are shown.

[0113] Figure 6 A block diagram of a server device according to some embodiments of the present disclosure is shown.

[0114] like Figure 6 As shown, the server device 62 includes an acquisition module 621 , an image processing module 622 , a determination module 623 and a sending module 624 .

[0115] The acquisition module 621 is configured to acquire a target video frame transmitted by the client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from the video frames not extracted from the original video at target time intervals. The operation of transmitting the target video frame from the client device to the server device is executed in response to the operation of the client device completing the extraction of the partial video frame of the original video, for example, executing Figure 2 In some embodiments, the acquisition module 621 may also execute Figure 2 The steps of any embodiment of step S210 are shown.

[0116] The image processing module 622 is configured to perform image processing on the target video frame in response to obtaining the target video frame, for example, performing the following steps: Figure 2 In some embodiments, the image processing module 622 may also perform Figure 2 The steps of any embodiment of step S220 are shown.

[0117] The determination module 623 is configured to determine the video processing result based on the image processing result, for example, Figure 2 In some embodiments, the determination module 623 may also perform Figure 2 The steps of any embodiment of step S230 are shown.

[0118] The sending module 624 is configured to send the video processing result to the client device, for example, Figure 2 In some embodiments, the sending module 624 may also perform Figure 2 The steps of any embodiment of step S240 are shown.

[0119] Figure 7 A block diagram of a video processing device according to some embodiments of the present disclosure is shown.

[0120] like Figure 7 As shown, the video processing device 7 includes a memory 71 and a processor 72 coupled to the memory 71. The memory 71 is used to store instructions for executing the corresponding embodiments of the video processing method. The processor 72 is configured to execute the video processing method in any of the embodiments of the present disclosure based on the instructions stored in the memory 71.

[0121] Memory 71 is used to store one or more computer-readable instructions. Memory 71 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 81 may, for example, store an operating system, application programs, a boot loader, a database, and other programs, as well as various application programs and data.

[0122] The processor 72 is used to run computer-readable instructions to implement the video processing method described in any of the above embodiments. The specific implementation of each step of the video processing method can be found in the above embodiments, and the repeated parts are not repeated here.

[0123] The processor 72 and the memory 71 can communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 82 and the memory 81 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0124] It should be noted that Figure 7 The components of the video processing device 7 shown are merely exemplary and non-limiting. The video processing device 7 may also have other components according to actual application requirements. The processor 72 may be combined with other components in the video processing device 7 to perform desired functions.

[0125] The video processing device can be implemented by software, firmware and / or hardware, and can be integrated into an electronic device installed with relevant application programs.

[0126] Figure 8 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0127] Figure 8 The electronic device 8 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.

[0128] Electronic devices include but are not limited to mobile terminals such as smart phones, laptops, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions, desktop computers, etc.

[0129] like Figure 8 As shown, the central processing unit (CPU) 81 performs various processes according to the program stored in the read-only memory (ROM) 82 or the program loaded from the storage part 88 to the random access memory (RAM) 83. In the RAM 83, data required when the CPU 81 performs various processes is stored as needed. The central processing unit is only exemplary and it can also be other types of processors, such as the various processors described above. The ROM 82, RAM 83 and the storage part 88 can be various forms of computer-readable storage media. It should be noted that although Figure 8 ROM 82, RAM 83 and storage portion 88 are shown separately in FIG, but one or more of them may be combined or located in the same or different memory or storage modules.

[0130] The CPU 81, the ROM 82, and the RAM 83 are connected to one another via a bus 84. An input / output interface 85 is also connected to the bus 84.

[0131] The following components are connected to the input / output interface 85: an input portion 86 such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 87 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 88 including a hard disk, a magnetic tape, etc.; and a communication portion 88 including a network interface card such as a LAN card, a modem, etc. The communication portion 88 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 8 The various devices or modules in the electronic device 8 are shown to communicate via a bus 84, but they may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0132] A drive 810 is also connected to the input / output interface 85 as needed. A removable medium 811 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 810 as needed so that a computer program read therefrom is installed in the storage section 88 as needed.

[0133] When the above-described series of processing is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 811 .

[0134] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which, when running on a computer, enables the computer to implement the video processing method described in any of the aforementioned embodiments. The computer program product includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 88, or installed from the storage part 88, or installed from the ROM 82. When the computer program is executed by the CPU 81, the video processing method of the embodiment of the present disclosure is executed.

[0135] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.

[0136] The computer readable medium may be a computer readable storage medium, or a computer readable signal medium, or any combination of the two.

[0137] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. A computer program is stored on a computer-readable storage medium that, when executed by a processor, implements the video processing method described in any of the aforementioned embodiments.

[0138] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0139] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0140] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the video processing method of any of the above embodiments. For example, the instructions may be embodied as computer program codes.

[0141] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure can be written in one or more programming languages ​​or combinations thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In situations involving a remote computer, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0143] The functions described above may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0144] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A video processing method, applied to a client device, comprising: Get the original video; At target time intervals, extracting a portion of video frames of the original video from video frames that have not been extracted from the original video as target video frames; In response to completing the operation of extracting the partial video frames of the original video, transmitting the target video frame to the server device; Receiving a video processing result obtained by the server device based on image processing performed on the target video frame; The video processing result is displayed.

2. The video processing method according to claim 1, further comprising: Performing audio separation on the original video to obtain target audio; In response to completing the audio separation operation on the original video, transmitting the target audio to the server device, wherein the transmission of the target audio is parallel to the transmission of the target video frame; Wherein, receiving the video processing result obtained by the server device based on image processing of the target video frame includes: receiving the video processing result obtained by the server device based on image processing of the target video frame and audio processing of the target audio.

3. The video processing method according to claim 2, wherein: The image processing includes optical character recognition, the audio processing includes automatic speech recognition, and the video processing result includes chapter information of the original video generated by the server device using a generation model based on the results of the optical character recognition and the automatic speech recognition.

4. The video processing method according to claim 2, wherein: In response to completing the audio separation operation on the original video, transmitting the target audio to the server device includes: In response to completing the audio separation operation on the original video, transmitting the target audio to an intermediate device; In response to completing the transmission of the target audio to the intermediate device, a first notification message is sent to the server device, where the first notification message is used to notify the server device to obtain the target audio from the intermediate device.

5. The video processing method according to claim 1, wherein: In response to completing the operation of extracting the partial video frames of the original video, transmitting the target video frame to the server device includes: In response to completing the operation of extracting the partial video frame of the original video, transmitting the target video frame to the intermediate device; In response to completing the transmission of the target video frame to the intermediate device, a second notification message is sent to the server device, where the second notification message is used to notify the server device to obtain the target video frame from the intermediate device.

6. The video processing method according to any one of claims 1 to 5, wherein the target duration is greater than or equal to the duration used to extract part of the video frames of the original video, and the difference between the target duration and the duration used to extract part of the video frames of the original video is less than a difference threshold.

7. A video processing method, applied to a server device, comprising: Obtaining a target video frame transmitted by a client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from video frames not extracted from the original video at target time intervals, and the client device transmitting the target video frame to the server device is performed in response to the client device completing the extraction of the partial video frame of the original video; In response to acquiring the target video frame, performing image processing on the target video frame; Determining a video processing result based on the image processing result; The video processing result is sent to the client device.

8. The video processing method according to claim 7, further comprising: Obtaining target audio transmitted by the client device, wherein the target audio is obtained by the client device performing audio separation on the original video, the client device transmitting the target audio to the server device in response to the client completing the audio separation operation on the original video, and the transmission of the target audio is performed in parallel with the transmission of the target video frame; In response to acquiring the target audio, performing audio processing on the target audio; Wherein, determining the video processing result based on the result of the image processing includes: determining the video processing result based on the result of the image processing and the result of the audio processing.

9. The video processing method according to claim 8, wherein: The image processing includes optical character recognition, the audio processing includes automatic speech recognition, and based on the results of the image processing and the audio processing, determining the video processing result includes: According to the result of the optical character recognition and the result of the automatic speech recognition, chapter information of the original video is generated as the video processing result by using a generation model.

10. The video processing method according to claim 8, wherein: Acquiring the target audio transmitted by the client device includes: In response to receiving a first notification message from the client device, the target audio is obtained from the intermediate device, wherein the target audio is transmitted to the intermediate device by the client device in response to completing the audio separation operation on the original video.

11. The video processing method according to claim 7, wherein: Acquiring the target video frame transmitted by the client device includes: In response to receiving a second notification message from the client device, the target video frame is obtained from the intermediate device, wherein the target video frame is transmitted to the intermediate device by the client device in response to completing the operation of extracting the partial video frame of the original video.

12. The video processing method according to any one of claims 7 to 11, wherein: The target duration is greater than or equal to a duration used to extract a portion of the video frames of the original video, and a difference between the target duration and a duration used to extract a portion of the video frames of the original video is less than a difference threshold.

13. A client device comprising: An acquisition module is configured to acquire original video; an extraction module configured to extract, at target time intervals, a portion of video frames of the original video from video frames that have not been extracted from the original video, as target video frames; a transmission module, configured to transmit the target video frame to a server device in response to completing the operation of extracting the partial video frame of the original video; A receiving module is configured to receive a video processing result obtained by the server device based on image processing performed on the target video frame; The display module is configured to display the video processing result.

14. A server device comprising: an acquisition module configured to acquire a target video frame transmitted by a client device, wherein the target video frame is a partial video frame of the original video extracted by the client device from video frames not extracted from the original video at target time intervals, and the operation of the client device transmitting the target video frame to the server device is performed in response to the client device completing the operation of extracting the partial video frame of the original video; an image processing module, configured to perform image processing on the target video frame in response to acquiring the target video frame; a determination module, configured to determine a video processing result based on the image processing result; The sending module is configured to send the video processing result to the client device.

15. A video processing device, comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the video processing method according to any one of claims 1 to 12 based on instructions stored in the memory. 16 . A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the video processing method according to claim 1 is implemented. 17 . A computer program product, which, when executed on a computer, enables the computer to implement the video processing method according to claim 1 .