Video processing method and device, electronic device and storage medium

The method uses playback behavior data and audio features to efficiently and accurately segment video content, addressing inefficiencies in manual marking and enhancing user experience by automating the skipping of unwanted segments.

JP7752720B2Active Publication Date: 2025-10-10BEIJING DUYOU INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024063553
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-04-10
Filing Date
2024-04-10
Publication Date
2025-10-10
Estimated Expiration
2044-04-10

AI Technical Summary

Technical Problem

Conventional methods for marking video content division points, such as openings, endings, and advertisements, are inefficient and labor-intensive, leading to high costs.

Method used

A video processing method that utilizes playback behavior data to determine approximate content division points and audio features to precisely locate these points, enabling efficient and accurate segmentation of video content.

Benefits of technology

Enables automatic and accurate identification of video content segments, improving user experience by allowing seamless skipping of uninteresting parts during playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752720000001
    Figure 0007752720000001
  • Figure 0007752720000002
    Figure 0007752720000002
  • Figure 0007752720000003
    Figure 0007752720000003
Patent Text Reader

Abstract

To provide a moving image processing method and apparatus for attaining efficient and accurate segmentation of moving image content, an electronic apparatus, a computer readable storage medium, and a computer program product.SOLUTION: A moving image processing method includes: obtaining playback behavior data of a moving image to be processed; determining, on the basis of the playback behavior data, a target moving image segment in which a content segmentation point of the moving image is located; extracting an audio feature of the target moving image segment; and determining the content segmentation point from the target moving image segment on the basis of the audio feature.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Detailed Description of the Invention

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of deep learning and computer vision. [Technical Field]

[0002] The present disclosure relates to the computer technology field, in particular to the multimedia technology field, and more particularly to a video processing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. [Background technology]

[0003] The same video may contain different types of content, such as an opening, a main part (the main body of the video), advertisements, an ending, etc. Users have different levels of interest in different types of content. Locating different types of content within a video can make it easier for users to view content that interests them.

[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or adopted. Unless otherwise noted, any approach described in this section should not be considered prior art merely because it is included in this section. Likewise, unless otherwise noted, the problems addressed in this section should not be considered to be acknowledged in the prior art. Summary of the Invention

[0005] The present disclosure provides a video processing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product. According to one aspect of the present disclosure, a video processing method is provided, including: obtaining playback behavior data of a video to be processed; determining a target video segment in which a content division point of the video is located based on the playback behavior data, wherein a type of video content located before the content division point is different from a type of video content located after the content division point; extracting audio features of the target video segment; and determining the content division point from the target video segment based on the audio features.

[0006] According to one aspect of the present disclosure, a video processing device is provided, including: an acquisition module configured to acquire playback behavior data of a video to be processed; a first determination module configured to determine a target video segment in which a content division point of the video is located based on the playback behavior data, wherein the type of video content located before the content division point is different from the type of video content located after the content division point; an extraction module configured to extract audio features of the target video segment; and a second determination module configured to determine the content division point from the target video segment based on the audio features.

[0007] According to one aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the above-described method.

[0008] According to one aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform the method described above.

[0009] According to one aspect of the present disclosure, there is provided a computer program product comprising computer program instructions which, when executed by a processor, implements the method described above.

[0010] According to one or more embodiments of the present disclosure, efficient and accurate segmentation of video content can be achieved. It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, and are not intended to limit the scope of protection of the present disclosure. Other features of the present disclosure will be easily understood from the following description. [Brief explanation of the drawings]

[0011] The drawings illustratively illustrate examples, constitute a part of the specification, and together with the written description serve to explain exemplary embodiments of the examples. The illustrated examples are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar, but not necessarily identical, elements.

[0012] [Figure 1] FIG. 1 is a schematic diagram illustrating an exemplary system capable of implementing the methods described herein, according to an embodiment of the present disclosure. [Figure 2] 1 is a flowchart illustrating a video processing method according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a schematic diagram illustrating a video processing system according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is an interactive timing diagram illustrating a video processing system according to an embodiment of the present disclosure. [Figure 5] 1 is a flowchart illustrating a video processing process according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a block diagram illustrating a configuration of a moving image processing device according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a block diagram illustrating the structure of an exemplary electronic device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013]

[0023] The following description will be made in conjunction with the drawings to illustrate exemplary embodiments of the present disclosure. Various details of the embodiments of the present disclosure are included to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the following description omits descriptions of known functions and structures.

[0014] In this disclosure, unless otherwise specified, the terms "first," "second," and the like, used to describe various elements are not intended to limit the location, timing, or importance of these elements. Such terms are used only to distinguish one element from another. In some instances, a first element and a second element may refer to the same instance of an element, or in some cases, may refer to different instances based on the context.

[0015] The terms used in the description of various examples of the present disclosure are intended only to describe particular examples and are not intended to be limiting. Unless the context clearly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any and all possible combinations of the listed items.

[0016] In the technical solution disclosed herein, the acquisition, storage and application of relevant user personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and good morals. When playing a video, users tend to skip over content they are not interested in, such as openings, endings, and advertisements, and only watch the main part. In conventional technology, video producers or video playback platform operators typically manually mark division points in video content (e.g., the positions of openings, endings, advertisements, etc.), which is inefficient and requires high labor costs.

[0017] In view of the above problems, an embodiment of the present disclosure provides a video processing method, which can realize efficient and accurate segmentation of video content. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.

[0018] 1 illustrates a schematic diagram of an exemplary system 100 in which various methods and apparatus described herein may be implemented, according to embodiments of the present disclosure. Referring to FIG. 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to run one or more applications.

[0019] In embodiments of the present disclosure, client devices 101, 102, 103, 104, 105, 106 and server 120 may run one or more services or software applications that may implement video processing methods of embodiments of the present disclosure.

[0020] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized and virtualized environments. In some embodiments, these services may be provided as web-based or cloud services, for example, provided to users of client devices 101, 102, 103, 104, 105, and / or 106 in a Software as a Service (SaaS) model.

[0021] In the configuration shown in FIG. 1 , server 120 may include one or more units that implement the functions performed by server 120. These assemblies may include software assemblies, hardware assemblies, or a combination thereof, executable on one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may, in turn, utilize one or more client applications to interact with server 120 to utilize services provided by these assemblies. It should be understood that a variety of different system configurations are possible and may differ from system 100. Thus, FIG. 1 is intended to be illustrative and not limiting of a system for implementing various methods described herein.

[0022] Client devices 101, 102, 103, 104, 105, and / or 106 may provide an interface through which a user of the client device interacts with the client device. The client device may also output information to the user through the interface. Although only six client devices are shown in FIG. 1, one skilled in the art will understand that the present disclosure can support any number of client devices.

[0023] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (e.g., personal computers or laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, in-vehicle equipment, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems, or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include mobile phones, intelligent phones, tablets, personal digital assistants (PDAs), and the like. Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, and the like. The client device may run a variety of applications, such as Internet-related applications, communication applications (eg, email applications), and short message service (SMS) applications, and may use a variety of communication protocols.

[0024] Network 110 may be any type of network known to those skilled in the art, which may use any one of several available protocols to support data communications (including, but not limited to, TCP / IP, SNA, IPX, etc.) By way of example, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token loop, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, Wi-Fi), and / or any combination of these and / or other networks.

[0025] Server 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 may also include one or more virtual machines running virtual operating systems or other computing architectures involving virtualization (e.g., one or more flexible pools of virtualized logical storage devices to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0026] The computing units in server 120 may run one or more operating systems, including any of the operating systems listed above and any commercial server operating system. Server 120 may also run any one of a variety of additional server and / or middle-tier applications, such as an HTTP server, an FTP server, a CGI server, a JAVA server, a database server, etc.

[0027] In some embodiments, server 120 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0028] In some embodiments, server 120 may be a server in a distributed system or a server incorporating blockchain. Server 120 may be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system that solves the drawbacks of traditional physical hosts and virtual private server (VPS) services, such as high management difficulty and poor business scalability.

[0029] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data or other information. For example, one or more of databases 130 may be used to store information such as audio files or video files. Databases 130 may be located in a variety of locations. For example, a database used by server 120 may be local to server 120 or may be remote from server 120 and in communication with server 120 over a network or dedicated connection. Databases 130 may be of a variety of types. In some embodiments, a database used by server 120 may be a relational database. One or more of these databases may store, update, and retrieve data from the databases in response to instructions.

[0030] In some embodiments, one or more of the databases 130 may be used by an application to store data for the application. The databases used by the application may be various types of databases, such as a key-value repository, an object repository, or a general-purpose repository supported by a file system.

[0031] The system 100 of FIG. 1 can be configured and operated in a variety of ways to accommodate the various methods and apparatus described in accordance with this disclosure. According to some embodiments, the client devices 101-106 may include a client application (e.g., a netdisk client, a video client, etc.) for playing videos. Accordingly, the server 120 is a server supporting the client application. The server 120 may determine content segmentation points of a video, such as an opening end point, an ending start point, or an advertisement end point, by executing a data processing method according to an embodiment of the present disclosure. The determined content segmentation points of a video may be associated with the video and stored in the database 130. When the client devices 101-106 initiate a playback request for the video, the server 120 may return video data of the video along with its content segmentation points to the user, allowing the client devices 101-106 to play the video. During the video playback process, the client devices 101-106 may automatically or manually (i.e., based on the user's selection) skip content such as an opening or ending based on the content segmentation points of the video, thereby improving the user's video playback experience.

[0032] In some embodiments, the client devices 101-106 may also determine content segmentation points of a video by executing a data processing method according to an embodiment of the present disclosure, which generally requires a higher hardware configuration and computing power of the client devices 101-106.

[0033] 2 shows a flowchart of a video processing method 200 according to an embodiment of the present disclosure. As described above, the execution entity of method 200 is typically a server (e.g., the above-described server 120). In some embodiments, the execution entity of method 200 may be a client device (e.g., the above-described client devices 101-106). As shown in FIG. 2, method 200 includes steps S210-S240.

[0034] In step S210, playback behavior data of the video to be processed is acquired. In step S220, a target video segment where a content division point of the video is located is determined based on the playback behavior data, and the type of the video content located before the content division point is different from the type of the video content located after the content division point.

[0035] In step S230, audio features of the target video segment are extracted. In step S240, content division points are determined from the target video segment based on the audio features.

[0036] According to an embodiment of the present disclosure, the approximate location of a content division point of a video, i.e., a target video segment, is determined based on user playback behavior data for the video, and then the exact location of the content division point is determined based on the audio characteristics of the target video segment, thereby realizing efficient and accurate division of video content.

[0037] Each step of method 200 is described in detail below. In step S210, playback behavior data of the video to be processed is acquired. The video to be processed may include different types of video content, including, but not limited to, an opening, a main part (i.e., the main body of the video), an ending, etc. In an embodiment of the present disclosure, a temporal boundary point between different types of video content is recorded as a content division point. The type of video content located before the content division point is different from the type of video content located after the content division point. The content division point includes, for example, an opening end point (i.e., a main part start point) for separating the opening and the main part, and an ending start point (i.e., a main part end point) for separating the main part and the ending. If an advertisement is inserted into the main part, the content division point may further include an advertisement end point (i.e., a main part start point) for separating the advertisement from the main part, etc.

[0038] In order to improve the accuracy of dividing video content, it is necessary to acquire playback behavior data from multiple users in step S210. Playback behavior data is behavior data that a user generates while playing a video.

[0039] According to some embodiments, the viewing behavior data includes, for example, the number of views of a video. According to some embodiments, playback behavior data includes the types of interactions a user has with the video, the playback point at which the video is located when a user interacts with the video, and the like.

[0040] The user's interactive operation on the video may be, for example, a playback selection operation. The playback selection operation is used to select the point at which to start playback of the video. In other words, the playback point corresponding to the playback selection operation is the playback start point selected by the user. For example, the playback selection operation may be an operation of dragging a pointing control on a progress bar, i.e., a drag operation. The playback point corresponding to the drag operation is the video point corresponding to the position of the pointing control when the user finishes the drag operation. Also, for example, the playback selection operation may be an operation of the user inputting the selected point into a text box, i.e., a time point input operation. The time point input by the user is the playback point corresponding to the time point input operation.

[0041] The user's interactive operation on the video may be, for example, a playback termination operation. The playback time corresponding to the playback termination operation is the time at which the video is located when the user terminates video playback, i.e., the playback end time. Specifically, the playback termination operation may be, for example, an operation to close or exit the video playback interface.

[0042] The user's playback behavior data can represent the user's level of interest in video content and, ultimately, changes in the video content. This allows the location of content division points to be quickly identified based on the playback behavior data of multiple users.

[0043] In step S220, a target video segment where a content division point of the video is located is determined based on the playback behavior data. According to some embodiments, as described above, the playback behavior data includes a playback time point at which the video is positioned when the user interacts with the video. Accordingly, step S220 may include the following steps S222 to S226.

[0044] In step S222, the video is divided into a plurality of video segments of target duration. In step S224, for any one of the plurality of video segments, an interactive count of the video segment is determined, where the interactive count is the number of interactive operations whose playback time point is located in the video segment.

[0045] In step S226, a target video segment is determined from the plurality of video segments based on the number of interactions of each of the plurality of video segments.

[0046] According to the above-described embodiment, a user's interactive operation on a video can reflect the user's level of interest in the video content and, ultimately, changes in the video content. For example, a user is usually not interested in the opening, ending, advertisements, etc. of a video content. When playing the opening or advertisement, the user is likely to perform a playback selection operation to skip these video contents and go straight to the main part. When playing the ending, the user is likely to perform a playback stop operation. This allows the approximate location of the content division point to be quickly determined based on the user's interactive operation.

[0047] In step S222, the video to be processed is divided into multiple video segments of the same duration. The duration of each video segment is a target duration. The target duration may be any value, such as 10 seconds, 30 seconds, or 1 minute. According to some embodiments, the target duration may have a positive correlation with the total duration of the video. That is, the target duration is set to be larger as the total duration of the video increases. This can improve video processing efficiency.

[0048] In step S224, the number of interactions for each video segment can be determined based on the user's playback behavior data, where the number of interactions for a video segment is the number of interactive operations whose playback time is located in the video segment.

[0049] According to some embodiments, different types of interactive operations can be selected to identify different types of content segmentation points, thereby improving the accuracy of content segmentation. Therefore, the specific implementation details of steps S224 and S226 are different for different types of content segmentation points.

[0050] According to some embodiments, the opening end point can be identified by selecting a playback selection operation. Thus, the playback behavior data includes a playback time point selected by the user in the playback selection operation and a continuous playback time length by the user from the playback time point. Accordingly, step S224 can further include steps S2242 and S2244.

[0051] In step S2242, an interactive operation whose continuous playback time is longer than a first threshold is determined as a valid interactive operation. The first threshold may be, for example, 3 seconds, 5 seconds, or the like.

[0052] In step S2244, the number of valid interactive operations whose playback time points are located in the video segment is determined as the number of interactions. Users are usually not interested in openings. If the continuous playback time after the user selects a certain playback time point is long, it indicates that the user has skipped the opening with this operation, that is, reached the opening end point. If the continuous playback time after the user selects a certain playback time point is short and the user immediately performs a next playback selection operation, it indicates that the user's current operation has not skipped the opening, that is, has not reached the opening end point. According to the above embodiment, it is possible to filter out invalid interactive operations (which do not skip the opening) with short continuous playback times, thereby improving the accuracy of opening identification.

[0053] According to some embodiments, if the content division point to be identified is an opening end point, step S226 may include determining a video segment with the largest number of interactions within a first time range as a target video segment where the opening end point is located, where the first time range is a time range from the start point of the video to a first time point, such as the fourth or fifth minute of the video.

[0054] According to the above embodiment, the target video segment where the opening end point is located can be determined from the start point of the video, thereby improving the accuracy of opening identification.

[0055] According to some embodiments, a playback end operation can be selected to identify an ending start point. Accordingly, the playback behavior data includes a playback end time corresponding to the playback end operation. Step S224 can include determining the number of playback end operations whose playback end time is located in a video segment as the number of interactions for the video segment.

[0056] According to some embodiments, if the content division point to be identified is the ending start point, step S226 may include determining the video segment with the highest number of interactions in the second time range as the target video segment where the ending start point is located. Here, the second time range is the time range from a second time point to the end point of the video. The second time point may be, for example, four minutes or five minutes from the end of the video.

[0057] Generally, users are not interested in the ending. When a video is played to the ending, the user usually ends the current video and continues to play the next video, or chooses not to play the video at all. According to the above embodiment, the target video segment where the ending start point is located can be determined from the ending of the video, thereby improving the accuracy of ending identification.

[0058] In some embodiments, if the content segmentation point to be identified is an advertisement end point, step S224 may include steps S2242 and S2244, and step S226 includes determining the video segment with the largest number of interactions within a third time range as the target video segment where the opening end point is located. The first time range is the time range from the first time point to the second time point.

[0059] Users are usually not interested in advertisements inserted into the main content. When an advertisement starts playing, the user usually drags the progress bar to skip the advertisement content. If the continuous playback time after the user selects a playback time point (dragging to a certain position on the progress bar) is long, it indicates that the user has skipped the advertisement with this operation, i.e., reached the advertisement end point. If the continuous playback time after the user selects a playback time point is short and the user immediately performs another drag operation, it indicates that the user has not skipped the advertisement with this operation, i.e., has not reached the advertisement end point. According to the above-mentioned embodiment, invalid interactive operations (which do not skip advertisements) with short continuous playback times are filtered out, and a target video segment where the advertisement end point is located is determined from an intermediate segment of the video, thereby improving the accuracy of identifying advertisements to be inserted into the main content.

[0060] According to some embodiments, the playback behavior data includes the number of times the video has been played. Accordingly, step S220 may be performed in response to the number of times the video has been played being greater than a second threshold. That is, in response to the number of times the video has been played being greater than the second threshold, determining a target video segment where a content division point of the video is located based on the playback behavior data. The second threshold may be, for example, 100, 500, etc.

[0061] When a video has been played a certain number of times, the obtained play behavior data is large and has more statistical significance. According to the above embodiment, when a video has been played a certain number of times, the video content division can be triggered, thereby improving the accuracy of the video content division.

[0062] According to some embodiments, the target duration for video segmentation and the second threshold for the video play count can be determined based on the content segmentation point labels of the sample videos. For example, the target video segments where the content segmentation points of each sample video are located can be determined through the above-described steps S222 to S226 based on the specified target duration and second threshold. The identification accuracy rate of the target video segment at the current target duration and second threshold of the sample video can be determined by comparing the determined target video segment (i.e., predicted value) with the true target video segment (i.e., true value) where the content segmentation point is located. The target duration and second threshold value that provide the highest identification accuracy rate are determined as the optimal target duration and second threshold value.

[0063] In step S230, audio features of the target video segment are extracted. According to some embodiments, step S230 may include steps S232 and S234.

[0064] In step S232, a Fourier transform is performed on the audio data of the target video segment to obtain a frequency spectrum corresponding to the audio data. In step S234, feature extraction is performed on the frequency spectrum to obtain voice features.

[0065] According to the above embodiment, by extracting the frequency domain audio features of the target video segment, the basic features of the audio data can be maintained while achieving data compression, thereby improving the efficiency and accuracy of video content segmentation.

[0066] According to some embodiments, the audio features may be, for example, Mel-frequency cepstral coefficients (MFCCs). Step S234 may include converting the frequency spectrum obtained in step S222 into a Mel spectrum using a Mel-filter bank with the same bank area, and performing cepstral analysis on the Mel spectrum to obtain Mel-frequency cepstral coefficients. Specifically, the cepstral analysis includes operations such as logarithmic operation and discrete cosine transform (DCT). The second to thirteenth coefficients after the discrete cosine transform are taken as Mel-frequency cepstral coefficients.

[0067] In step S240, content division points are determined from the target video segment based on the audio features. According to some embodiments, content segmentation points of a video are determined based on a mapping relationship between pre-defined audio features and content segmentation points. According to these embodiments, the precise positions of content segmentation points can be quickly determined. For example, the content segmentation points can be determined accurately to the second level.

[0068] The mapping relationship between the audio feature and the content segmentation point is the formula y=f(x), where the argument x is the audio feature and the variable y is the offset time of the content segmentation point in the target video segment.

[0069] According to some embodiments, the mapping relationship between audio features and content division points may be determined based on the content division point labels of the sample video and the audio features of the sample target video segments where the content division point labels are located. For example, the audio features of the sample target video segments of each sample video are extracted through the above steps S232 to S234. A data pair (x0, y0) consisting of the audio feature x0 of the sample target video segment and the offset time y0 of the content division point label in the sample target video segment is used as sample data, and fitting is performed to obtain a mapping equation y=f(x) between the audio feature x and the content division point y.

[0070] 3 shows a schematic diagram of a video processing system 300 according to an embodiment of the present disclosure. As shown in FIG. 3, the video processing system 300 includes a behavioral data collection module 310, a message queue 320, a behavioral data analysis module 330, a distributed cache 340, an audio analysis module 350, and a database 360.

[0071] The behavioral data collection module 310 is used to collect user playback behavior data. The playback behavior data can be recorded in a playback log by a video playback SDK (Software Development Kit), and the playback log can be written to an asynchronous message queue 320. The playback behavior data includes, for example, the start position and target position of the user's dragging of the progress bar, and the position where the user ends the video playback.

[0072] The behavior data analysis module 330 consumes the playback behavior data in the message queue 320 and temporarily caches the playback behavior data in the distributed cache 340. When the amount of cached data reaches a threshold (corresponding to the above-mentioned "second threshold"), it analyzes the playback behavior data and determines the target video segment where the content division point is located.

[0073] The audio analysis module 350 extracts audio features of the target video segment, determines the exact positions of content segmentation points based on the audio features, and writes the determined content segmentation points to the database 360. The content segmentation points of a video can be provided to a user who subsequently plays the video to be used to skip video content such as openings and endings based on the content segmentation points.

[0074] 4 illustrates an interactive timing diagram of a video processing system according to an embodiment of the present disclosure. In the embodiment illustrated in FIG. 4, the video processing system includes a behavioral data collection module 410, a message queue 420, a behavioral data analysis module 430, a distributed cache 440, an audio analysis module 450, and a database 460.

[0075] In step S471, the behavioral data collection module 410 collects playback behavior data of the user (user A) while the user is playing a video.

[0076] In step S472 , the behavioral data collection module 410 writes the playback behavioral data to the message queue 420 . In step S473, the behavioral data collection module 410 receives a message from the message queue 420 indicating that the playback behavioral data has been successfully written.

[0077] In step S474, the behavioral data analysis module 430 consumes the playback behavioral data in the message queue 420, and in step S475, temporarily stores the playback behavioral data in the distributed cache 440.

[0078] In step S476, the amount of data cached in the distributed cache 440 reaches a threshold (corresponding to the above-mentioned "second threshold"), and the behavioral data analysis module 430 analyzes the playback behavior data to determine the target video segment where the content division point is located.

[0079] In step S477, the audio analysis module 450 obtains the target video segment obtained by the behavior data analysis module 430, extracts audio features of the target video segment, and determines the precise position of the content division point based on the audio features.

[0080] In step S478, the audio analysis module 450 writes the determined content segmentation points to the database 460. In step S479, when another user (user B different from user A) watches the video, the video playback platform obtains the content division points of the video from the database 460 and provides them to the client device used by the user, so that the opening and ending can be automatically skipped for the user during the process of playing the video.

[0081] 5 shows a flowchart of a video processing process according to an embodiment of the present disclosure. In the embodiment shown in FIG. 5, the video processing system includes a behavioral data collection module 510, a message queue 520, a behavioral data analysis module 530, a distributed cache 540, an audio analysis module 550, and a database 560.

[0082] In step S591 , during the process in which the client 570 plays the video, the behavior data collection module 510 collects the playback behavior data of the client 570 via the gateway 580 and writes it to the message queue 520 .

[0083] In step S592 , the behavioral data analysis module 530 consumes the playback behavioral data in the message queue 520 and temporarily stores the playback behavioral data in the distributed cache 540 .

[0084] In step S593, the behavioral data analysis module 530 determines whether the amount of data cached in the distributed cache 540 has reached a threshold (corresponding to the above-mentioned "second threshold"), and if so, executes step S594; otherwise, returns to step S592 and continues consuming the playback behavioral data in the message queue 520.

[0085] In step S594, the behavioral data analysis module 530 analyzes the playback behavioral data to identify the target video segment in which the content division point is located.

[0086] In step S595, the audio analysis module 550 obtains the target video segment obtained by the behavioral data analysis module 530, extracts audio features of the target video segment, identifies precise locations of content division points based on the audio features, and writes the precise locations of the content division points into the database 560.

[0087]

[0033] According to an embodiment of the present disclosure, a video processing device is further provided. Fig. 6 shows a block diagram of a configuration of a video processing device 600 according to an embodiment of the present disclosure. As shown in Fig. 6, the device 600 includes an acquisition module 610, a first determination module 620, an extraction module 630, and a second determination module 640.

[0088] The acquisition module 610 is configured to acquire playback behavior data of the video to be processed. The first determination module 620 is configured to determine, based on the playback behavior data, a target video segment in which a content division point of the video is located, where the type of video content located before the content division point is different from the type of video content located after the content division point.

[0089] The extraction module 630 is configured to extract audio features of the target video segment. A second determining module 640 is configured to determine the content division points from the target video segments based on the audio features.

[0090] According to an embodiment of the present disclosure, the approximate location of a content division point of a video, i.e., a target video segment, is determined based on user playback behavior data for the video, and then the exact location of the content division point is determined based on the audio characteristics of the target video segment, thereby realizing efficient and accurate division of video content.

[0091] According to some embodiments, the playback behavior data includes a playback time point at which the video is located when a user interacts with the video, wherein the first determination module includes: a division unit configured to divide the video into a plurality of video segments of a target time length; a first determination unit configured to determine, for any one of the plurality of video segments, an interaction count for the video segment, wherein the interaction count is the number of interactive operations whose playback time point is located in the video segment; and a second determination unit configured to determine the target video segment from the plurality of video segments based on the interaction count of each of the plurality of video segments.

[0092] According to some embodiments, the interactive operation includes a playback selection operation, the playback behavior data further includes a continuous playback time length from the playback time point, and the content division point includes an opening end point, wherein the first determination unit is further configured to determine an interactive operation whose continuous playback time length is greater than a first threshold as a valid interactive operation, and to determine the number of valid interactive operations whose playback time point is located in the video segment as the number of interactions.

[0093] According to some embodiments, the second determination unit is further configured to determine the video segment with the highest number of interactions in a first time range as the target video segment, where the first time range is a time range from the start point of the video to a first time point.

[0094] According to some embodiments, the interactive operation includes a playback end operation, the content division point includes an ending start point, and the second determination unit is further configured to determine the video segment with the largest number of interactions in a second time range as the target video segment, and the second time range is a time range from a second time point to an end point of the video.

[0095] According to some embodiments, the playback behavior data includes a number of plays of the video, and the first determination module is further configured to determine a target video segment in which a content division point of the video is located based on the playback behavior data in response to the number of plays being greater than a second threshold.

[0096] According to some embodiments, the extraction module includes a transformation unit configured to perform a Fourier transform on audio data of the target video segment to obtain a frequency spectrum corresponding to the audio data, and an extraction unit configured to perform feature extraction on the frequency spectrum to obtain the audio features.

[0097] According to some embodiments, the second determination module is further configured to determine the content segmentation points based on a mapping relationship between a preset audio feature and a content segmentation point.

[0098] According to some embodiments, the mapping relationship is determined based on content division point labels of a sample video and audio features of a sample target video segment in which the content division point labels are located.

[0099] It should be understood that each module and unit of the apparatus 600 shown in Figure 6 may correspond to each step of the method 200 described with reference to Figure 2. Accordingly, the operations, features, and advantages described above with respect to the method 200 are equally applicable to the apparatus 600 and the modules and units included therein. For the sake of brevity, some operations, features, and advantages will not be described here.

[0100] Although particular functionality is discussed above with reference to particular modules, it should be noted that the functionality of each module discussed herein may be split into multiple modules and / or at least some of the functionality of multiple modules may be combined into a single module.

[0101] It should also be understood that various techniques may be described herein in the general context of software hardware elements or program modules. Each unit described in FIG. 6 above may be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units may be implemented as computer program code / instructions configured to be executed on one or more processors and stored on a computer-readable storage medium. Alternatively, these units may be implemented as hardware logic / circuitry. For example, in some embodiments, one or more of modules 610-640 may be implemented together in a System on Chip (SoC). An SoC may include an integrated circuit chip (e.g., a processor (e.g., including a Central Processing Unit (CPU), microcontroller, microprocessor, Digital Signal Processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components in other circuits) that may optionally execute received program code and / or include embedded firmware to perform functions.

[0102] According to an embodiment of the present disclosure, there is further provided an electronic device, the electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform a video processing method according to an embodiment of the present disclosure.

[0103] According to an embodiment of the present disclosure, there is further provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to perform a video processing method according to an embodiment of the present disclosure.

[0104] According to an embodiment of the present disclosure, there is further provided a computer program product including computer program instructions, which, when executed by a processor, implements a video processing method according to an embodiment of the present disclosure.

[0105] Next, referring to FIG. 7 , a block diagram of an electronic device 700 functioning as a server or client of the present disclosure will be described, which is an example of a hardware device applicable to various aspects of the present disclosure. The electronic device may represent various forms of digital electronic computers, such as laptop computers, desktop computers, stage computers, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components, their connections, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.

[0106] 7, electronic device 700 includes a computing unit 701 that can perform various appropriate operations and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or loaded from a storage unit 708 into a random access memory (RAM) 703. RAM 703 may also store various programs and data necessary for the operation of electronic device 700. Computing unit 701, ROM 702, and RAM 703 are connected to one another via a bus 704. An input / output (I / O) interface 705 is also connected to bus 704.

[0107] Several components of the electronic device 700, including an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709, are connected to the I / O interface 705. The input unit 706 may be any type of device capable of inputting information to the electronic device 700. The input unit 706 may receive input numeric or character information and generate key signal input for user settings and / or function control of the electronic device, including, but not limited to, a mouse, keyboard, touch screen, trackboard, trackball, joystick, microphone, and / or remote control. The output unit 707 may be any type of device capable of presenting information, including, but not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 708 may include, but is not limited to, a magnetic disk or an optical disk. The communication unit 709 enables the electronic device 700 to exchange information / data with other devices via a computer network, e.g., the Internet, and / or various telecommunications networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, e.g., a Bluetooth device, an 802.11 device, a Wi-Fi device, a WiMAX device, a cellular communication device, and / or the like.

[0108] The computing unit 701 can be various general-purpose and / or special-purpose processing components having processing and computational capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, some or all of the computer program may be loaded and / or installed into the electronic device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, it may perform one or more steps of method 200 described above. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the method 200 using any other suitable means (eg, firmware).

[0109] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being embodied in one or more computer programs that may be executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and may receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0110] Program code implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when executed by the processor or controller, the program code performs the functions / operations specified in the flowcharts and / or block diagrams. The program code may be entirely executed by machine, partially executed by machine, partially executed by machine and partially executed by a remote machine as a separate software package, or entirely executed on a remote machine or server.

[0111] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may include or store a program for use in or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection with one or more leads, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0112] To provide for user interaction, a computer may implement the systems and techniques described herein and include a display device (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction, for example, providing feedback to a user in any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and receiving input from a user in any form (including sound input, speech input, or tactile input).

[0113] The systems and techniques described herein may be implemented in a computing system including backstage components (e.g., as a data server), middleware components (e.g., as an application server), front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with the system or technique implementation), or any combination of backstage components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0114] The computer system may include a client and a server. The client and the server are generally remote from each other and usually interact via a communication network. The client-server relationship is created by running computer programs on corresponding computers. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.

[0115] It should be understood that the various forms of flow described above may be used to rearrange, add, or remove steps, and for example, the steps described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the technical solutions disclosed in this disclosure can achieve the desired results, and the present disclosure is not limited thereto.

[0116] Although embodiments or examples of the present disclosure have been described with reference to the drawings, it should be understood that the above-described methods, systems, and devices are merely exemplary embodiments or examples, and that the scope of the present disclosure is not limited by these embodiments or examples, but only by the scope of the appended claims and their equivalents. Various elements of the embodiments or examples may be omitted or replaced by equivalent elements. Furthermore, steps may be performed in an order different from that described in this disclosure. Furthermore, various elements of the embodiments or examples may be combined in various ways. Importantly, as technology evolves, many elements described herein may be replaced by equivalent elements that appear later in this disclosure.

Claims

1. A video processing method, comprising: Obtaining playback behavior data of a video to be processed, where the video includes different types of video content, including opening credits, main content, and ending credits, the playback behavior data being behavior data generated by multiple users during the playback process of the video, and the playback behavior data including playback times of the video when users perform interactive operations on the video; determining a target video segment where a content division point of the video is located based on the playback behavior data, wherein a type of video content located before the content division point is different from a type of video content located after the content division point, and wherein the content division point includes an opening end point used to separate an opening and a main part, or an ending start point used to separate a main part and an ending, wherein determining a target video segment where a content division point of the video is located based on the playback behavior data includes: Dividing the video into a plurality of video segments of the same target duration; For any one of the plurality of video segments, determining an interactive count of the video segment, where the interactive count is the number of interactive operations whose playback time points are located in the video segment; determining the target video segment from the plurality of video segments based on a number of interactions for each of the plurality of video segments; The video processing method includes: extracting audio features of the target video segment; determining the content division points from the target video segments based on the audio features.

2. The interactive operation includes a play selection operation, the play behavior data further includes a continuous play time length from the play time point, and the content division point includes an opening end point, wherein determining the number of interactions of the video segment includes: determining an interactive operation having a continuous playback time length greater than a first threshold as a valid interactive operation; The method of claim 1 , further comprising determining the number of valid interactive operations whose playback time points are located in the video segment as the number of interactions.

3. Determining the target video segment from the plurality of video segments based on the number of interactions of each of the plurality of video segments includes:

3. The method of claim 2, further comprising determining a video segment having the greatest number of interactions in a first time range as the target video segment, wherein the first time range is a time range from a start point of the video to a predetermined first time point.

4. The interactive operation includes a playback end operation, and the content division point includes an ending start point, and determining the target video segment from the plurality of video segments based on the number of interactions of each of the plurality of video segments includes:

2. The method of claim 1, further comprising determining a video segment having the greatest number of interactions in a second time range as the target video segment, wherein the second time range is a time range from a second predetermined point in time to an end point of the video.

5. The playback behavior data includes a number of times the video is played, and wherein determining a target video segment in which a content division point of the video is located based on the playback behavior data includes: In response to the number of plays being greater than a second threshold, The method of claim 1 , further comprising: determining a target video segment in which a content division point of the video is located.

6. Extracting audio features of the target video segment, performing a Fourier transform on the audio data of the target video segment to obtain a frequency spectrum corresponding to the audio data; and performing feature extraction on the frequency spectrum to obtain the audio features.

7. Determining the content division point from the target video segment based on the audio features comprises: The method of claim 1 , further comprising determining the content segmentation points based on a mapping relationship between a preset audio feature and a content segmentation point.

8. The method of claim 7 , wherein the mapping relationship is determined based on content segmentation point labels of a sample video and audio features of a sample target video segment in which the content segmentation point labels are located.

9. A video processing device, An acquisition module configured to acquire playback behavior data of a video to be processed, where the video includes different types of video content, including opening credits, main content, and ending credits, and the playback behavior data is behavior data generated by multiple users during the playback process of the video; A first determination module configured to determine a target video segment where a content division point of the video is located based on the playback behavior data, where the type of video content located before the content division point is different from the type of video content located after the content division point, where the content division point includes an opening end point used to separate an opening and a main part, or an ending start point used to separate a main part and an ending, where the first determination module: a segmentation unit configured to segment the video into a plurality of video segments of the same target duration; a first determination unit configured to determine, for any one of the plurality of video segments, an interactive count of the video segment, where the interactive count is a number of interactive operations whose playback time points are located in the video segment; a second determination unit configured to determine the target video segment from the plurality of video segments based on an interaction count for each of the plurality of video segments; The video processing device comprises: an extraction module configured to extract audio features of the target video segment; a second determining module configured to determine the content division point from the target video segment based on the audio feature.

10. The interactive operation includes a playback selection operation, the playback behavior data further includes a continuous playback time length from the playback point, and the content division point includes an opening end point, and wherein the first determination unit further includes: determining an interactive operation whose continuous playback time length is greater than a first threshold as a valid interactive operation; The device according to claim 9 , wherein the device is configured to determine the number of valid interactive operations whose playback time points are located in the video segment as the number of interactions.

11. The second determination unit further comprises:

11. The device of claim 10, configured to determine a video segment with the highest number of interactions in a first time range as the target video segment, wherein the first time range is a time range from a start point of the video to a predetermined first time point.

12. The interactive operation includes a playback end operation, and the content division point includes an ending start point, where the second determination unit further comprises:

10. The device of claim 9, further configured to determine the video segment with the highest number of interactions in a second time range as the target video segment, wherein the second time range is a time range from a second predetermined point in time to an end point of the video.

13. The playback behavior data includes the number of times the video has been played, and wherein the first determination module further: The device of claim 9 , configured to, in response to the number of plays being greater than a second threshold, determine a target video segment in which a content division point of the video is located based on the play behavior data.

14. The extraction module: a transform unit configured to perform a Fourier transform on the audio data of the target video segment to obtain a frequency spectrum corresponding to the audio data; an extraction unit configured to perform feature extraction on the frequency spectrum to obtain the audio features.

15. The second determination module further comprises: The apparatus according to claim 9 , configured to determine the content segmentation points based on a mapping relationship between a preset audio feature and a content segmentation point.

16. The apparatus of claim 15 , wherein the mapping relationship is determined based on content segmentation point labels of a sample video and audio features of a sample target video segment in which the content segmentation point labels are located.

17. An electronic device, at least one processor; and a memory communicatively coupled to the at least one processor, wherein: An electronic device, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to perform the method of any one of claims 1 to 8.

18. A non-transitory computer readable storage medium having stored thereon computer instructions, the computer instructions causing a computer to perform the method of any one of claims 1 to 8.

19. A computer program comprising computer program instructions which, when executed by a processor, implement the method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video segmentation and annotation method and device

    CN111757170A

  • Interaction method and system and electronic equipment

    CN114125566A

  • Recording and reproducing apparatus, and recording and reproducing method

    JP2007066409A

  • Device, method and program for editing information

    JP2007067522A

  • Content viewing apparatus

    JP2007208651A