Video processing method and device, equipment and storage medium

By acquiring consumer data and analyzing long-form video content, personalized short videos are generated, solving the problems of low efficiency and high cost in existing technologies, and achieving more effective long-form video guidance and consumer engagement.

CN121585848APending Publication Date: 2026-02-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511775555.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies for generating short videos corresponding to long videos are inefficient and labor-intensive, resulting in unsatisfactory guidance effects and difficulty in effectively stimulating consumer interest in long videos.

Method used

By acquiring consumer media consumption data and analyzing long-form video content, key video clips that match consumer interests are generated, and personalized short videos are automatically generated to guide consumers to watch long-form videos.

Benefits of technology

It improved the guiding effect of short videos on long videos, reduced the manual cost of generating short videos, and increased the click-through rate, completion rate, and user conversion rate of long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585848A_ABST
    Figure CN121585848A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and device, equipment and a storage medium, which can be applied to various scenes such as search and recommendation. The method comprises the following steps: acquiring media consumption data of N first consumers and video data of M long videos; determining interest feature information of the N first consumers based on the media consumption data of the N first consumers; analyzing video data of each long video in the M long videos to obtain a key video clip of each long video; and for each long video in the M long videos, based on the interest feature information of the N first consumers, editing the key video clip of the long video to generate a short video corresponding to the long video, the short video being used for guiding a second consumer to consume the long video corresponding to the short video. The content of the generated short video is more in line with interests and preferences of different consumers, the guiding effect of the short video on the long video is improved, the labor cost for generating the short video is reduced, and the generation efficiency of the short video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a video processing method, apparatus, device, and storage medium. Background Technology

[0002] With the rapid development of mobile internet and short video platforms, consumers' content consumption habits have changed significantly. People generally prefer watching shorter, concise videos, which poses an increasing challenge for longer videos, such as TV dramas, movies, and documentaries, in capturing consumer attention. While longer videos offer rich content and complete stories, their viewing threshold is higher, making it difficult for consumers to find parts that truly interest them within a short timeframe.

[0003] To address the sluggish consumption of long-form videos, many video platforms and content creators are attempting to leverage short videos to drive viewership of longer videos—a strategy known as "short-to-long." Currently, the primary method for generating short videos corresponding to long-form videos involves manually selecting and editing key segments from the long video. However, this method is not ideal in terms of boosting the viewing experience of long-form videos, and it is also labor-intensive and inefficient. Summary of the Invention

[0004] This application provides a video processing method, apparatus, device, and storage medium that can improve the guiding effect of short videos on long videos, reduce the manual cost of short video generation, and improve the generation efficiency of short videos.

[0005] In a first aspect, this application provides a video processing method, including: Obtain the media consumption data of each of the N first consumers and the video data of each of the M long videos, where N and M are both positive integers; Based on the media consumption data of each of the N first consumers, determine the interest characteristics of each first consumer; Analyze the video data of each of the M long videos to obtain the key video segments of each long video; For each of the M long videos, based on the interest feature information of the N first consumers, the key video segments of the long video are edited to generate a short video corresponding to the long video. The short video is used to guide the second consumer to consume the long video corresponding to the short video.

[0006] Secondly, this application provides a video processing apparatus, comprising: The acquisition unit is used to acquire media consumption data of each of the N first consumers and video data of each of the M long videos, where N and M are both positive integers. An interest determination unit is used to determine the interest characteristic information of each first consumer based on the media consumption data of each of the N first consumers; The selection unit is used to analyze the video data of each of the M long videos to obtain the key video segments of each long video; The processing unit is configured to, for each of the M long videos, edit key video segments of the long video based on the interest feature information of the N first consumers to generate a short video corresponding to the long video, and the short video is used to guide the second consumers to consume the long video corresponding to the short video.

[0007] In some embodiments, the generation unit is specifically used to cluster the N first consumers based on the interest feature information of each of the N first consumers to obtain P consumer groups, where P is a positive integer less than or equal to N; for the j-th consumer group among the P consumer groups, based on the interest feature information of the j-th consumer group, the key video segments of the long video are edited to generate a short video corresponding to the j-th consumer group, where j is a positive integer from 1 to P.

[0008] In some embodiments, the generation unit is specifically configured to, based on the interest feature information of the j-th consumer group, select at least one key video segment that the j-th consumer group is interested in from the key video segments of the long video; and, based on the interest feature information of the j-th consumer group, edit the at least one key video segment to generate a short video corresponding to the j-th consumer group.

[0009] In some embodiments, the generation unit is specifically configured to acquire content feature information of each key video segment in the at least one key video segment; determine short video generation strategy information for the j-th consumer group based on the interest feature information of the j-th consumer group and the content feature information of each key video segment in the at least one key video segment; and edit the at least one key video segment based on the short video generation strategy information of the j-th consumer group to generate a short video corresponding to the j-th consumer group.

[0010] In some embodiments, the generation unit is specifically used to edit the at least one key video segment based on the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment, to generate a short video corresponding to the j-th consumer group.

[0011] In some embodiments, the short video generation strategy information includes at least one of the following: the hook type of the short video, and the style requirements of the short video. The style requirements of the short video include at least one of the following: the first frame requirements of the short video, the editing requirements of the short video, the ending requirements of the short video, the special effects requirements, the background music requirements, the narration requirements, the style requirements, and the emotional rendering requirements.

[0012] In some embodiments, the acquisition unit is further configured to acquire consumption characteristic information of the second consumer, the consumption characteristic information of the second consumer including at least one of the following: interest characteristic information of the second consumer, media consumption data of the second consumer in a historical time period, and media consumption data of the second consumer in the current time period; the selection unit is further configured to select K short videos from the short videos corresponding to the M long videos based on the consumption characteristic information of the second consumer, where K is a positive integer; the processing unit is further configured to generate a short video recommendation list for the second consumer based on the K short videos.

[0013] In some embodiments, before acquiring the consumption characteristic information of the second consumer, the acquisition unit is further configured to acquire consumption data of multiple consumers on multiple videos, the multiple videos including long videos and short videos; based on the consumption data of the multiple consumers on the multiple videos, construct a heterogeneous graph of consumers and videos, the heterogeneous graph representing the consumption information of each of the multiple consumers on each of the multiple videos, and the association information between the multiple videos; based on the identification information of the second consumer, acquire the consumption characteristic information of the second consumer from the heterogeneous graph.

[0014] In some embodiments, the selection unit is specifically used to select Q short videos from the short videos corresponding to the M long videos based on the consumption characteristic information of the second consumer, where Q is less than M; determine the rating of each of the Q short videos based on the consumption characteristic information of the second consumer; and select K short videos from the Q short videos based on the rating.

[0015] In some embodiments, the selection unit is specifically used to, for the i-th short video among the Q short videos, obtain the content feature information of the i-th short video and the content feature information of the long video corresponding to the i-th short video; and determine the rating of the i-th short video based on the consumption feature information of the second consumer, the content feature information of the i-th short video, and the content feature information of the long video corresponding to the i-th short video.

[0016] In some embodiments, the selection unit is specifically used to construct state information based on the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video; and to determine the rating of the i-th short video under the state information through the strategy model.

[0017] In some embodiments, the processing unit is further configured to, after recommending the short video recommendation list to the second consumer, collect feedback information from the second consumer, the feedback information including the second consumer's consumption information of short videos in the short video recommendation list and consumption information of long videos corresponding to the short videos in the short video recommendation list; determine the reward of the strategy model based on the feedback information; and adjust the parameters in the strategy model based on the reward.

[0018] In some embodiments, the processing unit is specifically configured to determine the reward of the strategy model based on at least one of the following: the click-through rate of the short videos in the short video recommendation list by the second consumer, the playback completion rate of the short videos, the interaction rate of the short videos, and the click-through rate of the long videos corresponding to the short videos in the short video recommendation list, the playback completion rate of the long videos, the subscription rate of the related content of the long videos, and the paid conversion rate of the long videos.

[0019] In some embodiments, the processing unit is further configured to, in response to a triggering operation by the second consumer on a jump option in the playback interface of the first short video, obtain content feature information of the first short video and narrative structure information of the first long video corresponding to the first short video, wherein the first short video is a short video currently being played by the second consumer in the short video recommendation list; determine the jump video frame of the first long video based on the consumption feature information of the second consumer, the content feature information of the first short video, and the narrative structure information of the first long video; and jump to the jump video frame in the first long video based on the jump video frame of the first long video.

[0020] In some embodiments, the consumption feature information includes the narrative curiosity feature of the second consumer. The processing unit is specifically used to determine the first jump mode corresponding to the second consumer from multiple candidate jump modes based on the narrative curiosity feature of the second consumer and the content feature information of the first short video. The multiple candidate jump modes include an immediate plot continuation mode and a background plot supplement mode. Based on the first jump mode and the narrative structure information of the first long video, the processing unit determines the jump video frame of the first long video.

[0021] In some embodiments, the consumption feature information includes the narrative curiosity feature of the second consumer. The processing unit is further configured to generate guidance information based on the narrative curiosity feature of the second consumer and the video data of the first long video when the second consumer consumes the first long video. The guidance information is used to guide the second consumer to consume all or a specific part of the content of the first long video.

[0022] In some embodiments, the processing unit is specifically configured to: perform content parsing on the video data of the j-th long video among the M long videos to obtain multimodal information of the j-th long video, wherein the multimodal information includes at least one of visual information, audio information, and text information, and j is a positive integer from 1 to M; determine the semantic sentiment feature information of the j-th long video based on the multimodal information of the j-th long video; perform narrative structure analysis on the video data of the j-th long video to obtain the narrative structure information of the j-th long video; and select multiple key video segments from the j-th long video based on the semantic sentiment feature information and the narrative structure information of the j-th long video.

[0023] Thirdly, this application provides a computing device including a processor and a memory.

[0024] The memory is used to store computer programs; The processor is used to call and run the computer program stored in the memory to perform the method described in the first aspect.

[0025] Fourthly, a computer-readable storage medium is provided for storing a computer program, which is loaded and executed by a computing device to implement the method of the first aspect described above.

[0026] Fifthly, a computer program product is provided, including a computing computer program stored in a readable storage medium, which a computing device can read from the readable storage medium, causing the computing device to load and execute the computer program to implement the method of the first aspect described above.

[0027] In a sixth aspect, a computer program is provided that, when run on a computing device, causes the computing device to perform the method described in the first aspect.

[0028] In summary, this application obtains media consumption data from N first consumers and video data from M long videos, where N and M are both positive integers. Based on the media consumption data of the N first consumers, the interest characteristics of the N first consumers are determined. Simultaneously, the video data of each of the M long videos is analyzed to obtain key video segments for each long video. Thus, for each of the M long videos, based on the interest characteristics of the N first consumers, the key video segments are edited to generate a corresponding short video. This short video is used to guide second consumers to consume the corresponding long video. Therefore, to improve the guiding effect of short videos on long videos, the interest characteristics of different consumers are considered when generating the short videos corresponding to the long videos, making the content of the generated short videos more in line with the interests and preferences of different consumers. Generating personalized short videos can effectively stimulate audience interest, thereby guiding consumers to consume (e.g., watch) the complete long video, thus improving the click-through rate, completion rate, and user conversion rate of the long videos. Furthermore, this application embodiment generates short videos corresponding to long videos based on key video segments of long videos, which can further enhance the attractiveness of the generated short videos. Moreover, the entire short video generation process does not require manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the efficiency of short video generation. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of an implementation environment of an embodiment of this application; Figure 2 A schematic flowchart illustrating a video processing method provided in an embodiment of this application; Figure 3 A diagram illustrating the generation of a short video; Figure 4 Another illustration for generating short videos; Figure 5 Another illustration for generating short videos; Figure 6 A schematic flowchart illustrating a video processing method provided in an embodiment of this application; Figure 7A This is a schematic diagram for selecting K short videos; Figure 7B This is another illustration of selecting K short videos; Figure 8 This is a diagram illustrating the transition from a short video playback interface to a long video playback interface. Figure 9 A schematic diagram for generating guiding information; Figure 10 A schematic flowchart illustrating a video processing method provided in an embodiment of this application; Figure 11 A schematic diagram illustrating the framework of the video processing method provided in the embodiments of this application; Figure 12 This is a schematic block diagram of a video processing apparatus provided in an embodiment of this application; Figure 13 This is a schematic block diagram of a computing device provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. In embodiments of the invention, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0033] The video processing method provided in this application can be applied to various fields such as video promotion and video playback. When generating short videos corresponding to long videos, it considers the interest feature information of N first consumers, making the generated short videos more in line with the interests and preferences of each first consumer, thereby improving the guiding effect of short videos on long videos. Moreover, the entire short video generation process does not require manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the generation efficiency of short videos.

[0034] With the rapid development of mobile internet and short video platforms, consumers' content consumption habits have changed significantly. People generally prefer watching shorter, concise videos, which poses an increasing challenge for longer videos, such as TV dramas, movies, and documentaries, in capturing consumer attention. While longer videos offer rich content and complete stories, their viewing threshold is higher, making it difficult for consumers to find parts that truly interest them within a short timeframe.

[0035] To address the sluggish consumption of long-form videos, many video platforms and content creators are attempting to leverage short videos to drive viewership of longer videos—a strategy known as "short-to-long." Currently, the primary method for generating short videos corresponding to long-form videos involves manually selecting and editing key segments from the long video. However, this method is not ideal in terms of boosting the viewing experience of long-form videos, and it is also labor-intensive and inefficient.

[0036] To address the aforementioned technical problems, this application proposes a novel video processing method. This method involves acquiring media consumption data from N first consumers and video data from M long videos, where N and M are both positive integers. Based on the media consumption data of the N first consumers, the interest characteristics of these consumers are determined. Simultaneously, the video data of each of the M long videos is analyzed to obtain key video segments for each long video. For each of the M long videos, based on the interest characteristics of the N first consumers, the key video segments are edited to generate a corresponding short video. This short video is used to guide second consumers to consume the corresponding long video. Therefore, to enhance the guiding effect of the short video on the long video, the generation of the short video considers the interest characteristics of different consumers, making the content of the generated short video more aligned with the interests and preferences of different consumers. This personalized short video effectively stimulates viewer interest, thereby guiding consumers to consume (watch) the complete long video, thus improving the click-through rate, completion rate, and user conversion rate of the long video. Furthermore, this application embodiment generates short videos corresponding to long videos based on key video segments of long videos, which can further enhance the attractiveness of the generated short videos. Moreover, the entire short video generation process does not require manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the efficiency of short video generation.

[0037] The implementation environment of the embodiments of this application is described below.

[0038] Figure 1 This is a schematic diagram of an implementation environment of an embodiment of this application, such as... Figure 1 As shown, the implementation environment includes: terminal device 101 and server 102.

[0039] The terminal device 101 and the server 102 are connected via wired or wireless means. The terminal device 101 and the server 102 constitute a video recommendation platform. In this embodiment, the terminal device 101 is equipped with a client for the video recommendation platform, and the server 102 can be understood as the server-side or backend of the video recommendation platform. Consumers (e.g., users, or objects) can interact with the client of the video recommendation platform installed on the terminal device 101.

[0040] In some embodiments, the video processing method of this application is performed by server 102. For example, server 102 acquires media consumption data of N first consumers (also referred to as first objects) and video data of M long videos, where N and M are both positive integers. Based on the media consumption data of the N first consumers, server 102 determines the interest feature information of the N first consumers. Simultaneously, server 102 analyzes the video data of each of the M long videos to obtain key video segments for each long video. Thus, for each of the M long videos, server 102 edits the key video segments of the long video based on the interest feature information of the N first consumers to generate a short video corresponding to the long video. The short video is used to guide a second consumer (also referred to as a second object, which may be the same consumer as the first consumer or a different consumer) to consume the long video corresponding to the short video. Therefore, to enhance the guiding effect of short videos on long videos, this application's embodiments consider the interest characteristics of different consumers when generating short videos corresponding to long videos. This makes the content of the generated short videos more in line with the interests and preferences of different consumers, generating personalized short videos that can effectively stimulate viewers' interest and guide them to consume (watch) the complete long video, thereby increasing the click-through rate, completion rate, and user conversion rate of the long video. Furthermore, this application's embodiments generate short videos based on key video segments of the long video, which can further enhance the attractiveness of the generated short videos. Moreover, the entire short video generation process requires no manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the efficiency of short video generation.

[0041] In some embodiments, the implementation environment also includes a database 103, which is communicatively connected to a server 102. The server 102 can store the short videos corresponding to the M long videos generated above in the database 103.

[0042] In some embodiments, when video recommendations need to be made to a second consumer, server 102 can read short videos corresponding to M long videos from database 103. Then, server 102 selects K short videos from the short videos corresponding to the M long videos and sends them to the terminal device 101 corresponding to the second consumer. Terminal device 101 displays the K short videos to the second consumer through the installed video recommendation platform client.

[0043] In some embodiments, if a second consumer triggers a jump option on the first short video playback interface while watching the first short video out of K short videos, the terminal device 101 responds to the second consumer's trigger operation on the jump option on the first short video playback interface by jumping to and displaying the first long video corresponding to the first short video.

[0044] In some embodiments, the video processing method of this application can be performed by the terminal device 101 described above.

[0045] In some embodiments, the video processing method of this application can be jointly performed by terminal device 101 and server 102.

[0046] In some embodiments, the terminal device 101 includes, but is not limited to, desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices may include smartwatches, smart bracelets, and head-mounted devices. Terminal devices are often equipped with a display device, which may be a monitor, display screen, touchscreen, etc., and the touchscreen may be a touchscreen, touch panel, etc.

[0047] In some embodiments, the server 102 may be one or more servers. When there are multiple servers, at least two servers may be used to provide different services, and / or at least two servers may be used to provide the same service, such as providing the same service in a load-balanced manner. This application embodiment does not limit this. The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may also be a node in a blockchain.

[0048] It should be noted that the implementation environment of this application embodiment includes, but is not limited to, Figure 1 As shown.

[0049] The technical solutions of the embodiments of this application will be described in detail below through some examples. The following embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0050] Figure 2 This is a schematic flowchart illustrating a video processing method provided in an embodiment of this application. The method in this embodiment can be executed by a computing device, which may be... Figure 1 The server 102 or terminal device 101 shown.

[0051] like Figure 2 As shown, the video processing method of this application embodiment includes: S101. Obtain the media consumption data of each of the N first consumers and the video data of each of the M long videos.

[0052] Where N and M are both positive integers.

[0053] It should be noted that the data used in executing the embodiments of this application and the process of obtaining this data comply with the relevant laws and regulations.

[0054] In this embodiment, to improve the efficiency of the generated short video in guiding the corresponding long video, the interest characteristics of different consumers are considered when generating the short video corresponding to the long video. Based on this, when generating the short video, the computing device first obtains the media consumption data of N consumers. For ease of description, these N consumers are referred to as N first consumers.

[0055] Among them, these N first consumers are N different consumers, and the media consumption data of these N first consumers mainly includes the media consumption data of each of these N first consumers within a preset time period.

[0056] In this embodiment, the media consumption data of the first consumer mainly includes the consumer's consumption behavior data on short videos and long videos within a preset time period. For example, this consumption behavior data includes, but is not limited to: browsing, clicking, viewing time, likes, comments, sharing, searching, favorites, following, and payment behavior.

[0057] In the embodiments of this application, the media consumption data of the aforementioned N first consumers can be collected by the computing device from a single video playback platform, or the computing device can collect data from multiple video playback platforms.

[0058] The media consumption data of the first consumer can reveal their interests and preferences regarding video content. For example, it can show which types of videos, performers, or styles of videos the first consumer is interested in. In this way, the computing device can use the media consumption data of N first consumers to generate short videos that match the interests and preferences of different consumers. This personalized recommendation of short videos enhances the guiding effect of the generated short videos on longer videos.

[0059] This application does not impose specific limitations on the above-mentioned M long videos; they can be a single long video or multiple long videos.

[0060] In some embodiments, the aforementioned M long videos may be all long videos included in one or more video playback platforms.

[0061] In some embodiments, the M long videos mentioned above may be long videos that have recently been launched on one or more video playback platforms.

[0062] In some embodiments, the aforementioned M long videos may be long videos that one or more video playback platforms need to promote in the near future.

[0063] In this embodiment of the application, the computing device can obtain video data of M long videos from a video playback platform.

[0064] In this embodiment of the application, after the computing device obtains the media consumption data of N first consumers and the video data of M long videos through the above method, it executes the following step S102.

[0065] S102. Based on the media consumption data of each of the N first consumers, determine the interest characteristics of each first consumer.

[0066] In the embodiments of this application, the specific methods for determining the interest feature information of each of the N first consumers are basically the same. For ease of description, we will take determining the interest feature information of one of the N first consumers as an example, for example, determining the interest feature information of the i-th first consumer.

[0067] In this embodiment of the application, for each of the N first consumers, such as the i-th first consumer, the computing device analyzes the media consumption data of the i-th first consumer obtained above to obtain the interest feature information of the i-th first consumer.

[0068] In some embodiments, the interest feature information of the i-th first consumer includes at least one of the following: the i-th first consumer's multi-dimensional interest tags, the i-th first consumer's viewing intention for long videos, and the i-th first consumer's narrative curiosity features.

[0069] As described above, the media consumption data of the i-th first consumer obtained in this application embodiment includes various behaviors of the i-th first consumer on short videos and long videos, including but not limited to: browsing, clicking, viewing time, liking, commenting, sharing, searching, collecting, following, and paying behavior. The computing device analyzes the media consumption data of the i-th first consumer and constructs multi-level interest tags for the i-th first consumer. For example, it establishes interest tags for the i-th first consumer on different content types (e.g., dramas, science fiction, comedies, etc.), themes (e.g., historical themes, modern themes, family themes, campus themes, etc.), styles (e.g., realism, classicism, minimalism, cartoon style, documentary style, etc.), directors, actors, etc., and calculates the intensity and timeliness of the interests.

[0070] For example, the computing device analyzes the media consumption data of the i-th consumer and determines that among the short and long videos watched by the i-th consumer, 70% are science fiction, 25% are comedy, and 5% are drama. Therefore, the i-th consumer is interested in science fiction, comedy, and drama, with corresponding interest intensities of 0.7, 0.25, and 0.05, respectively. Simultaneously, it determines that among the short and long videos watched by the i-th consumer, 80% are contemporary themes, 15% are historical themes, and 5% are family themes. Therefore, the i-th consumer is interested in contemporary themes, historical themes, and family themes, with corresponding interest intensities of 0.8, 0.15, and 0.05, respectively. Furthermore, it was determined that among the short and long videos watched by the i-th first consumer, 75% were documentary style, 15% were realist style, and 10% were minimalist style. Therefore, the video styles the i-th first consumer is interested in are documentary, realist, and minimalist, with corresponding interest intensities of 0.75, 0.15, and 0.1, respectively. Additionally, it was determined that 85% of the short and long videos watched by the i-th first consumer featured actor A, and 15% featured actor B. Therefore, the actors the i-th first consumer is interested in are actor A and actor B, with corresponding interest intensities of 0.85 and 0.15, respectively. Finally, it was determined that 85% of the short and long videos watched by the i-th first consumer were directed by director C, and 15% were directed by actor D. Therefore, the directors the i-th first consumer is interested in are director C and actor D, with corresponding interest intensities of 0.85 and 0.15, respectively. In summary, the multi-dimensional interest tags of the i-th first consumer are shown in Table 1. Table 1

[0071] It should be noted that Table 1 above is just an example, and the multi-dimensional interest tags of the i-th first consumer include, but are not limited to, those shown in Table 1 above.

[0072] In this embodiment, the computing device further analyzes the media consumption data of the i-th first consumer. Based on the i-th first consumer's historical viewing and interaction data, and feedback on short videos corresponding to long videos, it predicts the i-th first consumer's viewing intention and payment tendency for specific long videos, long video series, or long videos of a specific type. For example, after analyzing the media consumption data of the i-th first consumer, it is found that the i-th first consumer has a clear intention to watch long videos of the science fiction genre. The consumer will likely jump to the corresponding long video through short videos of the science fiction genre and will pay to watch long videos of the science fiction genre. Therefore, the computing device can predict the i-th first consumer's long video viewing intention, for example, predicting that the i-th first consumer has a clear intention to watch long videos of the science fiction genre.

[0073] In this embodiment, to further enhance the guiding effect of the generated short videos on the long videos, the narrative curiosity characteristic of the first consumer is also determined when determining the interest characteristic information of the first consumer. Taking the i-th first consumer as an example, the computing device analyzes whether the i-th first consumer shows a desire to explore the subsequent plot development, character fate, and story background after watching the short video corresponding to a specific long video. For example, after watching a suspenseful short video, will the i-th first consumer search for relevant information, click on the jump entrance to the long video, or express curiosity about the ending in the comment section? Through these behaviors, the narrative curiosity characteristic of the i-th first consumer is quantified and used as an important interest characteristic of the i-th first consumer. For example, the narrative curiosity characteristic of the i-th first consumer can be represented by 1 or 0, where 1 indicates that the i-th first consumer has narrative curiosity, and 0 indicates that the i-th first consumer does not have narrative curiosity.

[0074] The above describes the process by which a computing device determines the interest feature information of the i-th first consumer among N first consumers. By referring to this method, the computing device can determine the interest feature information of each of the N first consumers.

[0075] S103. Analyze the video data of each of the M long videos to obtain the key video segments of each long video.

[0076] It should be noted that there is no specific order in which S103 and S102 are executed. That is, S103 can be executed after S102, before S102, or simultaneously with S102. This application embodiment does not impose any restrictions on this.

[0077] In this embodiment of the application, when the computing device generates one or more short videos for each of the M long videos, it analyzes the video data of each of the M long videos to obtain key video segments for each long video. The key video segments of the long videos are used to generate the corresponding short videos. Therefore, it can be seen that the accuracy of the selection of key video segments of long videos is directly related to the generation effect of the corresponding short videos.

[0078] In this application embodiment, the specific process of the computing device determining the key video segment of each of the M long videos is the same. For ease of description, the process of determining the key video segment of one of the M long videos is used as an example, such as determining the key video segment of the j-th long video among the M long videos.

[0079] In this embodiment, the key video segments of the j-th long video mainly include segments in the j-th long video that can create strong suspense, arouse curiosity, generate emotional resonance, showcase a unique style or humor, and quickly grab consumers' attention as the opening of a short video. In other words, these key video segments are segments in the j-th long video that have a "narrative hook," that is, segments that attract consumers to consume the long video.

[0080] This application embodiment does not limit the specific method by which a computing device analyzes the video data of the j-th long video to obtain the key video segments of the j-th long video.

[0081] In some embodiments, a key segment extraction model can be trained to extract key video segments from the j-th long video. For example, during model training, a manually labeled dataset is prepared, containing multiple data pairs, each pair including a long video and a key video segment of that long video. The long video from the data pair is used as input to the model, and the key video segment from the long video in the data pair is used as the label to train the model, enabling the trained model to identify the key video segments of the input long video. Thus, the computing device can input the video data of the j-th long video into the trained key segment extraction model for key video segment identification and output the key video segments of the j-th long video.

[0082] In some embodiments, the computing device determines the key video segment of the j-th long video through the following steps S103-A to S103-D (not shown in the figure): S103-A. For the j-th long video among M long videos, perform content parsing on the video data of the j-th long video to obtain the multimodal information of the j-th long video. The multimodal information includes at least one of visual information, audio information and text information, where j is a positive integer from 1 to M. S103-B: Based on the multimodal information of the j-th long video, determine the semantic sentiment feature information of the j-th long video; S103-C, Perform narrative structure analysis on the video data of the j-th long video to obtain the narrative structure information of the j-th long video; S103-D: Based on the semantic and emotional features and narrative structure information of the j-th long video, select multiple key video segments from the j-th long video.

[0083] In this implementation, when determining the key video segment of the j-th long video, the computing device first performs content parsing on the video data of the j-th long video to obtain the multimodal information of the j-th long video. The multimodal information includes at least one of visual information, audio information, and text information.

[0084] In one example, if the multimodal information of the j-th long video includes its visual information, the computing device classifies the j-th long video into an image sequence, analyzes each image in the sequence, and extracts visual features (such as composition, color, motion trajectory, and character recognition). Specifically, the computing device can extract feature vectors of the scene, objects, and faces in the j-th long video using a feature extraction module (e.g., ResNet, EfficientNet). The computing device uses a motion analysis module (e.g., 3D-Convolutional Neural Networks, 3D-CNN, or a transformer-based model) to understand the actions and motion patterns within the image sequence of the j-th long video. The computing device uses a recognition module (e.g., a face recognition module) to detect the presence of the main characters in the entire j-th long video and construct character trajectories. In this way, the computing device can obtain the visual information of the j-th long video, which mainly includes visual feature sequences, face label sequences, scene transition time points, motion intensity curves, etc.

[0085] In one example, if the multimodal information of the j-th long video includes its audio information, the computing device extracts audio information such as background music, sound effects, and human dialogue from the j-th long video. For example, the computing device uses a high-precision ASR engine to convert human voice into timestamped text. Another example is that the computing device identifies non-speech sounds in the j-th long video, such as laughter, applause, explosions, and tense background music. Yet another example is that the computing device extracts features such as volume, pitch, speech rate, and emotional tone from the j-th long video. Finally, this information is aggregated to obtain the audio information of the j-th long video.

[0086] In one example, if the multimodal information of the j-th long video includes the text information of the j-th long video, then the computing device extracts the metadata such as the video title, description, and subtitles of the j-th long video as the text information of the j-th long video.

[0087] Next, the computing device determines the semantic and sentiment features of the j-th long video based on its multimodal information. These semantic and sentiment features can include both semantic and sentiment characteristics of the j-th long video. For example, the computing device performs semantic and sentiment analysis on the multimodal information of the j-th long video to understand its core theme, plot development, and key events, thus obtaining its semantic features. It also analyzes the emotional fluctuations in different segments of the video to identify moments of intense emotion such as climaxes, plot twists, conflicts, and moments of warmth, thus obtaining its sentiment characteristics.

[0088] In one example, the computing device analyzes metadata such as video title, description, and subtitles in the text information of the j-th long video to obtain the core theme, plot development, and key events of the j-th long video.

[0089] In one example, the computing device analyzes the visual, audio, and textual information of the j-th long video to identify emotionally intense moments such as climaxes, plot twists, conflicts, and moments of warmth. For instance, it might determine emotional changes in characters by recognizing facial expressions and analyzing color saturation / contrast (e.g., brightness or darkness) in the visual information. Or, it might determine emotional changes by analyzing the type of background music (exhilarating or soothing), tone of voice (excitement or sadness), and volume in the audio information. Or, it might perform sentiment analysis on the text to determine emotional changes. In this way, the computing device can obtain a quantified emotional intensity curve spanning the entire j-th long video, marking emotional peaks (climax, plot twist, conflict, warmth).

[0090] In this embodiment of the application, in order to improve the accurate selection of key video segments in the j-th long video, in addition to determining the semantic and emotional feature information of the j-th long video, the video data of the j-th long video is also subjected to narrative structure analysis to obtain the narrative structure information of the j-th long video.

[0091] This application embodiment does not limit the specific method by which a computing device performs narrative structure analysis on the video data of the j-th long video to obtain the narrative structure information of the j-th long video.

[0092] In one example, the computing device inputs the video data of the j-th long video frame by frame into a deep learning model (such as a Transformer-based sequence model). This deep learning model identifies the temporal boundaries of structural plot points such as the beginning, development, climax, twist, and ending in the j-th long video. Specifically, for example, the computing device writes a prompt to instruct the deep learning model to identify these structural plot points in the input j-th long video. The computing device then inputs the prompt and the j-th long video into the deep learning model, which outputs a narrative structure graph of the j-th long video, clearly marking the temporal boundaries of these structural plot points.

[0093] In this embodiment, the narrative structure information of the j-th long video also includes the personality changes, emotional development, and character relationship network of the main characters in the j-th long video, as well as segments that highlight the charm or conflict of the main characters (i.e., highlight segments of the main characters). Specifically, the computing device can use facial recognition results from the visual information of the j-th long video to count the appearance time of each character in the j-th long video. Then, based on the appearance time and importance of each character in the structural plot points, it can calculate the "protagonist index" of each character, thereby obtaining the list of main characters in the j-th long video. The character relationship network is obtained through the frequency of characters appearing together and dialogue information (where dialogue information can be obtained from the text information of the j-th long video). The segments in the j-th video segment that highlight the charm or conflict of the main characters are obtained by analyzing the performance of the main characters in the structural plot points.

[0094] In this implementation, after the computing device determines the semantic sentiment feature information and narrative structure information of the j-th long video based on the above steps, it executes the above steps S103-D to select multiple key video segments from the j-th long video based on the semantic sentiment feature information and narrative structure information of the j-th long video.

[0095] In one possible implementation, the computing device selects segments with strong emotional impact, such as climaxes, plot twists, conflicts, and heartwarming moments, from the j-th long video based on its semantic and emotional features, and identifies these segments as key video segments of the j-th long video. Furthermore, based on the narrative structure information of the j-th long video, the computing device selects segments containing structural plot points such as the beginning, development, climax, plot twist, and ending, as well as segments from the j-th video segment that highlight the charm or conflict of the main characters, and identifies these segments as key video segments of the j-th long video.

[0096] In one possible implementation, the computing device divides the j-th long video into multiple candidate segments, for example, taking a 5-10 second segment from the j-th long video as a candidate segment. For each of these candidate segments, a score is determined based on the semantic sentiment features and narrative structure information of the j-th long video. Then, based on the scores, the highest-scoring candidate segments are selected from these candidate segments and identified as the key video segments of the j-th long video.

[0097] In this embodiment, the key video segment is a segment from the j-th long video that can create strong suspense, arouse curiosity, generate emotional resonance, showcase a unique style or humor, and quickly grab the user's attention as the opening of a short video. Based on this, the computing device determines the score of each candidate segment based on the semantic and emotional features and narrative structure information of the j-th long video in the following way: For each candidate segment, such as candidate segment 1, the computing device determines whether candidate segment 1 is located at a structural plot point in the j-th long video based on the narrative structure information, and whether the dialogue in candidate segment 1 contains interrogative sentences or unsolved mysteries, thereby determining the suspense and curiosity score of candidate segment 1. Simultaneously, the computing device determines the score of candidate segment 1 on the emotional intensity curve based on the semantic and emotional features of the j-th long video, obtaining the emotional resonance score of candidate segment 1. Furthermore, the computing device determines whether candidate segment 1 contains high-intensity motion, close-up shots, stunning visual effects, or unique compositions; whether it contains sudden silence, loud sound effects, iconic lines, or rousing music; whether it includes a main character or showcases their iconic personality traits / highlight moments; and thus, determines the unique style or humor of the candidate segment. Optionally, the computing device also determines whether candidate segment 1 can function as a relatively independent narrative unit, allowing consumers who have not seen the j-th long video to understand the basic context (avoiding abrupt transitions). Optionally, the computing device can also determine whether candidate segment 1, as the beginning of a short video, possesses visual and audio features that instantly grab attention (e.g., sudden visual changes, memorable lines, loud sound effects). Through the above comprehensive analysis, the computing device obtains a comprehensive score for candidate segment 1, which is then used as the score for candidate segment 1.

[0098] Using the method described above, the computing device can determine the score of each candidate segment in the j-th long video, and then identify the candidate segments with the highest scores as the key video segments of the j-th long video.

[0099] S104. For each of the M long videos, based on the interest feature information of N first consumers, edit the key video segments of the long video to generate the corresponding short video.

[0100] Among them, short videos are used to guide second consumers to consume (e.g., watch) the corresponding long videos.

[0101] In this embodiment, the computing device, based on the above steps, determines the interest feature information of each of the N first consumers and the key video segments of each of the M long videos. Then, based on the interest feature information of the N first consumers, the key video segments of each of the M long videos are edited to generate at least one short video corresponding to each long video.

[0102] In the embodiments of this application, the specific methods by which the computing device determines the short video corresponding to each of the M long videos are basically the same. For ease of description, the method of determining the short video corresponding to a long video will be used as an example for explanation.

[0103] In this embodiment, the computing device edits key video segments of the long video based on the interest feature information of N first consumers to generate one or more short videos corresponding to the long video. For example, the computing device can edit key video segments of the long video based on the interest feature information of each of the N first consumers to generate a short video corresponding to each of the N first consumers. That is, in this embodiment, the generated short videos corresponding to the long video include N short videos, and these N short videos correspond one-to-one with the N first consumers.

[0104] In some embodiments, the computing device may determine the short video corresponding to the long video through the following steps S104-A and S104-B (not shown in the figure): S104-A: Based on the interest feature information of each of the N first consumers, cluster the N first consumers to obtain P consumer groups, where P is a positive integer less than or equal to N; S104-B: For the j-th consumer group among P consumer groups, based on the interest feature information of the j-th consumer group, the key video segments of the long video are edited to generate the short video corresponding to the j-th consumer group, where j is a positive integer from 1 to P.

[0105] In this implementation, when the computing device generates a short video corresponding to the long video based on the interest feature information of N first consumers, it first clusters these N first consumers based on the interest feature information of each of the N first consumers. For example, it aggregates the first consumers with the same or similar interest feature information into a group. In this way, the N first consumers can be aggregated into P consumer groups. Usually, P is less than N. Under extreme conditions, P can be equal to N, that is, the interest feature information of the N first consumers is different.

[0106] Thus, for each of the P consumer groups, such as the j-th consumer group, the computing device can edit key video segments of the long video based on the interest feature information of the j-th consumer group to generate a short video corresponding to that consumer group. For example, the computing device inputs the interest feature information of the j-th consumer group into the short video generation model, which then edits the key video segments based on the interest feature information of the j-th consumer group to generate a short video corresponding to that consumer group.

[0107] In some embodiments, the computing device may determine the short video corresponding to the long video through the following steps S104-B1 and S104-B2 (not shown in the figure): S104-B1. Based on the interest feature information of the j-th consumer group, select at least one key video segment that the j-th consumer group is interested in from the key video segments of the long video, where j is a positive integer from 1 to P. S104-B2. Based on the interest characteristics of the j-th consumer group, edit at least one key video segment to generate a short video corresponding to the j-th consumer group.

[0108] Since the number of key video segments selected in the above steps may be large, while the duration of a short video is short, in order to improve the efficiency of generating short videos and reduce the amount of data that needs to be processed during the short video generation process, in this implementation, the computing device does not edit all the key video segments of the long video. Instead, based on the interest feature information of the j-th consumer group, it selects at least one key video segment that the j-th consumer group is interested in from the key video segments of the long video. For example, for each key video segment of the long video, the computing device obtains the content feature information of the key video segment (such as multimodal information such as visual information, audio messages, and text information of the video clip), and determines the degree of interest of the j-th consumer group in the key video segment based on the interest feature information of the j-th consumer group and the content feature information of the key video segment. One way to determine the degree of interest is to convert the interest feature information of the j-th consumer group into vector representation 1, convert the content feature information of the key video segment into vector representation 2, calculate the distance between vector representation 1 and vector representation 2, and determine the degree of interest of the j-th consumer group in the key video segment based on this distance. For example, the smaller the distance, the greater the degree of interest of the j-th consumer group in the key video segment. Following this method, the computing device can determine the degree of interest of the j-th consumer group in each key video segment of the long video, and then, based on the degree of interest, select at least one key video segment that the j-th consumer group is interested in from the key video segments of the long video.

[0109] Next, based on the interest characteristics of the j-th consumer group, the computing device edits at least one key video segment that the j-th consumer group is interested in, and generates a short video corresponding to the j-th consumer group.

[0110] This application embodiment does not limit the specific method by which a computing device edits at least one key video segment that the j-th consumer group is interested in based on the interest feature information of the j-th consumer group to generate a short video corresponding to the j-th consumer group.

[0111] In one possible implementation, such as Figure 3 As shown, the computing device inputs the interest feature information of the j-th consumer group and at least one key video segment that the j-th consumer group is interested in into the short video generation model. Based on the interest feature information of the j-th consumer group, the short video generation model edits and processes at least one key video segment that the j-th consumer group is interested in to generate a short video corresponding to the j-th consumer group.

[0112] In one possible implementation, the computing device can generate the short video corresponding to the long video through the following steps S104-B21 to S104-B23 (not shown in the figure): S104-B21. Obtain the content feature information of each key video segment in at least one key video segment; S104-B22. Based on the interest feature information of the j-th consumer group and the content feature information of each key video segment in at least one key video segment, determine the short video generation strategy information of the j-th consumer group. S104-B23. Based on the short video generation strategy information of the j-th consumer group, edit at least one key video segment to generate a short video corresponding to the j-th consumer group.

[0113] In this embodiment, the content feature information of the key video segment includes at least one of the following: multimodal information, semantic sentiment feature information, and narrative structure information. The multimodal information of the key video segment can be obtained from the multimodal information of the long video, and the semantic sentiment feature information and narrative structure information of the key video segment can be obtained from the semantic sentiment feature information and narrative structure information of the long video. The specific process can be referred to the detailed description of the above embodiments, and will not be repeated here. Thus, the computing device can determine the short video generation strategy information for the j-th consumer group based on the interest feature information of the j-th consumer group and the content feature information of each key video segment in at least one key video segment.

[0114] In one example, such as Figure 4 As shown in the embodiment of this application, a short video generation strategy generation model is trained. The computing device can input the interest feature information of the j-th consumer group and the content feature information of each key video segment in at least one key video segment into the short video generation strategy generation model. The short video generation strategy generation model analyzes the interest feature information of the j-th consumer group and each key video segment in at least one key video segment to generate short video generation strategy information for the j-th consumer group.

[0115] The short video generation strategy information of the j-th consumer group is used to guide the short video generation model to generate which type and style of short video.

[0116] In one example, the short video generation strategy information generated by this application embodiment includes at least one of the following: the hook type of the short video and the style requirements of the short video.

[0117] The hook types in short videos include: suspense, emotional impact, and plot twists.

[0118] The requirements for short video style include at least one of the following: requirements for the first frame of the short video, requirements for the editing of the short video, requirements for the ending of the short video, requirements for special effects, requirements for background music, requirements for narration, requirements for style, and requirements for emotional rendering.

[0119] For example, the first frame requirement for a short video could be to select the starting point from the at least one key video segment that is most likely to immediately grab the consumer's attention, possibly by using flashbacks, narration, or close-ups to quickly create conflict or suspense.

[0120] For example, the editing requirements for short videos can be to shorten, speed up, or omit secondary plots while maintaining the core narrative logic, and to highlight key information and emotional points.

[0121] For example, the ending of a short video could involve intelligently creating or reinforcing suspense or conflict at the end or a key point, such as abruptly ending at the most exciting part or posing a question through narration. The goal is not to "finish" the story, but to create a strong sense of "to be continued," compelling consumers to seek answers in the longer video.

[0122] For example, special effects requirements could include automatically adjusting camera language (such as zooming, panning, tilting), lighting, filters, and pacing according to the needs of the storyline to enhance the artistry and appeal of the short video.

[0123] For example, background music requirements could be to match or generate background music and sound effects based on the emotional tone and plot development of the short video to enhance its appeal.

[0124] For example, the narration requirement could be to automatically generate engaging narration or guiding text for segments that require explanation or guidance, based on the original audio, subtitles, and narrative purpose of the long video, such as "Want to know their fate? Click to watch the full version!"

[0125] For example, style requirements and emotional rendering requirements can be based on the original style of the long video and the preferences of the j-th consumer group, and the generated short video can be style-transferred or emotionally rendered to make it more in line with user expectations.

[0126] In this embodiment, after generating short video generation strategy information for the j-th consumer group based on the above steps, the computing device edits at least one key video segment based on the short video generation strategy information for the j-th consumer group to generate a short video corresponding to the j-th consumer group. For example, as... Figure 4As shown, the computing device inputs the short video generation strategy information of the j-th consumer group and the above-mentioned at least one key video segment into the short video generation model. Based on the short video generation strategy information of the j-th consumer group, the short video generation model edits the above-mentioned at least one key video segment to generate a short video corresponding to the j-th consumer group.

[0127] In some embodiments, to further improve the generation effect of the short video corresponding to the j-th consumer group, the computing device, when generating the short video corresponding to the j-th consumer group, considers not only the short video generation strategy information of the j-th consumer group, but also the interest feature information of the j-th consumer group and the content feature information of at least one key video segment mentioned above. For example, as... Figure 5 As shown, the computing device inputs the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment into the short video generation model. The short video generation model then edits the at least one key video segment based on the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment to generate a short video corresponding to the j-th consumer group.

[0128] Therefore, it can be seen that the computing device edits the at least one key video segment based on the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment, so that the generated short video corresponding to the j-th consumer group is more in line with the interests and preferences of the j-th consumer group, and can better retain the key content of the at least one key video segment, thereby improving the generation effect of the short video corresponding to the j-th consumer group.

[0129] The above embodiments describe how a computing device generates a short video corresponding to the j-th consumer group in one of M long videos based on the interest feature information of the j-th consumer group. Referring to this method, the computing device can determine the short video corresponding to each of the P consumer groups in each of the M long videos. In other words, for each of the M long videos, this embodiment can generate P short videos for that long video, and these P short videos correspond one-to-one with the P consumer groups.

[0130] In this embodiment, the short video corresponding to each of the M generated long videos is used to guide the second consumer to consume the long video corresponding to that short video. For example, when the second consumer consumes a certain short video, that short video can attract the second consumer to consume the long video corresponding to that short video. The second consumer can be any media content consumer.

[0131] The video processing method provided in this application acquires media consumption data of N first consumers and video data of M long videos, where N and M are both positive integers. Based on the media consumption data of the N first consumers, the interest characteristic information of the N first consumers is determined. Simultaneously, the video data of each of the M long videos is analyzed to obtain key video segments for each long video. Thus, for each of the M long videos, based on the interest characteristic information of the N first consumers, the key video segments of the long video are edited to generate a corresponding short video. This short video is used to guide second consumers to consume the corresponding long video. Therefore, to improve the guiding effect of the short video on the long video, the interest characteristic information of different consumers is considered when generating the short video corresponding to the long video, making the content of the generated short video more in line with the interests and preferences of different consumers. Generating personalized short videos can effectively stimulate the audience's interest, thereby guiding consumers to consume (watch) the complete long video, thus improving the click-through rate, completion rate, and user conversion rate of the long video. Furthermore, this application embodiment generates short videos corresponding to long videos based on key video segments of long videos, which can further enhance the attractiveness of the generated short videos. Moreover, the entire short video generation process does not require manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the efficiency of short video generation.

[0132] The foregoing section describes the specific process of generating short videos corresponding to M long videos in the video processing method provided in this application embodiment. Based on this, this application embodiment describes the specific process of recommending the short videos corresponding to the M long videos to a second consumer.

[0133] Figure 6 This is a schematic flowchart illustrating a video processing method according to an embodiment of this application. The method in this embodiment can be executed by a computing device, which can be the one described above. Figure 1 The terminal device 101 or server 102 in the middle.

[0134] like Figure 6 As shown, the method in this application embodiment includes: S201. Obtain the consumption characteristics information of the second consumer.

[0135] This application embodiment does not impose specific restrictions on the timing of when the computing device obtains the consumption characteristic information of the second consumer.

[0136] In one example, when a second consumer opens a video recommendation platform, the computing device obtains the second consumer's consumption characteristics information.

[0137] In one example, when a second consumer opens a specific video on a video recommendation platform, the consumption characteristics of that second consumer are obtained.

[0138] The consumption characteristic information of the second consumer includes at least one of the following: the second consumer's interest characteristic information, the second consumer's media consumption data in historical time periods, and the second consumer's media consumption data in the current time period.

[0139] In this application embodiment, the computing device obtains the consumption characteristic information of the second consumer in at least the following ways: Method 1: Obtain the media consumption data of the second consumer on the video recommendation platform during a historical time period from the platform the second consumer is currently browsing, as the second consumer's media consumption data for that historical time period; and obtain the second consumer's media consumption data on the video recommendation platform during the current time period, as the second consumer's media consumption data for that current time period. Further, the computing device analyzes the second consumer's media consumption data during the historical time period and the second consumer's media consumption data during the current time period to obtain the second consumer's interest characteristic information. The specific process can be referred to the relevant description in S102 above, and will not be repeated here.

[0140] Method 2: In this embodiment, before acquiring the consumption characteristic information of the second consumer, the computing device constructs a heterogeneous graph of consumers and videos. Specifically, the computing device acquires consumption data of multiple consumers for multiple videos, where the multiple videos include long videos and short videos. Then, based on the consumption data of the multiple consumers for the multiple videos, a heterogeneous graph of consumers and videos is constructed. This heterogeneous graph represents the consumption information of each of the multiple consumers for each of the multiple videos, as well as the association information between the multiple videos.

[0141] Specifically, the computing device acquires consumption data of multiple consumers for multiple videos over a recent period (e.g., a week or a month), including the aforementioned second consumer.

[0142] In one example, the aforementioned consumers could be those currently browsing the video recommendation platform for the second consumer. Correspondingly, the aforementioned videos could be the videos (including long and short videos) currently browsing the video recommendation platform for the second consumer.

[0143] In one example, the aforementioned multiple consumers can be consumers from different video recommendation platforms. Correspondingly, the aforementioned multiple videos can be videos from multiple video recommendation platforms (including long videos and short videos).

[0144] For example, consumer consumption data for videos includes relevant information about the video (such as whether it is a long or short video, the creator, the series, and hashtags) and consumer consumption behavior related to the video (such as watching, liking, commenting, sharing, subscribing, and paying for the video).

[0145] In this way, computing devices can construct a heterogeneous graph of consumers and videos based on the consumption data of these multiple consumers for these multiple videos. For example, the nodes of this heterogeneous graph are: consumers, short videos corresponding to long videos, long videos, creators, hashtags, long video series, etc. The edges of this heterogeneous graph are: short videos consumed by consumers (e.g., watching, liking, commenting, sharing), long videos consumed by consumers (e.g., watching, subscribing, paying), short videos originating from long videos, short videos created by creators, consumer interactions with hashtags, etc. The features of the nodes and edges in this heterogeneous graph are: content feature information of long videos (e.g., visual information, audio information, text information, and other multimodal information of long videos), consumer interest feature information, consumer demographic information, content feature information of short videos (e.g., visual information, audio information, text information, and other multimodal information of short videos), and metadata of the short video generation model when generating the short video, such as input prompts, parameters, model information, and evaluation data related to the generation process.

[0146] Therefore, the heterogeneous graph constructed in this embodiment can represent the consumption information of each of the multiple consumers for each of the multiple videos, as well as the association information between the multiple videos. Thus, in this method 2, the computing device can obtain the consumption characteristic information of the second consumer based on the heterogeneous graph. For example, based on the second consumer's identification information, the computing device can obtain the second consumer's media consumption data in historical time periods (e.g., the short and long videos consumed by the second consumer in historical time periods, and their consumption behavior for these short and long videos) from the heterogeneous graph, obtain the second consumer's media consumption data in the current time period (e.g., the short and long videos consumed by the second consumer in the current time period, and their consumption behavior for these short and long videos), and obtain the second consumer's interest characteristic information (e.g., the second consumer's multi-dimensional interest tags, the second consumer's viewing intention for long videos, and the second consumer's narrative curiosity characteristics).

[0147] In this embodiment of the application, after the computing device obtains the consumption characteristic information of the second consumer based on the above steps, it executes the following step S202.

[0148] S202. Based on the consumption characteristic information of the second consumer, select K short videos from the short videos corresponding to M long videos.

[0149] In this embodiment of the application, after obtaining the consumption characteristic information of the second consumer based on the above steps, the computing device selects K short videos from the short videos corresponding to the M long videos generated above and recommends them to the second consumer.

[0150] This application embodiment does not limit the specific method by which the computing device selects K short videos from M long videos corresponding to short videos based on the consumption characteristic information of the second consumer.

[0151] In some embodiments, such as Figure 7A As shown, for each short video among the M long videos, the computing device inputs the short videos corresponding to the M long videos and the consumption characteristic information of the second consumer into the recommendation model. This recommendation model can then select K short videos from the M long videos. For example, the computing device extracts the content feature information of the short video through an inference model, and then determines the second consumer's level of interest in the short video based on the content feature information and the second consumer's consumption characteristic information. For instance, the computing device determines the vector representation 'a' of the short video's content feature information and the vector representation 'b' of the second consumer's consumption characteristic information, calculates the distance between vector representation 'a' and vector representation 'b', and then determines the second consumer's level of interest in the short video based on this distance. For example, the smaller the distance, the more interested the second consumer is in the short video. In this way, the computing device can determine the second consumer's level of interest in each short video among the M long videos, and then select the K short videos with the highest level of interest from the M long videos.

[0152] In some embodiments, as described above, the data volume of the short videos corresponding to the M long videos is relatively large (for example, one long video corresponds to P short videos, and M long videos correspond to M*P short videos). In order to reduce the computational load on the computing device, the computing device selects K short videos through the following steps S202-A to S202-C (not shown in the figure): S202-A: Based on the consumption characteristic information of the second consumer, select Q short videos from the short videos corresponding to M long videos, where Q is less than M; S202-B: Based on the consumption characteristics information of the second consumer, determine the rating of each of the Q short videos; S202-C: Based on the rating, select K short videos from Q short videos.

[0153] In this implementation, the computing device first selects Q short videos from the M long videos based on the second consumer's consumption characteristics. For example, based on the second consumer's consumption characteristics, the computing device determines that the second consumer is interested in comedy films. Thus, the computing device can select short videos with comedy content from the M long videos as the Q short videos. Optionally, to further reduce the computational load on the computing device, when selecting Q short videos from the M long videos, more interest characteristics of the second consumer can be considered. For example, multi-dimensional interest tags such as content type, theme, style, director, and actors that the second consumer is interested in can be considered to select the Q short videos that the second consumer is likely to be interested in from the M long videos.

[0154] Next, the computing device determines the rating of each of the Q short videos based on the consumption characteristics information of the second consumer.

[0155] In one possible implementation, for each of the Q short videos, the computing device determines a vector representation 'a' of the content feature information of that short video and a vector representation 'b' of the consumption feature information of the second consumer. The distance between vector representation 'a' and vector representation 'b' is calculated, and based on this distance, the degree of interest of the second consumer in the short video is determined. Furthermore, based on the degree of interest of the second consumer in the short video, a rating for the short video is determined. For example, the degree of interest of the second consumer in the short video is directly determined as the rating for that short video.

[0156] In one possible implementation, the computing device determines the rating of each of the Q short videos through the following steps S202-B1 and S202-B2 (not shown in the figure): S202-B1. For the i-th short video among Q short videos, obtain the content feature information of the i-th short video and the content feature information of the long video corresponding to the i-th short video. S202-B2. Based on the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video, determine the rating of the i-th short video.

[0157] In this implementation, for each of the Q short videos, such as the i-th short video, the computing device first obtains the content feature information of the i-th short video and the content feature information of the corresponding long video. For example, the content feature information of the i-th short video and the corresponding long video can be obtained from the heterogeneous graph described above. Next, the computing device determines the rating of the i-th short video based on the consumption feature information of the second consumer, the content feature information of the i-th short video, and the content feature information of the corresponding long video.

[0158] In this embodiment of the application, the computing device determines the rating of the i-th short video in at least the following ways: Method 1: In this embodiment of the application, a scoring model is trained. The computing device inputs the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video into the scoring model. The scoring model outputs the score of the i-th short video.

[0159] Method 2: In this embodiment, the recommendation model can be a reinforcement learning model. The computing device can then use this model to predict the rating of the i-th short video. Specifically, the computing device constructs state information based on the consumption characteristics of the second consumer, the content characteristics of the i-th short video, and the content characteristics of the corresponding long video. Then, the rating of the i-th short video is determined using the policy model within the reinforcement learning model, given this state information.

[0160] In other words, in this example, such as Figure 7B As shown, the environment represents the user interaction ecosystem of the entire video recommendation platform. The consumption characteristics of the second consumer, the content characteristics of the i-th short video, and the content characteristics of the corresponding long video are used as state information. The i-th short video is used as the recommendation action. The policy model evaluates the long-term value (Q-value) of executing the recommendation action of the i-th short video under this state information, and this long-term value is determined as the rating of the i-th short video.

[0161] The embodiments of this application do not limit the specific network structure of the policy model, and can be deep Q-networks (DQN) or deep reinforcement learning networks such as Actor-Critic.

[0162] In this way, the computing device can determine the rating of each of the Q short videos, and then select K short videos from the Q short videos based on the ratings. Next, the following step S203 is executed.

[0163] S203. Based on K short videos, generate a short video recommendation list for the second consumer.

[0164] In this embodiment, based on the above steps, the computing device selects K short videos from the short videos corresponding to M long videos, and then generates a short video recommendation list for the second consumer based on these K short videos. For example, the computing device arranges the short video with the highest rating among these K short videos at the top to generate the short video recommendation list for the second consumer.

[0165] In some embodiments, if the computing device is the terminal device corresponding to the second consumer, the computing device can directly display the generated short video recommendation list on the current interface of the terminal device.

[0166] In some embodiments, if the computing device is a server, the computing device sends the generated short video recommendation list to the terminal device. The terminal device then displays the list on the current interface.

[0167] In some embodiments, if the computing device selects K short videos from Q short videos using method 2 described above, which evaluates the rating of each of the Q short videos through a strategy model and then selects K short videos from the Q short videos based on the rating, the method in this embodiment further includes: after recommending the short video recommendation list to a second consumer, the computing device collects feedback information from the second consumer, wherein the feedback information includes the second consumer's consumption information of the short videos in the short video recommendation list and consumption information of the long videos corresponding to the short videos in the short video recommendation list. Thus, the computing device determines the reward of the strategy model based on this feedback information.

[0168] For example, the reward of the strategy model is determined based on at least one of the following: the click-through rate (R_short_click), the completion rate (R_short_finish), and the interaction rate (R_short_interact) of the short videos in the short video recommendation list by the second consumer; the click-through rate (R_long_click), the completion rate (R_long_finish), the subscription rate (R_long_subscribe), and the paid conversion rate (R_long_pay) of the long videos corresponding to the short videos in the short video recommendation list.

[0169] For example, for each short video in the short video recommendation list for the second consumer, the computing device adds up the second consumer's click-through rate (R_short_click), completion rate (R_short_finish), interaction rate (R_short_interact) for that short video, as well as the click-through rate (R_long_click), completion rate (R_long_finish), related content subscription rate (R_long_subscribe), and paid conversion rate (R_long_pay) for the corresponding long video, to obtain the reward obtained by the strategy model for that short video. Then, the sum of the rewards obtained by the strategy model for each short video in the short video recommendation list is determined as the reward of the strategy model.

[0170] For example, for each short video in the short video recommendation list for the second consumer, the computing device calculates the weighted sum of the short video's click-through rate (R_short_click), completion rate (R_short_finish), interaction rate (R_short_interact), and the corresponding long video's click-through rate (R_long_click), completion rate (R_long_finish), related content subscription rate (R_long_subscribe), and paid conversion rate (R_long_pay). This sum is then used to determine the reward the strategy model receives for that short video. Finally, the sum of the rewards received by the strategy model for each short video in the recommendation list is determined as the reward for the strategy model.

[0171] For example, the computing device determines the reward R of the policy model using the following formula (1): R = (1) Where wi is a dynamically adjustable weight. Let be the reward for the i-th short video out of K short videos.

[0172] Next, based on the reward calculated from the above-mentioned strategy model, the computing device adjusts the parameters in the strategy model to improve the recommendation accuracy of the strategy model. For example, the computing device can use balancing strategies such as epsilon-greedy, Upper Confidence Bound (UCB), or Thompson Sampling to balance the exploration of new content and the utilization of known consumer preferences, thereby improving the adjustment effect of the strategy model. This allows the adjusted strategy model to improve recommendation accuracy in subsequent short video recommendation processes.

[0173] In this embodiment of the application, when the second consumer watches a short video from the aforementioned short video recommendation list, such as the first short video, ... Figure 8 The playback interface of the first short video shown includes a jump option. When triggered, the jump option allows the user to jump from the playback interface of the first short video to the playback interface of the corresponding first long video.

[0174] In some embodiments, to address the user experience of a second consumer jumping from a first short video to a corresponding first long video, a customized long video jump entry is provided when the second consumer watches the first short video. Specifically, if the second consumer triggers the jump option in the playback interface of the first short video while watching it, the computing device, in response to the second consumer's triggering operation, obtains the content feature information of the first short video and the narrative structure information of the corresponding first long video, where the first short video is the short video currently being played by the second consumer in the short video recommendation list. Next, based on the second consumer's consumption feature information, the content feature information of the first short video, and the narrative structure information of the first long video, the computing device determines the jump video frame of the first long video. Finally, based on the jump video frame of the first long video, the computing device jumps to that jump video frame within the first long video.

[0175] This application embodiment does not limit the specific method of determining the jump video frame of the first long video based on the consumption characteristic information of the second consumer of the computing device, the content characteristic information of the first short video, and the narrative structure information of the first long video.

[0176] In one possible implementation, the computing device trains a jump video frame prediction model. The computing device can input the consumption feature information of the second consumer, the content feature information of the first short video, the narrative structure information of the first long video, and the video data of the first long video into the jump video frame prediction model, and the jump video frame prediction model outputs the jump video frames of the first long video.

[0177] In one possible implementation, the consumption characteristic information of the second consumer includes the narrative curiosity characteristic of the second consumer. In this case, the computing device can determine the first jump mode corresponding to the second consumer from multiple candidate jump modes based on the narrative curiosity characteristic of the second consumer and the content characteristic information of the first short video. Then, based on the first jump mode and the narrative structure information of the first long video, the jump video frame of the first long video is determined. The multiple candidate jump modes include an immediate plot continuation mode and a background plot supplementation mode.

[0178] For example, if the first short video abruptly ends at a certain point, and the second consumer triggers the jump option, and the second consumer's narrative curiosity trait is 1 (i.e., they possess narrative curiosity), and the content characteristics of the first short video indicate that it is the opening segment of the first long video, then it means the second consumer wants to watch the subsequent video after the end of the first short video. In this case, the jump mode can be determined as the plot continuation mode. Thus, the computing device can determine the current video frame or the next video frame located at the end of the first short video within the first long video as the jump video frame of the first long video based on the narrative structure information of the first long video, achieving a seamless transition and satisfying the second consumer's desire to immediately know the subsequent plot development.

[0179] For example, if the second consumer's narrative curiosity trait is 1 (i.e., they possess narrative curiosity), and the content feature information of the first short video indicates that the first short video is a climax segment of the first long video, then the second consumer may not understand the preceding events. In this case, the jump mode can be determined as a background plot supplement mode. Thus, the computing device can, based on the narrative structure information of the first video, identify the video frames in the first long video containing important foreshadowing plot points preceding the climax segment as the jump frames for the first long video, thereby helping the second consumer establish contextual understanding.

[0180] In some embodiments, when the computing device detects that a second consumer has consumed the first long video, it also generates guidance information based on the second consumer's narrative curiosity characteristics and the video data of the first long video. For example, such as Figure 9 As shown in the embodiment of this application, a guidance model is trained. The computing device can input the video data of a first long video and the narrative curiosity characteristics of a second consumer into the guidance model. Based on the narrative curiosity characteristics of the second consumer, the guidance model can analyze the video data of the first long video and generate personalized guidance information, such as: "Want to know why they came here? Click to start watching from the 15th minute!" or "The full story is a hundred times more exciting than this, click here to unlock the whole series now!" This guidance information is used to further guide the second consumer to consume all or specific portions of the first long video.

[0181] In some embodiments, when the computing device detects that a second consumer is consuming a first short video, it may also generate guidance information based on the narrative curiosity characteristics of the second consumer and the video data of the first long video, so as to guide the second consumer to directly jump to consume all or a specific part of the content of the first long video when consuming the first short video.

[0182] In some embodiments, when the computing device detects that a second consumer is consuming the first short video, it may also generate interactive elements, such as dynamically presenting clickable buttons or cards like "Continue watching", "View the full series", or "Learn more" during or after the playback of the first short video. Optionally, it may also include information such as the number of episodes, duration, and rating of the first long video.

[0183] In some embodiments of this application, the computing device statistically analyzes various feedback information from consumers regarding short and long videos, including but not limited to: short video exposure, click-through rate, playback completion rate, and interaction data; long video jump click-through rate, playback completion rate after jump, payment rate, membership conversion rate, and retention rate. In one example, this application embodiment also conducts A / B testing based on the feedback data to verify the effectiveness of different solutions, primarily focusing on the impact on consumers' long-term activity and paid conversion. Subsequently, based on real-time feedback information and A / B test results, the computing device iteratively updates the aforementioned models (e.g., short video generation model, recommendation model, guidance model, etc.) periodically or irregularly, forming a continuously optimized, spiraling intelligent closed loop.

[0184] The video processing method provided in this application embodiment obtains the consumption characteristic information of a second consumer, which includes at least one of the following: the second consumer's interest characteristic information, the second consumer's media consumption data in historical time periods, and the second consumer's media consumption data in the current time period. Next, based on the second consumer's consumption characteristic information, K short videos are selected from the short videos corresponding to M long videos, where K is a positive integer. Finally, based on the K short videos, a short video recommendation list for the second consumer is generated. In this application embodiment, selecting K short videos from the short videos corresponding to M long videos based on the second consumer's consumption characteristic information makes the selected K short videos more likely to arouse the second consumer's interest, effectively guiding the second consumer to jump from short videos and continue watching long videos. Furthermore, to improve the transition experience from short videos to long videos in this embodiment, the computing device accurately determines the transition video frame of the first long video based on the second consumer's consumption characteristics, the content characteristics of the first short video, and the narrative structure of the first long video. Then, based on this accurate transition video frame, the user jumps to that frame in the long video, thus improving the transition experience. Further, the computing device generates personalized guidance information to more smoothly and naturally guide the second consumer to watch the long video, significantly improving the second consumer's experience and conversion efficiency.

[0185] The foregoing has described the short video generation and short video recommendation processes in the video processing method of this application. The following section will combine... Figure 10 and Figure 11 The overall process of the video processing method according to the embodiments of this application will be introduced.

[0186] Figure 10 This is a schematic flowchart of a video processing method provided in an embodiment of this application. Figure 11 This is a schematic diagram illustrating the framework of the video processing method provided in the embodiments of this application. The method in the embodiments of this application can be derived from the above... Figure 1 The terminal device 101 and server 102 shown are executed.

[0187] like Figure 10 As shown, the video processing method of this application embodiment includes: S301. The server obtains the media consumption data of each of the N first consumers and the video data of each of the M long videos.

[0188] Where N and M are both positive integers.

[0189] The specific implementation process of S301 can be referred to the relevant description of S101 above, and will not be repeated here.

[0190] S302. The server determines the interest characteristics of each of the N first consumers based on the media consumption data of each first consumer.

[0191] For example, such as Figure 11 As shown in the embodiment of this application, the module in the server used to process the media consumption data of each first consumer to obtain the interest feature information of each first consumer is denoted as the interest feature extraction module.

[0192] The specific implementation process of S302 can be referred to the relevant description of S102 above, and will not be repeated here.

[0193] S303. The server analyzes the video data of each of the M long videos to determine the key video segments of each long video.

[0194] For example, such as Figure 11 As shown in the embodiment of this application, the module that analyzes video data of long videos and determines key video segments of long videos on the server is referred to as the content analysis module.

[0195] For example, a server analyzes video data from a long video to obtain content feature information. Then, based on this content feature information, it identifies key video segments of the long video. Exemplarily, the content feature information includes at least one of the long video's multimodal information, semantic sentiment features, and narrative structure information.

[0196] The specific implementation process of S303 can be referred to the relevant description of S103 above, and will not be repeated here.

[0197] S304. For each of the M long videos, the server edits and processes the key video segments of the long video based on the interest feature information of N first consumers to generate a short video corresponding to the long video.

[0198] Among them, short videos are used to guide second consumers to consume long videos corresponding to short videos.

[0199] For example, such as Figure 11 As shown, the server uses a short video generation model to edit key video segments of a long video based on the interest feature information of N first consumers, and generates a short video corresponding to the long video.

[0200] The specific implementation process of S304 can be referred to the relevant description of S104 above, and will not be repeated here.

[0201] The above steps S301 to S304 describe the process of generating short videos corresponding to M long videos.

[0202] S305. The server obtains the consumption characteristic information of the second consumer.

[0203] The consumption characteristic information of the second consumer includes at least one of the following: the second consumer's interest characteristic information, the second consumer's media consumption data in historical time periods, and the second consumer's media consumption data in the current time period.

[0204] The specific implementation process of S305 can be referred to the relevant description of S201 above, and will not be repeated here.

[0205] S306. Based on the consumption characteristic information of the second consumer, the server selects K short videos from the short videos corresponding to M long videos.

[0206] Where K is a positive integer.

[0207] For example, such as Figure 11 As shown, the server uses a recommendation model to select K short videos from the short videos corresponding to M long videos, based on the consumption characteristics information of the second consumer.

[0208] The specific implementation process of S306 can be referred to the relevant description of S202 above, and will not be repeated here.

[0209] S307. The server generates a short video recommendation list for the second consumer based on K short videos.

[0210] The specific implementation process of S307 can be referred to the relevant description of S203 above, and will not be repeated here.

[0211] S308. The server sends the short video recommendation list to the terminal device corresponding to the second consumer.

[0212] S309. The terminal device responds to the second consumer's viewing operation on the first short video in the short video recommendation list and plays the first short video.

[0213] S310, In response to the second consumer's triggering operation of the jump option in the playback interface of the first short video, the terminal device sends a jump request to the server.

[0214] For example, the redirect request may include the identification information of the second consumer, and optionally, it may also include the identification information of the first short video.

[0215] S311. Based on the redirect request, the server obtains the content feature information of the first short video and the narrative structure information of the first long video corresponding to the first short video.

[0216] The first short video is the short video that the second consumer is currently playing in the short video recommendation list.

[0217] S312. The server determines the jump video frame of the first long video based on the consumption characteristic information of the second consumer, the content characteristic information of the first short video, and the narrative structure information of the first long video.

[0218] For example, such as Figure 11 As shown, the server uses a jump video frame prediction model to determine the jump video frame of the first long video based on the consumption characteristics of the second consumer, the content characteristics of the first short video, and the narrative structure of the first long video.

[0219] The specific implementation process of S312 can be referred to the relevant description of S203 above, and will not be repeated here.

[0220] S313. The server sends the video data of the first long video jump frame, as well as the video data after the jump frame, to the terminal device.

[0221] S314. The terminal device plays the first video starting from the jump video frame based on the video data of the jump video frame of the first long video and the video data after the jump video frame.

[0222] S315. When the second consumer consumes the first long video, the server generates guidance information based on the narrative curiosity characteristics of the second consumer and the video data of the first long video.

[0223] The guidance information is used to guide the second consumer to consume all or a specific portion of the first long video.

[0224] For example, such as Figure 11 As shown, the server generates guidance information based on the narrative curiosity characteristics of the second consumer and the video data of the first long video through a guidance model.

[0225] The specific implementation process of S315 can be referred to the relevant description of S203 above, and will not be repeated here.

[0226] S316. The server sends the boot information to the terminal device.

[0227] S317. The terminal device displays guidance information on the playback interface of the first video.

[0228] S318. The server collects feedback information from the second consumer and adjusts the model based on the feedback information.

[0229] For example, such as Figure 11 As shown, the server collects feedback information from the second consumer through the feedback module and adjusts the model based on the feedback information.

[0230] The feedback information includes the second consumer's consumption information for short videos in the short video recommendation list, as well as their consumption information for the corresponding long videos in the short video recommendation list. For example, the feedback information includes at least one of the following: click-through rate, completion rate, and interaction rate of the short videos in the short video recommendation list; click-through rate, completion rate, subscription rate of related content, and paid conversion rate of the corresponding long videos in the short video recommendation list.

[0231] Among them, the models that need to be adjusted include Figure 11 The model includes at least one of the following: short video generation model, recommendation model, jump video frame prediction model, and guidance model.

[0232] Therefore, to enhance the guiding effect of short videos on long videos, this application's embodiments consider the interest characteristics of different consumers when generating short videos corresponding to long videos. This makes the content of the generated short videos more in line with the interests and preferences of different consumers, generating personalized short videos that can effectively stimulate viewers' interest and guide consumers to consume (watch) the complete long video, thereby increasing the click-through rate, completion rate, and user conversion rate of the long video. Furthermore, this application's embodiments generate short videos corresponding to long videos based on key video segments of the long video, which can further enhance the attractiveness of the generated short videos. The entire short video generation process requires no manual intervention and can automatically generate short videos corresponding to long videos, thereby reducing the manual cost of short video generation and improving the efficiency of short video generation. When recommending short videos, based on the consumption characteristics of the second consumer, K short videos that are more likely to arouse the second consumer's interest are selected from M short videos corresponding to long videos and recommended to the second consumer, which can effectively guide the second consumer to jump from short videos and continue watching long videos. Furthermore, to improve the transition experience from short videos to long videos in this embodiment, the computing device accurately determines the transition video frame of the first long video based on the second consumer's consumption characteristics, the content characteristics of the first short video, and the narrative structure of the first long video. Then, based on this accurate transition video frame, the user jumps to that frame in the long video, thus improving the transition experience. Further, the computing device generates personalized guidance information to more smoothly and naturally guide the second consumer to watch the long video, significantly improving the second consumer's experience and conversion efficiency.

[0233] The above text combined Figures 2 to 11 The method embodiments of this application are described in detail below, in conjunction with... Figure 12 The following describes in detail the device embodiments of this application.

[0234] Figure 12 This is a schematic block diagram of a video processing apparatus provided in an embodiment of this application.

[0235] like Figure 12 As shown, the video processing device 10 includes: The acquisition unit 11 is used to acquire media consumption data of each first consumer among N first consumers, and video data of each long video among M long videos, where N and M are both positive integers; Interest determination unit 12 is used to determine the interest feature information of each first consumer based on the media consumption data of each first consumer among the N first consumers; Unit 13 is used to analyze the video data of each of the M long videos to obtain the key video segments of each long video; Processing unit 14 is configured to, for each of the M long videos, edit key video segments of the long video based on the interest feature information of the N first consumers to generate a short video corresponding to the long video, and the short video is used to guide the second consumers to consume the long video corresponding to the short video.

[0236] In some embodiments, the generation unit 14 is specifically used to cluster the N first consumers based on the interest feature information of each of the N first consumers to obtain P consumer groups, where P is a positive integer less than or equal to N; for the j-th consumer group among the P consumer groups, based on the interest feature information of the j-th consumer group, the key video segments of the long video are edited to generate a short video corresponding to the j-th consumer group, where j is a positive integer from 1 to P.

[0237] In some embodiments, the generation unit 14 is specifically used to select at least one key video segment that the j-th consumer group is interested in from the key video segments of the long video based on the interest feature information of the j-th consumer group; and to edit the at least one key video segment based on the interest feature information of the j-th consumer group to generate a short video corresponding to the j-th consumer group.

[0238] In some embodiments, the generation unit 14 is specifically used to obtain content feature information of each key video segment in the at least one key video segment; determine short video generation strategy information for the j-th consumer group based on the interest feature information of the j-th consumer group and the content feature information of each key video segment in the at least one key video segment; and edit the at least one key video segment based on the short video generation strategy information of the j-th consumer group to generate a short video corresponding to the j-th consumer group.

[0239] In some embodiments, the generation unit 14 is specifically used to edit the at least one key video segment based on the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment, to generate a short video corresponding to the j-th consumer group.

[0240] In some embodiments, the short video generation strategy information includes at least one of the following: the hook type of the short video, and the style requirements of the short video. The style requirements of the short video include at least one of the following: the first frame requirements of the short video, the editing requirements of the short video, the ending requirements of the short video, the special effects requirements, the background music requirements, the narration requirements, the style requirements, and the emotional rendering requirements.

[0241] In some embodiments, the acquisition unit 11 is further configured to acquire the consumption characteristic information of the second consumer, the consumption characteristic information of the second consumer including at least one of the following: the interest characteristic information of the second consumer, the media consumption data of the second consumer in a historical time period, and the media consumption data of the second consumer in the current time period; the selection unit 13 is further configured to select K short videos from the short videos corresponding to the M long videos based on the consumption characteristic information of the second consumer, where K is a positive integer; the processing unit 14 is further configured to generate a short video recommendation list for the second consumer based on the K short videos.

[0242] In some embodiments, before acquiring the consumption characteristic information of the second consumer, the acquisition unit 11 is further configured to acquire consumption data of multiple consumers on multiple videos, the multiple videos including long videos and short videos; based on the consumption data of the multiple consumers on the multiple videos, construct a heterogeneous graph of consumers and videos, the heterogeneous graph representing the consumption information of each of the multiple consumers on each of the multiple videos, and the association information between the multiple videos; based on the identification information of the second consumer, acquire the consumption characteristic information of the second consumer from the heterogeneous graph.

[0243] In some embodiments, the selection unit 13 is specifically used to select Q short videos from the short videos corresponding to the M long videos based on the consumption characteristic information of the second consumer, where Q is less than M; determine the rating of each of the Q short videos based on the consumption characteristic information of the second consumer; and select K short videos from the Q short videos based on the rating.

[0244] In some embodiments, the selection unit 13 is specifically used to obtain, for the i-th short video among the Q short videos, the content feature information of the i-th short video and the content feature information of the long video corresponding to the i-th short video; and to determine the rating of the i-th short video based on the consumption feature information of the second consumer, the content feature information of the i-th short video, and the content feature information of the long video corresponding to the i-th short video.

[0245] In some embodiments, the selection unit 13 is specifically used to construct state information based on the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video; and to determine the rating of the i-th short video under the state information through the strategy model.

[0246] In some embodiments, the processing unit 14 is further configured to, after recommending the short video recommendation list to the second consumer, collect feedback information from the second consumer, the feedback information including the second consumer's consumption information of short videos in the short video recommendation list and consumption information of long videos corresponding to short videos in the short video recommendation list; determine the reward of the strategy model based on the feedback information; and adjust the parameters in the strategy model based on the reward.

[0247] In some embodiments, the processing unit 14 is specifically configured to determine the reward of the strategy model based on at least one of the following: the click-through rate of the short videos in the short video recommendation list by the second consumer, the playback completion rate of the short videos, the interaction rate of the short videos, and the click-through rate of the long videos corresponding to the short videos in the short video recommendation list, the playback completion rate of the long videos, the subscription rate of the related content of the long videos, and the paid conversion rate of the long videos.

[0248] In some embodiments, the processing unit 14 is further configured to, in response to the second consumer's triggering operation of the jump option in the playback interface of the first short video, obtain the content feature information of the first short video and the narrative structure information of the first long video corresponding to the first short video, wherein the first short video is the short video currently being played by the second consumer in the short video recommendation list; determine the jump video frame of the first long video based on the consumption feature information of the second consumer, the content feature information of the first short video, and the narrative structure information of the first long video; and jump to the jump video frame in the first long video based on the jump video frame of the first long video.

[0249] In some embodiments, the consumption feature information includes the narrative curiosity feature of the second consumer. The processing unit 14 is specifically used to determine the first jump mode corresponding to the second consumer from multiple candidate jump modes based on the narrative curiosity feature of the second consumer and the content feature information of the first short video. The multiple candidate jump modes include an immediate plot continuation mode and a background plot supplement mode. Based on the first jump mode and the narrative structure information of the first long video, the processing unit 14 determines the jump video frame of the first long video.

[0250] In some embodiments, the consumption feature information includes the narrative curiosity feature of the second consumer. The processing unit 14 is further configured to generate guidance information based on the narrative curiosity feature of the second consumer and the video data of the first long video when the second consumer consumes the first long video. The guidance information is used to guide the second consumer to consume all or a specific part of the content of the first long video.

[0251] In some embodiments, the processing unit 14 is specifically configured to: perform content parsing on the video data of the j-th long video among the M long videos to obtain multimodal information of the j-th long video, wherein the multimodal information includes at least one of visual information, audio information, and text information, and j is a positive integer from 1 to M; determine the semantic sentiment feature information of the j-th long video based on the multimodal information of the j-th long video; perform narrative structure analysis on the video data of the j-th long video to obtain the narrative structure information of the j-th long video; and select multiple key video segments from the j-th long video based on the semantic sentiment feature information and the narrative structure information of the j-th long video.

[0252] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, further details will not be provided here. Specifically, Figure 12 The apparatus shown can perform the embodiments of the above-described method, and the foregoing and other operations and / or functions of each module in the apparatus are for implementing the embodiments of the above-described method, which will not be described in detail here for the sake of brevity.

[0253] The apparatus of this application embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this application can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this application embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0254] Figure 13 This is a schematic block diagram of a computing device provided in an embodiment of this application. The computing device can be a terminal device or a server. The example will use the computing device as a server to execute the above-described video processing method.

[0255] like Figure 13 As shown, the computing device 40 may include: The system includes a memory 41 and a processor 42. The memory 41 stores a computer program 43 and transfers the computer program 43 to the processor 42. In other words, the processor 42 can retrieve and run the computer program 43 from the memory 41 to implement the methods described in the embodiments of this application.

[0256] For example, the processor 42 can be used to execute the steps in the above method according to the instructions in the computer program 43.

[0257] In some embodiments of this application, the processor 42 may include, but is not limited to: General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0258] In some embodiments of this application, the memory 41 includes, but is not limited to: Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0259] In some embodiments of this application, the computer program 43 may be divided into one or more modules, which are stored in the memory 41 and executed by the processor 42 to perform the method provided in this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 43 in the computing device.

[0260] like Figure 13 As shown, the computing device 40 may further include: Transceiver 34, which can be connected to processor 42 or memory 41.

[0261] The processor 42 can control the transceiver 34 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 34 may include a transmitter and a receiver. The transceiver 34 may further include antennas, and the number of antennas may be one or more.

[0262] It should be understood that the various components in the computing device 40 are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.

[0263] This application also provides a computer storage medium storing a computer program thereon, which, when loaded and executed by a computing device, enables the computing device to implement the above-described method embodiments.

[0264] This application also provides a computer program product comprising a computer program stored in a readable storage medium. At least one processor of a computing device can read the computer program from the readable storage medium, load and execute the computer program, and cause the computing device to implement the method embodiments described above.

[0265] In other words, when implemented using software, it can be implemented wholly or partially in the form of a computer program product. This computer program product includes a computer program. When the computer program is loaded and executed on a computing device, it generates, wholly or partially, the processes or functions according to the embodiments of this application. The computer program can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program can be transferred from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to the computing device or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0266] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0267] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0268] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0269] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Furthermore, reference can be made between the various method embodiments and between the various device embodiments; identical or corresponding content in different embodiments can be mutually referenced, without further elaboration.

Claims

1. A video processing method, characterized in that, include: Obtain the media consumption data of each of the N first consumers and the video data of each of the M long videos, where N and M are both positive integers; Based on the media consumption data of each of the N first consumers, determine the interest characteristics of each first consumer; Analyze the video data of each of the M long videos to obtain the key video segments of each long video; For each of the M long videos, based on the interest feature information of the N first consumers, the key video segments of the long video are edited to generate a short video corresponding to the long video. The short video is used to guide the second consumer to consume the long video corresponding to the short video.

2. The method according to claim 1, characterized in that, The step of editing key video segments of the long video based on the interest feature information of the N first consumers to generate a short video corresponding to the long video includes: Based on the interest characteristics of each of the N first consumers, the N first consumers are clustered to obtain P consumer groups, where P is a positive integer less than or equal to N; For the j-th consumer group among the P consumer groups, based on the interest feature information of the j-th consumer group, the key video segments of the long video are edited to generate a short video corresponding to the j-th consumer group, where j is a positive integer from 1 to P.

3. The method according to claim 2, characterized in that, The step of editing key video segments of the long video based on the interest feature information of the j-th consumer groups to generate a short video corresponding to the j-th consumer group includes: Based on the interest characteristics of the j-th consumer group, at least one key video segment that the j-th consumer group is interested in is selected from the key video segments of the long video. Based on the interest characteristics of the j-th consumer group, the at least one key video segment is edited to generate a short video corresponding to the j-th consumer group.

4. The method according to claim 3, characterized in that, The step of editing at least one key video segment based on the interest feature information of the j-th consumer group to generate a short video corresponding to the j-th consumer group includes: Obtain the content feature information of each key video segment in the at least one key video segment; Based on the interest characteristics of the j-th consumer group and the content characteristics of each key video segment in the at least one key video segment, the short video generation strategy information of the j-th consumer group is determined. Based on the short video generation strategy information of the j-th consumer group, the at least one key video segment is edited to generate a short video corresponding to the j-th consumer group.

5. The method according to claim 4, characterized in that, The process of editing at least one key video segment based on the short video generation strategy information of the j-th consumer group to generate a short video corresponding to the j-th consumer group includes: Based on the short video generation strategy information of the j-th consumer group, the interest feature information of the j-th consumer group, and the content feature information of the at least one key video segment, the at least one key video segment is edited to generate a short video corresponding to the j-th consumer group.

6. The method according to claim 4, characterized in that, The short video generation strategy information includes at least one of the following: the hook type of the short video, and the style requirements of the short video. The style requirements of the short video include at least one of the following: the first frame requirements of the short video, the editing requirements of the short video, the ending requirements of the short video, the special effects requirements, the background music requirements, the narration requirements, the style requirements, and the emotional rendering requirements.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain the consumption characteristic information of the second consumer, which includes at least one of the following: the interest characteristic information of the second consumer, the media consumption data of the second consumer in a historical time period, and the media consumption data of the second consumer in the current time period. Based on the consumption characteristics information of the second consumer, K short videos are selected from the short videos corresponding to the M long videos, where K is a positive integer; Based on the K short videos, a short video recommendation list is generated for the second consumer.

8. The method according to claim 7, characterized in that, Before obtaining the consumption characteristic information of the second consumer, the method further includes: Acquire consumption data of multiple consumers for multiple videos, including long videos and short videos; Based on the consumption data of the multiple consumers on the multiple videos, a heterogeneous graph of consumers and videos is constructed. The heterogeneous graph represents the consumption information of each consumer on each of the multiple videos, as well as the association information between the multiple videos. The acquisition of the second consumer's consumption characteristic information includes: Based on the identification information of the second consumer, the consumption characteristic information of the second consumer is obtained from the heterogeneous graph.

9. The method according to claim 7, characterized in that, The step of selecting K short videos from the short videos corresponding to the M long videos based on the consumption characteristic information of the second consumer includes: Based on the consumption characteristics information of the second consumer, Q short videos are selected from the short videos corresponding to the M long videos, where Q is less than M; Based on the consumption characteristics information of the second consumer, determine the rating of each of the Q short videos; Based on the scores, K short videos are selected from the Q short videos.

10. The method according to claim 9, characterized in that, The step of determining the rating of each of the Q short videos based on the consumption characteristic information of the second consumer includes: For the i-th short video among the Q short videos, obtain the content feature information of the i-th short video and the content feature information of the long video corresponding to the i-th short video; The rating of the i-th short video is determined based on the consumption characteristics of the second consumer, the content characteristics of the i-th short video, and the content characteristics of the long video corresponding to the i-th short video.

11. The method according to claim 10, characterized in that, The step of determining the rating of the i-th short video based on the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video includes: Based on the consumption characteristic information of the second consumer, the content characteristic information of the i-th short video, and the content characteristic information of the long video corresponding to the i-th short video, state information is constructed; The rating of the i-th short video is determined using a strategy model under the given state information.

12. The method according to claim 11, characterized in that, The method further includes: After recommending the short video recommendation list to the second consumer, the feedback information of the second consumer is collected. The feedback information includes the consumption information of the second consumer for the short videos in the short video recommendation list, and the consumption information for the long videos corresponding to the short videos in the short video recommendation list. Based on the feedback information, the reward of the strategy model is determined; Based on the reward, the parameters in the strategy model are adjusted.

13. The method according to claim 12, characterized in that, Determining the reward of the strategy model based on the feedback information includes: The reward of the strategy model is determined based on at least one of the following: the click-through rate of the short videos in the short video recommendation list by the second consumer, the completion rate of the short videos, the interaction rate of the short videos, the click-through rate of the long videos corresponding to the short videos in the short video recommendation list, the completion rate of the long videos, the subscription rate of the related content of the long videos, and the paid conversion rate of the long videos.

14. The method according to claim 7, characterized in that, The method further includes: In response to the second consumer's triggering operation of the jump option in the playback interface of the first short video, the content feature information of the first short video and the narrative structure information of the first long video corresponding to the first short video are obtained, wherein the first short video is the short video that the second consumer is currently playing in the short video recommendation list; Based on the consumption characteristics of the second consumer, the content characteristics of the first short video, and the narrative structure of the first long video, the jump video frame of the first long video is determined. Based on the jump video frame of the first long video, jump to the jump video frame in the first long video.

15. The method according to claim 14, characterized in that, The consumption characteristic information includes the narrative curiosity characteristic of the second consumer. The step of determining the jump video frame of the first long video based on the consumption characteristic information of the second consumer, the content characteristic information of the first short video, and the narrative structure information of the first long video includes: Based on the narrative curiosity characteristics of the second consumer and the content characteristics of the first short video, the first jump mode corresponding to the second consumer is determined from multiple candidate jump modes. The multiple candidate jump modes include the real-time plot continuation mode and the background plot supplement mode. Based on the first jump mode and the narrative structure information of the first long video, the jump video frame of the first long video is determined.

16. The method according to claim 14, characterized in that, The consumption characteristic information includes the narrative curiosity characteristic of the second consumer, and the method further includes: When the second consumer consumes the first long video, guidance information is generated based on the narrative curiosity characteristics of the second consumer and the video data of the first long video. The guidance information is used to guide the second consumer to consume all or a specific part of the content of the first long video.

17. The method according to any one of claims 1-6 and 8-16, characterized in that, The process of analyzing the video data of each of the M long videos to obtain key video segments for each long video includes: For the j-th long video among the M long videos, the video data of the j-th long video is parsed to obtain the multimodal information of the j-th long video. The multimodal information includes at least one of visual information, audio information and text information, where j is a positive integer from 1 to M. Based on the multimodal information of the j-th long video, determine the semantic sentiment feature information of the j-th long video; Narrative structure analysis is performed on the video data of the j-th long video to obtain the narrative structure information of the j-th long video; Based on the semantic and emotional features and narrative structure information of the j-th long video, multiple key video segments are selected from the j-th long video.

18. A video processing apparatus, characterized in that, include: The acquisition unit is used to acquire media consumption data of each of the N first consumers and video data of each of the M long videos, where N and M are both positive integers. An interest determination unit is used to determine the interest characteristic information of each first consumer based on the media consumption data of each of the N first consumers; The selection unit is used to analyze the video data of each of the M long videos to obtain the key video segments of each long video; The processing unit is configured to, for each of the M long videos, edit key video segments of the long video based on the interest feature information of the N first consumers to generate a short video corresponding to the long video, and the short video is used to guide the second consumers to consume the long video corresponding to the short video.

19. A computing device, characterized in that, Including processor and memory; The memory is used to store computer programs; The processor is configured to invoke and run a computer program stored in the memory to implement the method as described in any one of claims 1 to 17.

20. A computer-readable storage medium, characterized in that, Used to store computer programs; The computer program is loaded and executed by a computing device to implement the method as described in any one of claims 1 to 17.