Video explanation method, electronic device, computer storage medium and computer program product

By analyzing user behavior and creating various types of narration roles, the problem of lack of emotional expression and personalization in existing intelligent narration technologies has been solved, enabling personalized video narration and enhancing user experience and interactivity.

CN121509750APending Publication Date: 2026-02-10MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511485914.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing intelligent commentary technology lacks emotional expression and personalized features, cannot simulate the interaction and communication between real commentators, and cannot achieve personalized commentary that is unique to each individual.

Method used

By acquiring target videos and user behavior data, analyzing target feature types, creating multiple types of commentary roles, and playing commentary content based on commentary strategies, personalized video commentary can be achieved.

Benefits of technology

It enhances the user experience, enables intelligent multi-role commentary, simulates the interaction between real commentators, and provides personalized commentary services tailored to each individual.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509750A_ABST
    Figure CN121509750A_ABST
Patent Text Reader

Abstract

The invention provides a video explanation method, electronic equipment, a computer storage medium and a computer program product. The method comprises the following steps: acquiring a target video and a user behavior operation corresponding to the target video; analyzing a target feature of the target video based on the user behavior operation to obtain a target feature type; in response to a watching operation of a user for the target video, creating a multi-type explanation role based on the target feature type; and determining an explanation strategy corresponding to the multiple types of explanation roles, and playing the explanation content of the target video based on the explanation strategy. According to the invention, multi-role intelligent explanation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a video commentary method, an electronic device, a computer storage medium and a computer program product. BACKGROUND

[0002] With the increasing consumption of video content, video commentary technology has been widely used in sports events, film and television entertainment and other fields. Video commentary aims to describe and analyze video content in the form of voice or text to enhance the viewing experience of users. The current mainstream intelligent commentary technology mainly relies on data analysis models, combines event data and real-time picture recognition, and objectively states the key events and player information in the game process.

[0003] In the prior art, a single role voice broadcast method is usually used to present commentary content, and standardized commentary sentences are generated based on preset rules or machine learning models. This method can realize basic game progress restoration and data interpretation, but lacks emotional expression and personalized features, and cannot simulate the interactive communication between real human commentators. SUMMARY

[0004] The embodiments of the present application provide a method, device, computer readable storage medium and computer program product, which can realize personalized video commentary services, improve user experience, and expand the application boundary of intelligent commentary.

[0005] The technical solutions of the embodiments of the present application are as follows: The video commentary method provided by the embodiments of the present application comprises: obtaining a target video and a user behavior operation corresponding to the target video; analyzing a target feature of the target video based on the user behavior operation to obtain a target feature type; creating multiple types of commentary roles based on the target feature type in response to a user's viewing operation on the target video; determining a commentary strategy corresponding to the multiple types of commentary roles, and playing commentary content of the target video based on the commentary strategy.

[0006] The embodiments of the present application provide an electronic device, which comprises: a memory for storing computer executable instructions; a processor for executing the computer executable instructions stored in the memory to implement the video commentary method provided by the embodiments of the present application.

[0007] The embodiment of the present application provides a computer readable storage medium, which stores a computer program or computer executable instructions, and is used for realizing the video explanation method provided by the embodiment of the present application when the processor executes.

[0008] The embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions, and the computer program or computer executable instructions are executed by a processor to realize the video explanation method provided by the embodiment of the present application.

[0009] In the technical scheme of the embodiment of the present application, the target video and the user behavior operation corresponding to the target video are acquired; the target feature of the target video is analyzed based on the user behavior operation, and the target feature type is obtained; in response to the watching operation of the user for the target video, the multi-type explanation role is created based on the target feature type; the explanation strategy corresponding to the multi-type explanation role is determined, and the explanation content of the target video is played based on the explanation strategy; in this way, the target video and the user behavior operation are acquired, the target feature type is analyzed based on the user behavior, and then the multi-type explanation role is created according to the type and the corresponding explanation strategy is formulated, so that the personalized video explanation is realized. Compared with the single and fixed template intelligent explanation mode in the prior art, the explanation role with different positions and styles can be dynamically generated according to the user's preference, so that the explanation content is closer to the emotional needs of the user, and the user experience is improved. In addition, through the matching mechanism combining the user behavior and the explanation strategy, the personalized presentation of thousands of people can be realized, and the problem that the traditional intelligent explanation cannot reflect the user's personalized preference is solved. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 FIG. 1 is a structural schematic diagram of a video explanation system architecture provided by the embodiment of the present application; Figure 2 FIG. 2 is a structural schematic diagram of an apparatus provided by the embodiment of the present application; Figure 3 FIG. 3 is a flow schematic diagram of a video explanation method provided by the embodiment of the present application; Figure 4 FIG. 4 is a principle schematic diagram of a video explanation system provided by the embodiment of the present application. DETAILED DESCRIPTION

[0011] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by a person skilled in the art without making creative labor belong to the protection scope of the present application.

[0012] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or a different subset of all possible embodiments, and can be combined with each other, without conflict.

[0013] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects, and it is understood that the "first\second\third" can be interchanged in a specific order or sequence as allowed, so that the application embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0014] In the embodiments of the application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0015] Unless otherwise defined, all technical and scientific terms used in the embodiments of the application have the same meanings as commonly understood by one of ordinary skill in the art. The terms used in the embodiments of the application are only for the purpose of describing the embodiments of the application, and are not intended to limit the application.

[0016] In the application embodiments, the relevant data collection process should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.

[0017] At present, the existing intelligent commentary technology creates a model, identifies the official competition data and live pictures, surrounds the players on the court and the events occurring in the game, objectively restores the attack and defense and scoring situation in the game, and describes the live information of the players and the events. The intelligent commentary technology is difficult to achieve emotional presentation in game live commentary, and the information conveyed by voice is true but lacks praise and criticism. Therefore, it is impossible to show the individualized position of human commentators. The intelligent commentary on the market is a single voice broadcast explanation, there is no dialogue type commentary in multiple voice forms, and there is no effective interactive commentary effect for user demands. In addition, the intelligent commentary focuses more on real-time analysis of competition data and review of past data and objective description, which will lead to the fact that the intelligent commentary cannot achieve the individualized presentation of real commentary, and cannot realize the individualized commentary form for users as most users of intelligent commentary services cannot achieve the individualized commentary form.

[0018] Most AI-powered commentary systems do not support human-computer interaction, and foreseeable technological advancements will likely rely on user-initiated voice or text interaction. Furthermore, AI commentary often presents a single-role voice broadcast, whereas human commentary involves a host and multiple guest commentators in a conversational format. Therefore, AI commentary cannot replicate the real-time communication and intellectual exchange between human commentators.

[0019] To address the aforementioned issues, embodiments of this application provide a video narration method, an electronic device, a computer storage medium, and a computer program product, capable of enabling multi-role intelligent narration based on user preferences.

[0020] The following describes exemplary applications of the devices provided in the embodiments of this application. The electronic devices provided in the embodiments of this application can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and vehicle terminals, or they can be implemented as servers.

[0021] See Figure 1 , Figure 1 This is a schematic diagram of the video narration system architecture provided in the embodiments of this application, exemplified. Figure 1 The system involves server 101, terminal device 102, and network 103. Terminal device 102 is connected to server 101 through network 103, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.

[0022] In some embodiments, the present application embodiments can be implemented collaboratively by a server and a terminal device. For example, terminal device 102 sends a target video and the corresponding user behavior operation to server 101. Server 101 analyzes the target features of the target video based on the user behavior operation using the video narration method provided in the present application embodiments, obtains the target feature type, determines the narration strategy based on the target feature type, and stores it locally. After responding to the user's viewing operation on the target video, terminal device 102 triggers a multi-type narration role creation request and sends a multi-type narration role creation instruction to server 101. Server 101 creates multi-type narration roles based on the target feature type, determines the narration strategy corresponding to the multi-type narration roles based on the stored narration strategy, and sends the narration strategy to terminal device 102. Terminal device 102 then plays the narration content of the target video.

[0023] In other embodiments, the embodiments of this application can be implemented independently by a terminal device. Terminal device 102 acquires the target video and the corresponding user behavior operations; based on the user behavior operations, it analyzes the target features of the target video to obtain the target feature type. In response to a user's viewing operation on the target video, terminal device 102 sends a multi-type narration role creation request to server 101. Server 101 receives the multi-type narration role creation request and sends the model (narration strategy) used for video narration provided in the embodiments of this application to terminal device 102. Terminal device 102 receives the model sent by the server and downloads it locally. Using the model and the current target video data, it creates multi-type narration roles, determines the narration strategy corresponding to each multi-type narration role, and terminal device 200 plays the narration content of the target video based on the narration strategy.

[0024] In some embodiments, the terminal device or server can implement the video narration method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application program, module, or plug-in. Terminal devices include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft.

[0025] See Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The electronic device 400 shown can be either the server 100 or the terminal device 200 mentioned above. Figure 2 The illustrated electronic device 400 includes at least one processor 410, a memory 430, and at least one network interface 420. The various components of the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 440.

[0026] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0027] The memory 430 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 430 may optionally include one or more storage devices physically located away from the processor 410.

[0028] The memory 430 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 430 described in this application embodiment is intended to include any suitable type of memory.

[0029] In some embodiments, memory 430 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0030] Operating system 431 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 432 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2 A video narration device 433 stored in memory 430 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a data acquisition module 4331 and a data processing module 4332. These modules are logically integrated and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.

[0031] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video narration method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0032] Figure 3 This is a flowchart illustrating the video narration method provided in the embodiments of this application. The following will be combined with... Figure 3 The steps shown will be explained. It should be noted that... Figure 3 The video narration method in this example is illustrated using the user terminal as the executing entity. Figure 3 As shown, the method includes the following steps 301 to 304: Step 301: Obtain the target video and the corresponding user behavior operations.

[0033] In this embodiment, the target video can be video data obtained from the user's playback history. Specifically, the target video can be a live video, such as a live sports event, or a recorded video, such as a classic match. The user behavior corresponding to the target video can be certain actions performed by the user on specific content in the video based on the user's attitude. User behavior includes liking, commenting, sending bullet comments, blocking, giving negative reviews, repeatedly playing, pausing, and skipping, among other interactive behaviors during the video viewing process. User behavior can reflect the user's attitude towards certain specific content in the video, such as an athlete, the occurrence of an event, or related information (such as the match venue, weather conditions, etc.).

[0034] In some embodiments, after obtaining the target video from the user's playback history, the matches appearing in the target video can be analyzed to obtain the target matches. The rule for analyzing the target matches can be that matches that the user continuously watches or repeatedly watches are defined as target matches.

[0035] In its implementation, the system collects various user actions during video viewing via the client application and matches this data with timestamps from the video content to determine the user's attitude towards specific content within a given timeframe. This method of matching action data with timestamps not only improves the accuracy of user behavior analysis but also provides a reliable data foundation for the subsequent creation of various narration roles.

[0036] In practice, by collecting and analyzing user behavior in real time, the system can quickly identify users' emotional inclinations towards different video content, thus providing a basis for subsequent personalized intelligent commentary. This process of collecting and analyzing user behavior in real time and identifying users' emotional inclinations towards different video content effectively enhances user engagement and lays the foundation for generating multiple types of commentary characters with differing perspectives.

[0037] Step 302: Analyze the target features of the target video based on user behavior operations to obtain the target feature types.

[0038] In this embodiment, target features may include target objects, target events, or related information. For example, target features may be athletes in a target competition, events generated by athletes, competition-related information, etc. After identifying the target competition from the target video, the user's emotional tendency towards the target features can be analyzed based on user behavior, resulting in target feature types. Target feature types may include three categories: liking, disliking, and accepting. Specifically, the analysis of users' emotional tendencies towards athletes, events generated by athletes, and competition-related information while watching the target competition, using user operation data, yields three target feature types: liking, disliking, and accepting. The results relate to the evaluation definitions of athletes, events, and related information, respectively.

[0039] In some embodiments, target features include at least one of the following: target object features, target event features, and related information features. Here, target features refer to target objects, target events, or related information appearing in the target video. For example, target object features could be athletes appearing in the target competition. Target event features could be events created by athletes in the target competition, mainly involving competitive events under the competition rules, scoring events, win / loss determination events, violation events, exciting competitive actions, record-breaking events, injury events, interruption events, withdrawal events, etc. Related information features could include information such as coaches, competition venues, participating countries, athlete matchups, and local weather conditions.

[0040] For example, natural language processing (NLP) technology can be used to perform semantic analysis on user comments and bullet screen content to determine their sentiment. Simultaneously, by combining user behavior data such as clicks, dwell time, and replay frequency, a comprehensive assessment of user attitudes towards different target characteristics can be made. This multi-dimensional analysis method enables the system to more accurately identify users' true preferences, providing a basis for creating multiple types of commentary roles with different perspectives. Thus, by accurately identifying users' attitudes towards video content, the system can provide commentary services that are more tailored to users' interests. This feature not only enhances the user's immersion in the video content but also improves the overall viewing experience.

[0041] In some embodiments, the target feature type includes at least one of the following: a first type of target feature, a second type of target feature, and a third type of target feature; the user preference level corresponding to the first type of target feature is higher than the user preference level corresponding to the second type of target feature, and the user preference level corresponding to the second type of target feature is higher than the user preference level corresponding to the third type of target feature.

[0042] Here, target feature type refers to the sentiment classification result of user behavior actions on the target features (such as athletes, events, and related information) of the target video. Based on user actions such as attention, likes, comments, and bullet comments on the target features of the target video, target feature types are divided into three categories: The first category refers to users exhibiting a positive attitude towards the specific features contained in the target video while watching it. The second category refers to users not showing a clear preference for the target features in the target video, or not taking any significant action towards the target features. The third category refers to users holding a negative attitude towards the target features corresponding to the target video while watching it. By classifying target features into the first, second, and third categories, we can more accurately understand user preferences and adjust the interpretation strategy accordingly.

[0043] In this embodiment, by setting a hierarchical structure for target feature types, the system can more precisely differentiate users' preferences for target features in a target video, thus providing a basis for subsequent multi-role intelligent commentary. By setting a hierarchical structure for target feature types, the system can improve the accuracy of personalized commentary, thereby enhancing the user experience and further increasing the user's immersion and interactivity with the live stream content.

[0044] Specifically, the target features of the target video are analyzed based on user behavior operations to obtain target features, including: when the user behavior operation is a positive operation on the target features of the target video, the target feature type is the first type of target feature; when the user behavior operation is no operation on the target features of the target video, the target feature type is the second type of target feature; and when the user behavior operation is a negative operation on the target features of the target video, the target feature type is the third type of target feature.

[0045] Here, user behavior refers to the interactive behaviors exhibited by users towards the target features (such as athletes, events, and related information) of a target video while watching it. Specifically, positive user actions can include clicking the follow button, liking content, posting comments on the page, and sending positive bullet comments. No action means that the user does not react to the target features of the target video; the user neither likes nor performs negative actions. Negative user actions can include setting a negative rating for the target object, adding it to a blacklist, sending negative comments to the platform, and sending negative bullet comments on the video playback interface. Through the identification and classification of these actions, the system can determine the user's attitude towards the target features of the target video and categorize the target features into the corresponding target feature types. For example, if a user frequently likes an athlete, the system categorizes that athlete as a first-category target feature; if a user does not react to an event, the system categorizes that event as a second-category target feature; if a user repeatedly posts negative comments, the system categorizes the information involved in the negative comments as a third-category target feature.

[0046] For example, based on user behavior, the characteristics of the target object in the target video are analyzed to obtain the target object feature types. Specifically, the athletes appearing in the competition are analyzed to obtain the first type of target feature (liked athletes), the third type of target feature (disliked athletes), and the second type of target feature (accepted athletes). The analysis rules are as follows: for athletes appearing in the target competition, if a user follows, likes, or comments with positive feedback, it is considered a liked athlete; if a user blocks, dislikes, or comments with negative feedback, it is considered a disliked athlete; if a user does not follow, block, like, dislike, or comment on the athletes, it is considered an accepted athlete.

[0047] For example, based on user behavior, the target event characteristics of the target video are analyzed to obtain the target event characteristic types. Specifically, user attitude analysis is performed on the events generated by athletes to obtain the first type of target characteristic, namely, liked events; the third type of target characteristic, namely, disliked events; and the second type of target characteristic, namely, accepted events. The analysis rules are events created by athletes in the target competition, mainly involving competitive events under the competition rules, scoring events, win / loss determination events, violation events, exciting competitive actions, record-breaking events, injury events, interruption of the competition events, and withdrawal events. When the above events occur, if users like, comment, or leave positive comments, it is considered a liked event; if users dislike, comment, or leave negative comments, it is considered a disliked event; and if users do not like, dislike, comment, or leave comments, it is considered an accepted event.

[0048] For example, based on user behavior, the relevant information features of the target video are analyzed to obtain the relevant information feature types. Specifically, user attitude analysis is performed on the relevant information in the competition to obtain the first type of target feature, namely, liking-related information; the third type of target feature, namely, disliking-related information; and the second type of target feature, namely, accepting-related information. The analysis rules are as follows: the relevant information in the target competition includes information such as coaches, competition venues, participating countries, athlete matchups, and local weather. When the above information appears, if users follow, like, comment, or leave a positive evaluation in the bullet screen, it is considered liking-related information; if users block, dislike, comment, or leave a negative evaluation in the bullet screen, it is considered disliking-related information; if users do not follow, block, like, dislike, comment, or leave a bullet screen, it is considered accepting-related information.

[0049] By identifying and classifying the target features of a video through the above operations, we can determine the user's attitude towards those features and categorize them into the corresponding feature types. For example, if a user frequently likes a particular athlete, that athlete is categorized as a first-type feature; if a user does not react to an event, that event is categorized as a second-type feature; and if a user repeatedly posts negative comments, the information in those negative comments is categorized as a third-type feature. In this way, by classifying the target features of a video based on user behavior, we can accurately identify the user's true emotional inclination towards the target features of various types of videos. Classifying the target features of a video based on user behavior enables personalized intelligent commentary, improving user satisfaction and enhancing platform user stickiness and activity.

[0050] In this embodiment, by setting target feature types and analyzing user behavior, user preferences can be effectively identified and differentiated commentary content can be generated. By setting target feature types and analyzing user behavior, and generating differentiated commentary content, a highly customized intelligent commentary experience can be achieved, thereby meeting diverse user needs and further promoting the intelligent development of sports event live streaming.

[0051] Step 303: In response to the user's viewing action on the target video, create multiple types of narration roles based on the target feature type.

[0052] In some embodiments, prior to responding to a user's viewing action for a target video, the method further includes identifying whether the video the user is watching is the target video. Here, if the video is a live stream, it is necessary to identify whether the match in the live stream is the target match; if so, the creation of intelligent commentary multi-role is initiated.

[0053] In some embodiments, once a user begins watching a target video, various types of intelligent commentary characters can be created based on the target feature type.

[0054] In some embodiments, the types of commentary roles include at least one of the following: a first type of commentary role, a second type of commentary role, and a third type of commentary role. Here, the created commentary roles can be divided into three categories according to the type of target characteristic. For example, the first type of commentary role can be a liking commentary role, the second type of commentary role can be an accepting commentary role, and the third type of commentary role can be an averse commentary role.

[0055] For example, create a first type of interpreter role, such as the "liking" role. This role uses a strong liking tone and approach to explain content that the user likes and accepts (i.e., the first and second target characteristics), and a strong dislike tone and approach to explain content that the user dislikes (i.e., the third target characteristic). This represents an interpreter role that fully caters to the user, with a clear stance of liking and disliking. Create a second type of interpreter role, such as the "accepting" role. This role uses a calm tone to explain content that the user likes, accepts, and dislikes (i.e., the first, second, and third target characteristics), pursuing objective, truthful, and unbiased explanations. Create a third type of interpreter role, such as the "disliking" role. This role uses a strong dislike tone and approach to explain content that the user likes and accepts (i.e., the first and second target characteristics), and a strong liking tone and approach to explain content that the user dislikes (i.e., the third target characteristic). This represents an interpreter role that strongly opposes what the user likes while strongly supporting what the user dislikes. In this way, these three characters together construct a diverse commentary scenario, allowing users to experience interactions and exchanges similar to those between multiple real commentators. In the specific implementation, each intelligent commentary character has independent voice characteristics, tone style, and expression mode, so that users can clearly distinguish the positions and viewpoints of the three intelligent commentary characters.

[0056] In practice, by creating multiple intelligent commentary characters with different perspectives, it is possible to simulate the conflict and clashes between human commentators to a certain extent, thereby enhancing the emotional expression of the commentary content and user engagement. This multi-type commentary character mechanism not only enriches the presentation of video content but also provides more options for how such content is presented, offering users a more personalized viewing experience.

[0057] In some embodiments, the method further includes: generating an explanation strategy based on target features and target feature types, wherein the explanation strategy includes at least one of the following: explanation form, explanation intensity, and explanation duration.

[0058] It should be noted that the explanation strategy can also be called an intelligent explanation capability model or an intelligent explanation capability module; this application does not specifically limit it in this regard. The intelligent explanation capability model includes explanation format, explanation intensity, and explanation duration.

[0059] Here, based on the three target feature types of likes, dislikes, and acceptance, the intelligent commentary capability model will generate commentary strategies around the user's preferences. The capability module is the specific method for multi-role intelligent commentary in the subsequent live broadcast of the event, which includes setting three dimensions: commentary format, commentary intensity, and commentary duration, and determining different levels for each of these three dimensions.

[0060] For example, the commentary style is categorized into three types based on the live broadcast: preferred, disliked, and accepted commentary styles. Preferred commentary: Presents the match, athletes, events, and related information with detailed, comprehensive praise, encouragement, and a combination of forward-looking and retrospective language, delivered at a positive, emotionally charged, and engaging pace. Disliked commentary: Presents the match, athletes, events, and related information with concise, localized criticism and narrow descriptive language, delivered at a negative, depressed, and emotionally suppressed pace. Acceptable commentary: Presents the match, athletes, events, and related information with objective and neutral descriptive language, remaining true to all objective information about the match without elaboration or extrapolation, describing the true situation as seen on the live broadcast.

[0061] For example, the commentary duration refers to the time during a live broadcast when a multi-role intelligent commentator starts speaking when a user's preferred, disliked, or accepted athlete, event, or related information appears, and when the multi-role intelligent commentator stops speaking when the next user-preferred, disliked, or accepted athlete, event, or related information appears. The commentary duration is from the start of speaking to the end of speaking. The commentary duration will be evenly distributed according to the number of multi-role commentators, and each commentator will initially work at 100% of their allocated commentary time. The commentary duration is divided into five tiers, each tier representing 20% ​​of the total working time. See Table 1 below for details.

[0062] Table 1

[0063] For example, when athlete A appears, the multi-role intelligent commentary starts speaking; when athlete B appears, the multi-role intelligent commentary stops speaking. The commentary time surrounding athlete A is 2 minutes. During this period, two roles appear in the multi-role intelligent commentary, with each role allocated an average of 1 minute of commentary time. The initial working time rule for each role is the fifth level of commentary time, using 100% of the commentary time, that is, using the full 1 minute.

[0064] For example, the intensity of narration involves variations across seven dimensions: tone, speaking speed, volume, pauses, vocabulary, interjections, and quotations. It is divided into five intensity levels, ranging from 1 to 5, with 1 being the weakest and 5 the strongest. See Table 2 below for details.

[0065] Table 2

[0066] Step 304: Determine the commentary strategy corresponding to the various types of commentary roles, and play the commentary content of the target video based on the commentary strategy.

[0067] In this embodiment, each intelligent commentary character determines its own commentary strategy based on its position. Each intelligent commentary character then plays the commentary content of the target video based on its commentary strategy.

[0068] In some embodiments, before playing the narration content of the target video based on the narration strategy, the method further includes: acquiring data information of the target video; and generating narration content of the target video based on the data information of the target video, target object characteristics, target event characteristics, and related information characteristics.

[0069] Here, the data information of the target video includes data related to the target video. For example, the data information of the target video can be publicly available event data from the event organizing committee or professional event operation organization. After completing the analysis of user behavior, the target competition and the athletes, events, and related information that appear in the competition, such as likes, dislikes, and acceptance, are obtained. At the same time, the visual content of the target video is analyzed to obtain the target characteristic information conveyed by the visuals. Commentary is generated based on the target characteristic information and data information. For example, the visual content of the live broadcast of the target competition is analyzed to generate commentary on the objective information such as athletes and events conveyed by the visuals. Combined with the publicly available event data from the event organizing committee or professional event operation organization, the commentary content, i.e., the commentary, is generated based on the objective and authentic data of athletes, events, and related information obtained above.

[0070] In some embodiments, determining the commentary strategy corresponding to multiple types of commentary roles includes: determining the commentary strategy corresponding to the first type of commentary role as the first type of commentary strategy; determining the commentary strategy corresponding to the second type of commentary role as the second type of commentary strategy; and determining the commentary strategy corresponding to the third type of commentary role as the third type of commentary strategy.

[0071] Here, the interpretation strategy includes the interpretation format, interpretation intensity, and interpretation duration.

[0072] The forms of explanation include: the first type of explanation, the second type of explanation, and the third type of explanation; among them, the first type of explanation can be a liked explanation; the second type of explanation can be an accepted explanation; and the third type of explanation can be an unpleasant explanation.

[0073] The commentary intensity is divided into five levels: Level 1, Level 2, Level 3, Level 4, and Level 5. Among them, Level 1 commentary intensity is less than Level 2, Level 2 commentary intensity is less than Level 3, Level 3 commentary intensity is less than Level 4, and Level 4 commentary intensity is less than Level 5.

[0074] The commentary duration is divided into five tiers: Tier 1, Tier 2, Tier 3, Tier 4, and Tier 5. Tier 5 uses 100% of the commentary time, Tier 4 uses 80%, Tier 3 uses 60%, Tier 2 uses 40%, and Tier 1 uses 20%.

[0075] In some embodiments, the treatment of commentary in three commentary styles—favorite, disliked, and accepting—is based on changes in emotional expression, setting a stance of liking, disliking, or neutrality towards the athletes and their performances as portrayed in the commentary. The commentary is processed in terms of emotion, vocal tone, and other emotional aspects to achieve a clear auditory distinction between the commentators.

[0076] Narration intensity refers to the details of tone, speed, volume, pauses, vocabulary, interjections, and quotations in the specific audio generation of the narration. Five intensity levels represent different detail processing schemes. For the tone established by the three narration formats, variations in tone, speed, volume, pauses, vocabulary, interjections, and quotations are used to create variations in intensity, achieving character differentiation so that each character has a different performance style.

[0077] Commentary duration is determined by calculating the commentary time between two commentary points, assuming the commentary roles and their corresponding commentary strengths are in effect. This time is then allocated to different commentary roles to ensure each role has their designated commentary time. This system ensures the commentary is delivered within the allocated time. Since there are three scenarios regarding commentary time and script: the time is sufficient to finish the script, the time is insufficient, and there is remaining time after finishing the script, the presentation of the commentary time focuses on the calculation and definition of each commentator's allotted time, rather than requiring the script to be broadcast within the allotted time. For example, when an intelligent commentator is commentating on a World Cup match, if the screen shows a player running and looking for an opportunity, and the intelligent commentator introduces player A's position, but the screen switches to the opposing team scoring a goal, the intelligent commentator will not continue playing the commentary about player A's position; instead, it will commentate on the goal itself. It enables the commentary time to flexibly change around the real-time footage and timing of the game.

[0078] In some embodiments, the first type of explanation strategy includes: when explaining the first type of target features and / or the second type of target features, the explanation form is the first type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration; when explaining the third type of target features, the explanation form is the third type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration.

[0079] Here, for athletes, events, and related information that users like during the live stream, the commentary format for those who like the athletes, events, and related information will be set to the fifth level of intensity, with the same commentary duration. For athletes, events, and related information that users dislike, the commentary format for those who dislike them will be set to the fifth level of intensity, with the same commentary duration.

[0080] The first type of commentary typically involves enthusiastic descriptions of athletes, events, or related incidents that users like or accept, with a clear positive tone that evokes user resonance; conversely, it uses strong negative sentiment and tone when commenting on athletes, events, or related incidents that users dislike. For example, using the first type of commentary when a user's favorite athlete scores a crucial point can enhance user engagement and satisfaction. The fifth level of commentary intensity refers to the strongest combination of tone, speed, volume, pauses, vocabulary, interjections, and quotations. For example, using the fifth level of commentary intensity when a user dislikes an athlete commits a foul can evoke stronger negative emotions, thus enhancing emotional resonance. The fifth level of commentary duration means that the commentator corresponding to this duration will occupy all available time during the commentary. The settings or mechanisms represented by the fifth level of commentary duration indicate that the commentary content will be elaborated as thoroughly as possible to ensure that users receive a complete information experience. For example, after a user's favorite athlete delivers a brilliant performance, using the fifth level of commentary duration allows the user to fully understand the significance and background of the performance, thus deepening their impression.

[0081] In some embodiments, the second type of explanation strategy includes: when explaining the second type of target features, the explanation form is the second type of explanation form, the explanation intensity is the first level of explanation intensity, and the explanation duration is the fifth level of explanation duration.

[0082] Here, all comments related to users' likes / dislikes / acceptance of athletes, events, and related information appearing in the live stream are presented in an acceptance-oriented commentary style. All commentary for this role is delivered at the highest intensity level, using the fifth-highest commentary duration. The second type of commentary uses neutral, objective descriptive language, free from obvious emotional bias, and is suitable for the target characteristics of user acceptance. This second type of commentary emphasizes factual statements and avoids subjective evaluations to ensure the accuracy and fairness of information transmission. For example, when users view a match event with a neutral attitude, using the second type of commentary can guide users to focus on the event itself, thereby avoiding interference from emotional factors. In this embodiment, by adopting the second type of commentary, a strong expressive force can be maintained while maintaining a neutral stance, thus balancing users' emotional needs and information acquisition needs, thereby improving the diversity and adaptability of the overall viewing experience. In practice, combining the second type of explanation with the fifth level of explanation intensity can enhance users' attention to and understanding of relevant events by adopting a high-intensity language rhythm and expression style while maintaining a neutral stance.

[0083] In some embodiments, the third type of explanation strategy includes: when explaining the first type of target features and / or the second type of target features, the explanation form is the third type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration; when explaining the third type of target features, the explanation form is the first type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration.

[0084] Here, for athletes, events, and related information that users like during the live stream, the fifth level of negative commentary, with the fifth level of commentary duration, is used for both acknowledging and accepting the athletes, events, and related information. Conversely, for athletes, events, and related information that users dislike, the fifth level of positive commentary, with the fifth level of commentary duration, is used. The third type of commentary typically involves a strong negative tendency and tone when commenting on athletes, events, or related events that users like or accept; and an enthusiastic description of athletes, events, or related events that users dislike.

[0085] It's important to note that when initially creating multiple types of commentary characters, the fifth level of commentary intensity is used for the first and third types. That is, both favored and disliked commentary characters initially use the fifth level of intensity. This presents the most emotionally charged perspectives when users first listen, making it easier for them to decide whether to downgrade the intensity later. If users approve of the highest-intensity commentary character, it will consistently provide a full range of commentary, addressing both their favorite and disliked points. If users find the highest-intensity commentary character unacceptable, abandoning it will prompt a downgrade in intensity and duration, reducing user dissatisfaction.

[0086] In some embodiments, in response to a user's selection of any type of commentator role, the commentary strategy corresponding to any type of commentator role is updated; and the commentary content of the target video is played based on the updated commentary strategy.

[0087] In some embodiments, in response to a user's selection operation for any type of commentator role, the commentary strategy corresponding to any type of commentator role is updated, including: when the user's selection operation for any type of commentator role is a rejection operation, the commentary duration corresponding to any type of commentator role is reduced; wherein, when any type of commentator role is a first type of commentator role and / or a third type of commentator role, the commentary intensity corresponding to the first type of commentator role and / or the third type of commentator role is reduced.

[0088] Here, when a user enters the live stream, the intelligent commentary for the liked, disliked, and accepted commentary characters will all use the initial commentary strategy, which is 100% of the commentary time at the fifth level. The liked and disliked commentary characters will also continue with the fifth level of intensity, following the initial strategy. For the accepted commentary characters, the lowest intensity level (first level) is selected, and the highest duration level (fifth level) is selected. The user will then decide whether to keep or discard any of the three characters based on their viewing experience and the commentary. If a user chooses to keep any commentary character, the commentary time and intensity remain unchanged. It's worth noting that when a user discards a character and a new commentary character is generated, the change in intensity will be auditory enough to distinguish the different characters, allowing for a more subtle adjustment to the intensity of a fixed character.

[0089] When a user chooses to abandon any role, that role's commentary time will decrease by 20%, from 100% of the fifth-tier commentary time to 80% of the fourth-tier commentary time. The commentary intensity of liked and disliked roles will also decrease, from the fifth-tier intensity to the fourth-tier intensity. If a role is abandoned to the first-tier commentary time and the user chooses to abandon it again, that role will cease functioning and exit the multi-role intelligent commentary system. Acceptable commentary roles always maintain the first-tier intensity. When a user chooses to abandon a role, its commentary time only decreases by 20%, from the fifth-tier to the fourth-tier, while its intensity remains unchanged.

[0090] As can be seen from the above, the video commentary method provided by this application creates commentary rules of different levels for each dimension and analyzes the target historical matches watched by the user to determine the target that the user likes, dislikes, and accepts. When the user is watching the current live event video, if the current live event video is the target match, three roles are created for the user, and the commentary rules of each role are formulated by selecting the corresponding level of commentary rules from each dimension according to preset rules. The commentary is then provided for the corresponding target appearing in the current live event video according to the commentary rules of each role, so as to provide personalized live commentary for the user. Furthermore, for characters that users like or dislike, the strongest level is selected from each dimension when formulating the initial explanation rules. For characters that users accept, the lowest level is selected from the dimensions corresponding to explanation format and intensity, and the strongest level is selected from the dimension corresponding to explanation duration. After obtaining the user's target choice (such as keeping or discarding) for any of the three characters, the levels of the corresponding dimensions for each character are adjusted according to the target choice and preset adjustment rules to obtain the adjusted explanation rules for each character, thereby improving the user experience.

[0091] Figure 4 This is a schematic diagram of the video narration system provided in the embodiments of this application, such as... Figure 4As shown, to address the technical problem in existing technologies that use the same commentary template to provide commentary for users, which fails to meet the personalized needs of different users, a method for multi-role intelligent commentary based on user preferences is provided. This method creates different levels of commentary rules for each dimension (commentary style, commentary intensity, and commentary duration), and analyzes the historical matches watched by the user to determine the user's favorite, disliked, and accepted targets (such as athletes, events, and related information). When the user is watching the current live match video, if the current live match video is the target match, three roles are created for the user (such as a favorite role, a disliked role, and an accepted role). The commentary rules for each role are formulated by selecting the corresponding level of commentary rules from each dimension according to preset rules, and the commentary is provided for the corresponding target appearing in the current live match video according to the commentary rules of each role. For characters that are liked or disliked, the strongest level is selected from each dimension when formulating the initial explanation rules. For characters that are accepted, the lowest level is selected from the dimensions corresponding to explanation format and intensity, and the strongest level is selected from the dimension corresponding to explanation duration. After obtaining the user's target choice (e.g., keep or discard) for any of the three characters, the levels of the corresponding dimensions for each character are adjusted according to the target choice and preset adjustment rules to obtain the adjusted explanation rules for each character. The details are explained below: Step 1: Analyze user viewing preferences to identify target matches.

[0092] The system analyzes matches appearing in a user's viewing history to identify target matches. Analysis rule: Matches that a user continuously watches or repeatedly watches are defined as target matches.

[0093] Step 2: Analyze user behavior to obtain three results: liking, disliking, and accepting.

[0094] After identifying the target competition, the emotional biases exhibited by users regarding the competition, athletes, events, and related information they watched are analyzed. Based on user viewing behavior analysis, three results are derived: liking, disliking, and accepting. These results are used to define the evaluations of athletes, events, and related information.

[0095] 1) Analyze the athletes in the competition to identify those who are liked, disliked, or accepted.

[0096] The analysis rules are based on the athletes appearing in the target competition. If a user follows, likes, or comments positively on the athlete, it is considered that they like the athlete; if a user blocks, dislikes, or comments negatively on the athlete, it is considered that they dislike the athlete; if a user does not follow, block, like, dislike, or comment on the athlete, it is considered that they accept the athlete.

[0097] 2) Users analyze their attitudes toward events involving athletes to identify events they like, events they dislike, and events they accept.

[0098] The analysis rules govern events created by athletes in a target competition, primarily involving competitive events, scoring events, outcome determination events, violations, spectacular athletic moves, record-breaking events, injury events, interruption of the match, and withdrawal events under the competition rules. When these events occur, positive user actions such as liking, commenting, and sending bullet comments are considered "like" events; negative user actions such as disliking, commenting, and sending bullet comments are considered "dislike" events; and the absence of user actions such as liking, disliking, commenting, or sending bullet comments is considered "acceptance" events.

[0099] 3) Analyze users' attitudes toward relevant information in the competition to obtain information that they like, dislike, or accept.

[0100] The analysis rules define relevant information in the target competition as including coaches, competition venues, participating countries, athlete matchups, local climate, and other related information. When users see the above information, actions such as following, liking, commenting, and sending positive feedback in the bullet comments are considered as liking the relevant information; actions such as blocking, disliking, commenting, and sending negative feedback in the bullet comments are considered as disliking the relevant information; and no actions such as following, blocking, liking, disliking, commenting, or sending bullet comments are considered as accepting the relevant information.

[0101] Step 3: Generate a multi-role intelligent commentary module, including commentary format, commentary intensity, and commentary duration.

[0102] Based on the three results obtained in Step 2—likes, dislikes, and acceptance—the intelligent commentary module will generate a multi-role intelligent commentary module based on user preferences. This module represents the specific method for multi-role intelligent commentary during subsequent live broadcasts, including settings for commentary style, intensity, and duration, with corresponding levels for each of these dimensions.

[0103] 1) Generation of narration format: Commentary style refers to the generation of commentary styles based on the competition, athletes, events, and related information that occur during a live broadcast. People either like, dislike, or accept commentary styles.

[0104] Preferred commentary style: The commentary on the game, athletes, events and related information is presented in a detailed, comprehensive manner, combining praise and encouragement with a combination of outlook and review, delivered at a positive, emotionally charged, and expressive pace. Dislike of commentary style: Commentary on matches, athletes, events and related information is presented in a restrained, localized critical and questioning manner and narrow descriptive language, delivered at a negative, depressed and emotionally suppressed pace; Acceptance of commentary format: The commentary on the game, athletes, events and related information is presented in objective and neutral descriptive language, faithful to all objective information of the game, without extension or extrapolation, and describes the true situation of everything seen on the live broadcast through language.

[0105] 2) Explanation of intensity generation: The intensity of narration involves variations across seven dimensions: tone, speaking speed, volume, pauses, vocabulary, interjections, and quotations. It is divided into five intensity levels, ranging from 1 to 5, with 1 being the weakest and 5 the strongest. See Table 1 above for details.

[0106] 3) Generation of narration duration: Commentary duration refers to the time during a live broadcast when the multi-role intelligent commentary starts speaking when a user's favorite, disliked, or accepted athlete, event, or related information appears, and when the multi-role intelligent commentary stops speaking when the next user's favorite, disliked, or accepted athlete, event, or related information appears. The time from when the commentary starts to when it stops is the commentary duration.

[0107] The commentary time will be evenly distributed based on the number of commentator roles. Each commentator will initially dedicate 100% of their allocated commentary time to their work. The commentary time is divided into five tiers, each representing 20% ​​of the total work time. See Table 2 above for details.

[0108] Step 4: Create multiple intelligent commentary characters in the initial state, and obtain the liked character, disliked character, and accepted character.

[0109] Once the live stream begins, users enter to watch. The system identifies whether the match being streamed is the target match defined in the first step of analysis. If so, it begins creating multiple intelligent commentary roles. Roles are categorized into three types: preferred commentary roles, disliked commentary roles, and accepting commentary roles.

[0110] A "favorite" commentator uses a strong liking and tone to comment on athletes, events, and related information that users like and accept, and a strong dislike and tone to comment on athletes, events, and related information that users dislike. This represents a strong attempt to cater to users, with a clear stance of liking and disliking.

[0111] An "anti-fan" commentator uses a strong anti-fan attitude and tone when commenting on athletes, events, and related information that users like and accept, and a strong liking attitude and tone when commenting on athletes, events, and related information that users dislike. This clearly represents a stance of strongly opposing what users like while strongly supporting what users dislike.

[0112] An accepting commentator role refers to commentary on athletes, events, and related information that users like, accept, or dislike, and strives for objectivity, truthfulness, and impartiality.

[0113] The simultaneous appearance of three commentary roles in a live stream creates a dynamic interplay of opposing, conflicting, and mutually disagreeing personas among the intelligent commentators. This comprehensively caters to a wide range of user preferences and aversions. This makes the value orientations of the various roles clearly defined, more closely resembling the realistic commentary scenarios where multiple human commentators clash and exchange ideas.

[0114] Favorite commentator role: For athletes, events, and related information that users like during the live stream, adopt a favorable commentary style with a strength of 5 and use the fifth-level commentary duration. For athletes, events, and related information that users dislike, adopt a dislike commentary style with a strength of 5 and use the fifth-level commentary duration.

[0115] Disliked Commentator Role: For athletes, events, and related information that users like during the live stream, accept the athletes, events, and related information, and use a disliked commentary style with a intensity of 5, using the fifth level of commentary duration. For athletes, events, and related information that users dislike, use a liked commentary style with an intensity of 5, using the fifth level of commentary duration.

[0116] Acceptance-oriented commentator role: This role uses an acceptance-oriented commentary style for all comments made during the live stream that relate to user preferences / dislikes / acceptance of athletes, events, or related information. All commentary for this role is at intensity level 1 and uses the fifth-highest commentary duration.

[0117] Both the "liked" and "disliked" commentary characters are initially created with a commentary intensity of 5. This presents the most emotionally charged perspective when users first listen, making it easier for them to decide whether to downgrade the character later. If a user approves of the highest-intensity commentary character, it will consistently provide commentary on both their preferred and disliked content. If a user finds the highest-intensity commentary character unacceptable, abandoning it will cause the character's commentary intensity and duration to decrease, reducing user dissatisfaction.

[0118] Step 5: After the user retains or discards a role, the multi-role intelligent narration is generated again.

[0119] Once the live stream begins and users enter the live stream, the intelligent commentary for the liked, disliked, and accepted characters will all work using the initial state level 5, 100% of the commentary duration. Among them, the intelligent commentary for the liked and disliked characters will also follow the commentary format of the initial state intensity 5.

[0120] At this point, users will decide whether to keep or discard any of the three characters based on their experience watching the live stream and listening to the commentary. If a user chooses to keep any character, the initial state of 100% commentary duration and the commentary format of initial state strength 5 will remain unchanged.

[0121] When a user chooses to discard any character, that character's commentary time will decrease by 20%, from 100% in the fifth tier to 80% in the fourth tier. The commentary strength of liked and disliked characters will also decrease, from strength 5 to strength 4. Once a character's commentary time has been reduced to the first tier, discarding it again will cause that character to stop working and exit the multi-character intelligent commentary mode.

[0122] The accepting commentator role always uses a commentary intensity of 1. When a user chooses to abandon the role, the commentary time is reduced by only 20%, that is, from 100% in the fifth level to 80% in the fourth level. The commentary intensity remains unchanged.

[0123] Step 6: Use multi-role intelligent commentary to present and interpret the commentary of the live match.

[0124] After completing the user behavior analysis, the target competition and the athletes, events, and related information that the participants liked, disliked, or accepted during the competition were also obtained. Simultaneously, the live broadcast content of the sports competitions mentioned in the plan was analyzed to generate commentary on the objective information conveyed by the footage, such as athletes and events. Combined with publicly available competition data from the organizing committee or professional event management organization, commentary was generated based on the objective and authentic data of the athletes, events, and related information obtained above.

[0125] The text explores the handling of commentary scripts based on three commentary styles: liking, disliking, and accepting. It uses emotional expression to establish a stance of liking, disliking, or neutrality regarding the athletes and their performances as portrayed in the commentary. The commentary is processed emotionally, including mood and vocal delivery, to achieve a clear auditory distinction between the commentators.

[0126] Narration intensity refers to the detail in the tone, speed, volume, pauses, vocabulary, interjections, and quotations of the narration text when generating audio. As described in Step 3, the five intensities represent different detail processing schemes. Based on the tone established by the three narration formats mentioned in the previous paragraph, variations in tone, speed, volume, pauses, vocabulary, interjections, and quotations create variations in intensity, achieving character differentiation where each character has a different performance style. When the user discards the narration in Step 5 and generates a new narration character, the changes in narration intensity allow for auditory differentiation of the different characters, enabling the user to perceive subtle adjustments in the narration intensity of a fixed character.

[0127] The commentary duration is determined after the commentator's role and corresponding commentary strength are activated. It is calculated by allocating the commentary time between two commentary points to different commentator roles, thus ensuring each role has their designated commentary time. This system ensures the commentary is delivered within the allocated time. Since there are three scenarios regarding the relationship between commentary time and the commentary: the commentary time is sufficient to finish the commentary, the commentary time is insufficient to finish the commentary, and the commentary time is sufficient with remaining time, the presentation of the commentary focuses on the calculation and definition of each commentator's allotted time within the plan, rather than requiring the commentary to be delivered within the allotted time.

[0128] The following description continues to illustrate the exemplary structure of the video narration device 433 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the video narration device 433 in the memory 430 may include: The data acquisition module 4331 is used to acquire the target video and the user behavior operation corresponding to the target video.

[0129] The data processing module 4332 is used to analyze the target features of the target video based on the user behavior operation to obtain the target feature type; in response to the user's viewing operation on the target video, create multiple types of narration roles based on the target feature type; determine the narration strategy corresponding to the multiple types of narration roles, and play the narration content of the target video based on the narration strategy.

[0130] In some embodiments, the target feature type includes at least one of the following: a first type of target feature, a second type of target feature, and a third type of target feature; the user preference level corresponding to the first type of target feature is higher than the user preference level corresponding to the second type of target feature, and the user preference level corresponding to the second type of target feature is higher than the user preference level corresponding to the third type of target feature.

[0131] In some embodiments, the data processing module 4332 is further configured to: obtain the target feature type as the first type of target feature when the user behavior operation is to perform a positive operation on the target feature of the target video; obtain the target feature type as the second type of target feature when the user behavior operation is not to perform any operation on the target feature of the target video; and obtain the target feature type as the third type of target feature when the user behavior operation is to perform a negative operation on the target feature of the target video.

[0132] In some embodiments, the data processing module 4332 is further configured to determine an explanation strategy based on the target feature and the target feature type; wherein the explanation strategy includes at least one of the following: explanation form, explanation intensity, and explanation duration.

[0133] In some embodiments, the types of the multiple types of commentary roles include at least one of the following: a first type of commentary role, a second type of commentary role, and a third type of commentary role.

[0134] In some embodiments, the data processing module 4332 is further configured to determine that the explanation strategy corresponding to the first type of explanation role is a first type of explanation strategy; wherein, the first type of explanation strategy includes: when explaining the first type of target feature / or the second type of target feature, the explanation form is a first type of explanation form, the explanation intensity is a fifth level of explanation intensity, and the explanation duration is a fifth level of explanation duration; when explaining the third type of target feature, the explanation form is a third type of explanation form, the explanation intensity is a fifth level of explanation intensity, and the explanation duration is a fifth level of explanation duration.

[0135] In some embodiments, the data processing module 4332 is further configured to determine that the explanation strategy corresponding to the second type of explanation role is a second type of explanation strategy; wherein, the second type of explanation strategy includes: when explaining the second type of target features, the explanation form is a second type of explanation form, the explanation intensity is a first level of explanation intensity, and the explanation duration is a fifth level of explanation duration.

[0136] In some embodiments, the data processing module 4332 is further configured to determine that the explanation strategy corresponding to the third type of explanation role is a third type of explanation strategy; wherein, the third type of explanation strategy includes: when explaining the first type of target features and / or the second type of target features, the explanation form is a third type of explanation form, the explanation intensity is a fifth-level explanation intensity, and the explanation duration is a fifth-level explanation duration; when explaining the third type of target features, the explanation form is a first type of explanation form, the explanation intensity is a fifth-level explanation intensity, and the explanation duration is a fifth-level explanation duration.

[0137] In some embodiments, the data processing module 4332 is further configured to, in response to the user's selection operation for any type of commentary role, update the commentary strategy corresponding to the any type of commentary role; and play the commentary content of the target video based on the updated commentary strategy.

[0138] In some embodiments, the data processing module 4332 is further configured to reduce the commentary duration corresponding to any type of commentary role when the user's selection operation for any type of commentary role is a discard operation; wherein, when the any type of commentary role is a first type of commentary role and / or a third type of commentary role, the commentary intensity corresponding to the first type of commentary role and / or the third type of commentary role is reduced.

[0139] In some embodiments, the target features include at least one of the following: target object features, target event features, and related information features.

[0140] In some embodiments, the data acquisition module 4331 is further configured to acquire data information of the target video.

[0141] In some embodiments, the data processing module 4332 is further configured to generate narration content for the target video based on the data information of the target video, the characteristics of the target object, the characteristics of the target event, and the related information characteristics.

[0142] This application provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the video narration method described in this application.

[0143] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the video narration method provided in this application.

[0144] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0145] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0146] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0147] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0148] In summary, this application's embodiments classify target features of target videos based on user behavior, identifying the user's degree of liking, disliking, or accepting of specific features (such as athletes, events, or related information). Secondly, based on the classification results, multiple types of commentary roles with different stances (such as liking, disliking, and accepting) are created. Then, a commentary strategy matching each role's stance is formulated, including dimensions such as commentary style, intensity, and duration. Finally, the commentary strategy is dynamically adjusted during user viewing, supporting user selection and feedback on roles, achieving a personalized and highly interactive video commentary experience. This entire mechanism effectively simulates the clash of opposing stances and viewpoints between real commentators, enhancing the emotional expression of the commentary content and user participation, thereby overcoming problems such as single-voice broadcasting, lack of stance differentiation, and lack of interaction in traditional intelligent commentary.

[0149] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A video narration method, characterized in that, The method includes: Obtain the target video and the corresponding user behavior operations; The target features of the target video are analyzed based on the user behavior to obtain the target feature type; In response to a user's viewing action on the target video, multiple types of narration roles are created based on the target feature type; Determine the commentary strategies corresponding to the various commentary roles, and play the commentary content of the target video based on the commentary strategies.

2. The method according to claim 1, characterized in that, The target feature type includes at least one of the following: a first type of target feature, a second type of target feature, and a third type of target feature; the user preference level corresponding to the first type of target feature is higher than that corresponding to the second type of target feature, and the user preference level corresponding to the second type of target feature is higher than that corresponding to the third type of target feature; wherein... The analysis of target features of the target video based on the user behavior operation to obtain target features includes: When the user action is to perform a positive operation on the target feature of the target video, the target feature type is the first type of target feature; When the user action is to not operate on the target features of the target video, the target feature type is the second type of target feature; When the user action is to perform a negative operation on the target feature of the target video, the target feature type is the third type of target feature.

3. The method according to claim 2, characterized in that, The method further includes: Based on the target features and the target feature types, an explanation strategy is determined; wherein the explanation strategy includes at least one of the following: explanation form, explanation intensity, and explanation duration.

4. The method according to claim 1, characterized in that, The types of the multiple commentary roles include at least one of the following: a first type of commentary role, a second type of commentary role, and a third type of commentary role; The determination of the commentary strategies corresponding to the multiple types of commentary roles includes: The explanation strategy corresponding to the first type of explanation role is determined as the first type of explanation strategy; and / or, the explanation strategy corresponding to the second type of explanation role is determined as the second type of explanation strategy; and / or, the explanation strategy corresponding to the third type of explanation role is determined as the third type of explanation strategy; wherein, The first type of explanation strategy includes: when explaining the first type of target features and / or the second type of target features, the explanation form is the first type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration; when explaining the third type of target features, the explanation form is the third type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration; The second type of explanation strategy includes: when explaining the second type of target features, the explanation form is the second type of explanation form, the explanation intensity is the first level of explanation intensity, and the explanation duration is the fifth level of explanation duration; The third type of explanation strategy includes: when explaining the first type of target features and / or the second type of target features, the explanation form is the third type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration; when explaining the third type of target features, the explanation form is the first type of explanation form, the explanation intensity is the fifth level of explanation intensity, and the explanation duration is the fifth level of explanation duration.

5. The method according to claim 1, characterized in that, The method further includes: In response to the user's selection of any type of commentary role, update the commentary strategy corresponding to that type of commentary role; The narration content of the target video is played based on the updated narration strategy.

6. The method according to claim 5, characterized in that, The step of updating the commentary strategy corresponding to any type of commentary role in response to the user's selection operation includes: When the user selects any type of commentary role as an abandonment operation, the commentary duration corresponding to that type of commentary role is reduced; wherein, when the any type of commentary role is a first type of commentary role and / or a third type of commentary role, the commentary intensity corresponding to the first type of commentary role and / or the third type of commentary role is reduced.

7. The method according to any one of claims 1 to 6, characterized in that, The target features include at least one of the following: target object features, target event features, and related information features; before playing the narration content of the target video based on the narration strategy, the method further includes: Obtain the data information of the target video; Based on the data information of the target video, the characteristics of the target object, the characteristics of the target event, and the relevant information characteristics, the narration content of the target video is generated.

8. An electronic device, characterized in that, include: A processor and a memory for storing a computer program capable of running on the processor, wherein the processor, when running the computer program, performs the steps of the method according to any one of claims 1 to 7.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.